A compression system for engineering information data

By analyzing and screening the duplicate substrings in engineering information, grouping and arranging, and combining the compression principle of run coding, the problem that traditional technology cannot effectively compress and share engineering information is solved, and efficient data compression and sharing is achieved.

CN119788089BActive Publication Date: 2025-05-30JIANGXI PROVINCIAL EXPRESSWAY INVESTMENT GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510286093.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-05-30
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

Traditional run coding cannot effectively reduce storage space and improve information sharing efficiency when processing engineering information, especially when there is a large number of dispersed and repeated content.

Method used

By obtaining the repeated substrings in the engineering information, analyzing their distribution characteristics, filtering out substrings with better uniformity and length, grouping and arranging, generating the initial compression sequence, and obtaining the final compression sequence through shift adjustment, adapting to the compression principle of run-code.

Benefits of technology

It improves the compression rate of engineering information, reduces storage space usage, improves the speed and efficiency of information sharing, and avoids the increase in additional data volume.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119788089B_ABST
    Figure CN119788089B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of engineering information data compression processing, and specifically relates to a compression system for engineering information data. The system includes: a repeated substring acquisition module for acquiring different types of repeated substrings; a filtered substring acquisition module for filtering the repeated substrings to obtain reference substrings, then calculating the filtering values of the reference substrings, and obtaining filtered substrings; an optimal grouping length acquisition module for dividing the string to be processed using each grouping length within the grouping length range, and then obtaining the optimal grouping length; a compressed sequence acquisition module for dividing the string to be processed using the optimal grouping length to obtain character substrings; placing other character substrings based on the target character substring, and performing shift adjustment on the characters in the character substrings to obtain a compressed sequence; an information sharing module for compressing the compressed sequence using run-length encoding and transmitting the compressed data to the sharing party. This application can improve the sharing efficiency of engineering information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of engineering information data compression processing, and particularly relates to a compression system for engineering information data. Background Art

[0002] With the rapid development of the construction industry, the scale and complexity of engineering projects are increasing day by day. The traditional engineering information management method has been unable to meet the needs of modern engineering management. Especially in large-scale engineering projects, involving multiple parties, multiple data types and complex business processes, how to achieve effective sharing and management of engineering information has become an urgent problem to be solved.

[0003] Due to the large amount of engineering information content and a large number of repeated text information, the traditional run-length encoding is used for data compression processing. The run-length encoding is lossless compression, which can ensure the accurate consistency of engineering information. However, for data with low repetition, it often increases the amount of data additionally. Although there is a large amount of repeated content in engineering information, these repeated contents are relatively scattered. Therefore, directly using run-length encoding for compression will result in an increase in additional data volume, and cannot well reduce the storage space, and the data packet is too large during transmission, thereby affecting the efficiency of information sharing. Summary of the Invention

[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a compression system for engineering information data, and the specific technical solutions adopted are as follows:

[0005] An embodiment of the present invention provides a compression system for engineering information data, and the system includes:

[0006] A repeated substring acquisition module, configured to convert engineering information into a string to be processed; obtain the occurrence frequencies of different characters in the string to be processed, and obtain different types of repeated substrings according to the occurrence frequencies and the intervals between characters;

[0007] A screened substring acquisition module, configured to obtain the uniformity of each type of repeated substring according to the variance of the interval distances between the repeated substrings in each type of repeated substring and the length of each type of repeated substring; screen the repeated substrings to obtain reference substrings; calculate the screening values of the reference substrings according to the uniformity and length of the reference substrings, and screen the reference substrings according to the screening values to obtain screened substrings;

[0008] An optimal grouping length acquisition module, configured to obtain a grouping length range according to the minimum value and the maximum value of the interval distances between every two screened substrings in each type of screened substrings; divide the string to be processed by using each grouping length within the grouping length range, and obtain the optimal grouping length based on the division result;

[0009] The compression sequence acquisition module is used to split the string to be processed using the optimal grouping length to obtain character substrings; use the first character substring where each repeated substring appears as the target character substring; arrange the target character substrings in order to obtain a reference line; arrange the other character substrings in the columns corresponding to the target character substrings in the reference line respectively according to the presence of repeated substrings in the other character substrings to obtain an initial compression sequence; perform a shift adjustment on the characters in each character substring in the initial compression sequence to obtain a compression sequence.

[0010] The information sharing module is used to compress the compression sequence using run-length encoding and transmit the compressed data to the sharing party.

[0011] Preferably, different types of repeated substrings are obtained according to the occurrence frequency and the interval between characters, including:

[0012] Sort the characters according to the occurrence frequency of each character in the string to be processed to obtain a character sequence, and there is only one character of each type in the character sequence; record the first character in the character sequence as the first character; obtain the interval distance between every two first characters in the string to be processed. If the interval distance is greater than or equal to the preset value, select the two first characters with the smallest interval distance and appearing at the front of the string to be processed, and record them as the first type of characters. Use the first character in the first type of characters and the characters between the two first characters to form a matching template, and find the character substring identical to the matching template in the string to be processed to obtain the same type of repeated substring, which is recorded as the first type of repeated substring; remove the first type of repeated substring from the string to be processed to obtain the first remaining string to be processed.

[0013] Find two first characters with the same interval distance as between the two first characters of the first type of characters but different interval characters and located at the front of the first remaining string to be processed, and record them as the second type of characters; use the first character in the second type of characters and the characters between the two first characters in the second type of characters to form a second matching template, and find the character substring identical to the matching template in the first remaining string to be processed to obtain the same type of repeated substring, which is recorded as the second type of repeated substring; remove the second type of repeated substring from the first remaining string to be processed to obtain the second remaining string to be processed.

[0014] After all types of repeated substrings corresponding to the two first characters with the smallest interval distance are found, and so on, find all types of repeated substrings corresponding to the two first characters with the second smallest interval distance until all types of repeated substrings corresponding to the first character are found; and so on, according to the order of the characters in the character sequence, find different types of repeated substrings corresponding to each character in the character sequence.

[0015] Preferably, the formula for uniformity is:

[0016] ,

[0017] Among them, represents the uniformity of the J-th type of repeated substring; represents the variance of the interval distance between repeated substrings in the J-th type of repeated substring; represents the number of the J-th type of repeated substrings; n represents the number of characters in the string to be processed; represents the length of the J-th type of repeated substring; represents rounding down.

[0018] Preferably, screening the repeated substrings to obtain reference substrings, including:

[0019] The ratio of the number of each type of repeated substring to the number of all repeated substrings is the repetition frequency of each type of repeated substring; the repetition threshold is the reciprocal of the number of types of repeated substrings. If the repetition frequency of a type of repeated substring is greater than the repetition threshold, then this type of repeated substring is a reference substring.

[0020] Preferably, calculating the screening value of the reference substring according to the uniformity and length of the reference substring, and screening the reference substring according to the screening value to obtain the screened substring, including:

[0021] Obtain the reciprocal of the length of the reference substring, and add it to the uniformity of the reference substring to get the addition result; normalize the addition result to obtain the screening value of the reference substring; set a screening threshold, and the reference substring with a screening value greater than the screening threshold is the screened substring.

[0022] Preferably, obtaining the grouping length range according to the minimum and maximum values of the interval distance between every two screened substrings in each type of screened substring, including:

[0023] Respectively obtain the minimum value of the interval distance between every two screened substrings in each type of screened substring, and form a minimum interval distance sequence after adding each minimum interval distance value to the first preset value; respectively obtain the maximum value of the interval distance between every two screened substrings in each type of screened substring, and form a maximum interval distance sequence after adding each maximum interval distance value to the first preset value; the least common multiple of each value in the minimum interval distance sequence is the lower limit value of the grouping length range, and the least common multiple of each value in the maximum interval distance sequence is the upper limit value of the grouping length range.

[0024] Preferably, dividing the string to be processed by using each grouping length within the grouping length range, and obtaining the optimal grouping length based on the division result, including:

[0025] Take values at a preset step length within the grouping length range starting from the lower limit value of the grouping length range until the upper limit value of the grouping length range according to the preset step length to obtain different grouping lengths; obtain the average value of the ratio of the number of reference substrings of each type in the division result corresponding to a grouping length to the original number of reference substrings of each type, and record it as the division completeness rate corresponding to this grouping length. Select the grouping length with the largest division completeness rate and record it as the optimal grouping length.

[0026] Preferably, arrange the target character substrings in order to obtain a reference line, including:

[0027] Record the number of characters in the repeated substrings in each target character substring as the repetition rate of the target character substring; arrange the target character substrings in descending order of the repetition rate to obtain a reference line.

[0028] Preferably, obtain an initial compression sequence, including:

[0029] Place the character substrings containing the same repeated substrings in the columns corresponding to the target character substrings according to their order in the string to be processed. After completion, obtain the unplaced character substrings and form an unplaced substring set;

[0030] Obtain the similarity between each character substring in the unplaced substring set and each target character substring. The similarity is the number of identical letters between the character substring and the target character substring. Place each character substring in the unplaced substring set in the column corresponding to the target character substring with the highest similarity, and the order in the row direction during placement is determined by the order of each character substring in the unplaced substring set in the string to be processed;

[0031] Place the unplaced character substrings in the unplaced substring set according to their order of appearance in the string to be processed to obtain an initial compression sequence.

[0032] Preferably, perform a shift adjustment on the characters in each character substring in the initial compression sequence to obtain a compression sequence, including:

[0033] Calculate the similarity between the current character substring and the character substring in the first row or the previous row of its column respectively, and select the character substring with a greater similarity in the character substring of the first row or the previous row as the adjustment target; obtain the string with the same characters in the current character substring and the adjustment target, and record it as the same string; if the characters at the same positions of the same string corresponding to the current character substring and the same string corresponding to the adjustment target are all the same, the current character substring can be shifted, and if they are not the same, it cannot be shifted; obtain the adjustment efficiency of the character substring that can be shifted according to the number of shift operations when the character substring that can be shifted is shifted and the number of characters of the character substring that can be shifted; if the adjustment efficiency is greater than or equal to the adjustment threshold, shift the character substring that can be shifted.

[0034] The calculation formula of the adjustment efficiency is:

[0035] ,

[0036] where, δ represents the adjustment efficiency of the character substring that can be shifted; R represents the number of shift operations when the character substring that can be shifted is shifted; max() represents the operation of taking the maximum value; represents the number of overlapping characters in the character substring that can be shifted and the adjustment target; represents the number of overlapping characters in the character substring that can be shifted after shifting and the adjustment target; Q represents the number of characters in the character substring that can be shifted.

[0037] The embodiments of the present invention have at least the following beneficial effects: In this application, engineering information is converted into a string to be processed, and duplicate substrings in the string to be processed are obtained; the distribution characteristics of the duplicate substrings in the string to be processed are used for analysis to obtain the uniformity of each type of duplicate substring, and then the reference substrings are obtained through screening, and the screened substrings are further screened to obtain the screened substrings; furthermore, the interval distances between the screened substrings are analyzed to obtain the most suitable grouping length, that is, the optimal grouping length, and then the string to be processed is segmented to obtain character substrings, which can ensure that there are more consecutive repeated characters in the sequence obtained after such segmentation and adjustment, thereby improving the compression ratio; further, a reference line composed of target character substrings is obtained, and then other character substrings are placed based on the reference line to obtain an initial compression sequence, and then the characters in each character substring in the initial compression sequence are shifted and adjusted to obtain a compression sequence, so that the repeatability of the data to be compressed is increased, which can better adapt to the compression principle of run-length encoding, avoid additional data volume, and obtain less compression space as much as possible under the condition of not losing information, realizing the purpose of lossless and less compression space occupation and higher compression ratio, thereby improving the speed of engineering information sharing and increasing the efficiency of engineering information sharing. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0039] Figure 1 It is a system block diagram of a compression system for engineering information data provided by an embodiment of the present invention;

[0040] Figure 2 It is an example diagram of possible duplicate strings of a compression system for engineering information data provided by an embodiment of the present invention;

[0041] Figure 3 It is a schematic diagram of character substrings of a compression system for engineering information data provided by an embodiment of the present invention;

[0042] Figure 4 It is a schematic diagram of a reference line of a compression system for engineering information data provided by an embodiment of the present invention;

[0043] Figure 5 It is a schematic diagram of a first sequence of a compression system for engineering information data provided by an embodiment of the present invention;

[0044] Figure 6The second sequence schematic diagram of a compression system for engineering information data provided by an embodiment of the present invention;

[0045] Figure 7 The initial compression sequence schematic diagram of a compression system for engineering information data provided by an embodiment of the present invention;

[0046] Figure 8 The shift operation schematic diagram of a compression system for engineering information data provided by an embodiment of the present invention. Detailed implementation manners

[0047] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following combines the accompanying drawings and preferred embodiments to detail the specific scheme, specific implementation manners, structures, features and effects of a compression system for engineering information data proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0049] The following specifically describes the specific scheme of a compression system for engineering information data provided by the present invention with reference to the accompanying drawings.

[0050] Embodiment:

[0051] The main application scenario of the present invention is as follows: When using run-length encoding to compress engineering information to be shared, due to the large number of repeated fields in text information but a relatively high probability of discrete distribution, the compression effect of run-length encoding may not reach the desired effect, and compression cannot be directly used, thereby affecting the efficiency of information sharing and reducing the transmission efficiency during engineering information sharing.

[0052] Please refer to Figure 1 , which shows a block diagram of a compression system for engineering information data provided by an embodiment of the present invention. The system includes the following modules:

[0053] A repeated substring acquisition module, configured to convert engineering information into a string to be processed; obtain the occurrence frequencies of different characters in the string to be processed, and obtain different types of repeated substrings according to the occurrence frequencies and the intervals between characters.

[0054] The main purpose of this application is to sort and transform the character string converted from the text information according to the distribution of repeated fields in the text information, increase the probability of continuous repetition of characters, so as to better use run-length encoding for compression. First, it is necessary to read the engineering information file to be shared from the database, convert the Chinese part in the engineering information file into English, and use the converted file content as the character string to be processed, wherein after conversion to English, punctuation marks and numbers are also characters, thereby obtaining the character string to be processed. The text information in this application is compressed and transmitted to the sharing party, which can improve the efficiency of data transmission, thereby improving the efficiency of data sharing.

[0055] Since there are repeated contents in the read engineering information, but the repeated contents may not appear continuously, but more often appear scattered, which is not conducive to the use of run-length encoding, it is necessary to sort and adjust the converted character string to be processed. The adjustment principle is to group the character substrings according to the repetitive characteristics, and arrange the substrings in similar combinations as much as possible within a small range. Therefore, it is necessary to first find the repeated substrings in the character string to be processed converted from the engineering information.

[0056] Based on the above analysis, the occurrence frequencies of different characters in the string to be processed are obtained, and different types of repeated substrings are obtained according to the occurrence frequencies and the intervals between characters. Specifically, the process of obtaining each type of repeated substring is as follows:

[0057] First, count the occurrence frequency of each character in the character string to be processed, and sort the characters from large to small to obtain a character sequence, in which there is only one character of each type in the character sequence; record the first character in the character sequence as the first character; obtain the interval distance between every two first characters in the character string to be processed, if the interval distance is greater than or equal to a preset value, select the two first characters with the smallest interval distance and appearing at the front of the character string to be processed, and record them as first type characters, and use the first character of the first type of characters and the characters between the two first characters to form a matching template, and search for a character substring identical to the matching template in the character string to be processed to obtain a repeated substring of the same type, which is recorded as a first type repeated substring; remove the first type repeated substring in the character string to be processed to obtain a first remaining character string to be processed;

[0058] Find two first characters that have the same spacing distance as the first type of characters but different spacing characters and are located at the front of the first remaining character string to be processed, and record them as second type of characters; use the first character of the second type of characters and the characters between the two first characters of the second type of characters to form a second matching template, find a character substring that is the same as the matching template in the first remaining character string to be processed, and obtain a repeated substring of the same type, which is recorded as a second type of repeated substring; remove the second type of repeated substring in the first remaining character string to be processed, and obtain a second remaining character string to be processed;

[0059] After all types of repeated substrings corresponding to the two first characters with the smallest interval distance are found, search for all types of repeated substrings corresponding to the two first characters with the second smallest interval distance until all types of repeated substrings corresponding to the first characters are found; and so on, in the order of the characters in the character sequence, search for different types of repeated substrings corresponding to each character in the character sequence.

[0060] For example, if the character with the highest frequency is the A character, that is, the first character is A, obtain the interval distance between A characters in the string to be processed. The interval distance is the number of characters between the intervals of two characters. If there is an interval distance greater than or equal to the preset value of 2, the position where this interval appears may be the position where the combination of A and adjacent characters repeats. Find the positions of A with the same interval, and the manifestation is as Figure 2 shown Figure 2 In, 2 and 6 represent the interval distance between two characters. The small box is the currently searched character, that is, the first character, and the large box is the possible repeated combination within two first characters with the same interval distance.

[0061] Such as Figure 2 , the interval distance between two first characters is 2. When this interval appears for the first time, the characters in the searched character are BC, then the matching template is ABC. Search for character substrings in the string to be processed that are the same as the matching template. These character substrings and the matching template are repeated substrings of the same type, and the successfully matched character substrings are recorded as repeated substrings , i represents the i-th repeated substring 1. The search for the remaining repeated substrings with the same interval for the A character is the same. The already marked characters are no longer matched, that is, after finding a type of repeated substring, it needs to be removed from the string to be processed, and then the remaining string to be processed is used to find the next type of repeated substring. For example, the first remaining string to be processed and the second remaining string to be processed obtained after the above removal.

[0062] After finishing searching for the A character, start repeating the above steps with the character with the second highest frequency, that is, the second character in the character sequence, such as the B character, until no repeated character substrings can be found. Obtain the set of repeated substrings {repeated substring (J = 1, 2, 3…N, i = 1, 2, 3…M)}, where J represents the type of repeated substring, i represents the i-th repeated substring of the same type, N represents the number of types of repeated substrings, and M represents the number of repeated substrings of the same type.

[0063] It should be noted that this method can only find most of the repeated character substrings in the string to be processed and cannot ensure finding all repeated substrings, but most repeated substrings are statistically already of reference significance.

[0064] The substring screening acquisition module is used to obtain the uniformity of each repeated substring according to the variance of the interval distance between repeated substrings and the length of each repeated substring in each type of repeated substring; screen the repeated substrings to obtain reference substrings; calculate the screening value of the reference substrings according to the uniformity and length of the reference substrings, and screen the reference substrings according to the screening value to obtain the screened substrings.

[0065] The distribution uniformity of the repeated substrings can to a certain extent reflect the repeated regions in the string, the positions and substring sizes where repeated changes are more likely to occur. Therefore, the uniformity of each repeated substring is obtained according to the variance of the interval distance between repeated substrings and the length of each repeated substring in each type of repeated substring. Specifically, the calculation formula for the uniformity of the repeated substring is:

[0066] ,

[0067] where, represents the uniformity of the Jth type of repeated substring; represents the variance of the interval distance between repeated substrings in the Jth type of repeated substring; represents the number of the Jth type of repeated substring; n represents the number of characters in the string to be processed; represents the length of the Jth type of repeated substring; represents rounding down.

[0068] The smaller the variance, the closer the interval data values of the Jth type of repeated substring are, that is, the more uniform the distribution of the Jth type of repeated substring; represents the ratio of the number of the Jth type of repeated substring to the number of parts after the string is divided. The larger this ratio, the greater the distribution density of the Jth type of repeated substring. The larger it is, the more uniform the distribution of the Jth type of repeated substring in the string to be processed and the greater the distribution density, and the better the uniformity.

[0069] The greater the repetition frequency of the repeated substring, the greater the probability that the characters in the string to be processed form this type of repeated substring. Therefore, it is necessary to use the repeated substring with a relatively large repetition probability as the reference substring. Since when the combination length is short, there are fewer characters to choose from and the total number of possible combinations is also small, it is easier to have repeated combinations under multiple combinations of the same length. At the same time, if a certain type of reference substring is more uniformly distributed in the entire string, it means that this type of repeated substring is more likely to appear in various places in the string to be processed. When splitting the string to be processed, the possibility of splitting these repeated substrings is lower. Therefore, based on the above reasons, the screening value of the reference substring can be calculated, and the intervals of multiple reference substrings with a screening value greater than the screening threshold are used as the string grouping range. The optimal grouped character substring length is obtained by calculating the integrity rate of whether the reference substring is not split after different numbers of groupings.

[0070] Screen the repeated substrings to obtain reference substrings. Specifically, the ratio of the number of each type of repeated substring to the number of all repeated substrings is the repetition frequency of each type of repeated substring; the repetition threshold is the reciprocal of the number of types of repeated substrings. If the repetition frequency of a type of repeated substring is greater than the repetition threshold, then this type of repeated substring is a reference substring.

[0071] Furthermore, calculate the screening value of the reference substring according to the uniformity and length of the reference substring. Specifically, obtain the reciprocal of the length of the reference substring and add it to the uniformity of the reference substring to get the addition result; normalize the addition result to obtain the screening value of the reference substring. It should be noted that the screening value of each reference substring is the screening value of the reference substrings of the type to which the reference substring belongs.

[0072] The calculation formula for the screening value of each type of reference substring is: , where β is the normalized screening value of the selected reference substring, α is the uniformity of the selected reference substring, is the length of the selected reference substring. The larger the screening value, the greater the probability that the combination of the reference substring generates repetition and the lower the possibility of being split.

[0073] Screen the reference substrings according to the screening value to obtain screened substrings. Set a screening threshold, and the reference substrings with a screening value greater than the screening threshold are screened substrings. Thus, the repeated substrings can be screened to obtain various types of screened substrings.

[0074] The optimal grouping length obtaining module is used to obtain the grouping length range according to the minimum and maximum interval distances between every two screened substrings in each type of screened substring; use each grouping length within the grouping length range to divide the string to be processed, and obtain the optimal grouping length based on the division result.

[0075] The above has obtained the screened substrings. Further, it is necessary to obtain the grouping length range according to the screened substrings in order to determine the most suitable splitting length when splitting the string to be processed.

[0076] Furthermore, respectively obtain the minimum interval distance between every two screened substrings in each type of screened substring, and form a minimum interval distance sequence after adding each minimum interval distance to the first preset value respectively; respectively obtain the maximum interval distance between every two screened substrings in each type of screened substring, and form a maximum interval distance sequence after adding each maximum interval distance to the first preset value respectively; the least common multiple of each value in the minimum interval distance sequence is the lower limit value of the grouping length range, and the least common multiple of each value in the maximum interval distance sequence is the upper limit value of the grouping length range. Thus, the grouping length range can be obtained.

[0077] It should be noted that for each type of filtered substring, a minimum interval distance and a maximum interval distance will be obtained. The first preset value is 1. The reason for adding the first preset value to each of the minimum interval distance and the maximum interval distance respectively is that the interval distance does not include the first character of the filtered substring.

[0078] Furthermore, according to a preset step size, the preset step size is 1. Within the grouping length range, values are taken starting from the lower limit value of the grouping length range according to the preset step size until the upper limit value of the grouping length range, obtaining different grouping lengths. Each grouping length is used to divide the string to be processed respectively, obtaining the division result corresponding to each grouping length. Based on the division result, the optimal grouping length is obtained. Specifically, the average value of the ratio of the number of reference substrings of each type in the division result corresponding to a grouping length to the original number of reference substrings of each type is recorded as the division completeness rate corresponding to this grouping length. The grouping length with the largest division completeness rate is selected and recorded as the optimal grouping length. The calculation formula for the division completeness rate of each grouping length is as follows:

[0079] ,

[0080] where, represents the division completeness rate corresponding to a grouping length, K represents the number of types of reference substrings, represents the number of the k-th type of reference substring after dividing the string to be processed using this grouping length; represents the original number of the k-th type of reference substring. represents the completeness rate of the k-th type of reference substring after grouping with the selected grouping length. The larger the value, the fewer the number of the k-th type of reference substrings split after grouping. The larger the , the fewer the number of reference substrings split after grouping, and the better the division effect. Thus, the optimal grouping length can be obtained.

[0081] The compressed sequence acquisition module is used to split the string to be processed using the optimal grouping length to obtain character substrings; use the first character substring where each repeated substring appears as the target character substring; arrange the target character substrings in order to obtain a reference line; arrange the other character substrings in the columns corresponding to the respective target character substrings in the reference line according to the presence of repeated substrings in the other character substrings, obtaining an initial compressed sequence; perform a shift adjustment on the characters in each character substring in the initial compressed sequence to obtain a compressed sequence.

[0082] After obtaining the optimal grouping length, the string to be processed is split using the optimal grouping length to obtain character substrings, and the character substrings are as Figure 3 shown. Figure 3The numbers on each character substring are the numbers of the character substrings. Since the premise of repetition is that the characters in the character substring are the same characters and in the same positions, but in fact the positions are not the same and the characters may not be the same either. Therefore, it is necessary to first determine some target character substrings. Based on the character and position distribution of the target character substrings, adjust other character substrings to make the character substrings and the target character substrings have more repeated positions and the character substrings in the same column can be repeated as much as possible. Therefore, for the remaining character substrings, the target character substring with the highest character similarity should be selected for adjustment. This is because only when there are repeated characters can the position be adjusted to increase the overlap probability of the two character substrings.

[0083] Since the transformation process needs to be recorded when adjusting the character substrings for restoring the string information during subsequent decompression, and recording this process will occupy storage space. Therefore, it is necessary to analyze and calculate the shifting efficiency of each character substring to be adjusted. If the shifting consumes too much storage space, that is, the shifting efficiency is low, it may lead to an unsatisfactory compression effect. Therefore, the character substrings with low shifting efficiency are not adjusted to avoid excessive increase in the stored content.

[0084] Take the first character substring where each repeated substring appears as the target character substring. It should be noted that if there are two repeated substrings in the first character substring, the target character substrings corresponding to these two repeated substrings are both this first character substring. That is, each repeated substring corresponds to a target character substring, and there may also be cases where two, three or more repeated substrings commonly correspond to a target character substring.

[0085] Then sort the target character substrings. Specifically, record the number of characters of the repeated substrings in each target character substring as the repetition rate of the target character substring; arrange the target character substrings in descending order of the repetition rate to obtain the reference row. The reference row is as Figure 4 shown. The target character substring with the largest repetition rate has the smallest column number, and the target character substring with the smallest repetition rate has the largest column number.

[0086] Then, arrange the other character substrings in the columns corresponding to the target character substrings in the reference row according to the existence of the repeated substrings in the other character substrings. Specifically, place the character substrings containing the same repeated substring in the column corresponding to the target character substring according to their order in the string to be processed. After completion, obtain the unplaced character substrings and form an unplaced substring set; the sequence obtained after placement is recorded as the first sequence. The first sequence is as Figure 5 shown.

[0087] Obtain the similarity between each character substring in the set of unplaced substrings and each target character substring. The similarity is the number of identical letters between the character substring and the target character substring. Place each character substring in the set of unplaced substrings in the column of the target character substring with the highest similarity. When placing, the order in the row direction is determined by the order of each character substring in the set of unplaced substrings in the string to be processed; the sequence obtained after placement is denoted as the second sequence, and the second sequence is as Figure 6 shown.

[0088] Place the unplaced character substrings in the set of unplaced substrings according to their order of appearance in the string to be processed to obtain the initial compression sequence, as Figure 7 shown. Then, it is necessary to record the position of each character substring for subsequent decompression data use, as Figure 7 shown, and the recorded position set is {1, 2, 5, 6, 3, 10, 7, …}.

[0089] It should be noted that when placing the unplaced character substrings in the set of unplaced substrings according to their order of appearance in the string to be processed, it is necessary to obtain the number of character substrings in each column, and then fill in the character substrings in each column in ascending order of the number of character substrings in each column. First, fill in the column with the smallest number of character substrings, then the second smallest, and so on.

[0090] Finally, it is necessary to perform a shift operation on each character substring in the initial compression sequence. Calculate the similarity between the current character substring and the character substring in the first row or the previous row of its column respectively, and select the character substring with a higher similarity in the character substring in the first row or the previous row as the adjustment target; obtain the string with the same characters in the current character substring and the adjustment target, and denote it as the same string; if the characters at the same positions in the same string corresponding to the current character substring and the same string corresponding to the adjustment target are all the same, then the current character substring can perform a shift operation, otherwise it cannot perform a shift operation; according to the number of characters whose positions change during the shift operation of the character substring that can perform a shift operation, the number of shift operations, and the number of characters of the character substring that can perform a shift operation, obtain the adjustment efficiency of the character substring that can perform a shift operation; if the adjustment efficiency is greater than or equal to the adjustment threshold, then perform a shift operation on the character substring that can perform a shift operation. The process of the shift operation is as Figure 8 shown.

[0091] The calculation formula for the adjustment efficiency of the character substring that can perform a shift operation is:

[0092] ,

[0093] Among them, δ represents the adjustment efficiency of the character substring that can perform the shift operation; R represents the number of shift operations when the character substring that can perform the shift operation performs the shift operation; max() represents the operation of taking the maximum value; represents the number of overlapping characters in the character substring that can perform the shift operation and the adjustment target; represents the number of overlapping characters between the character substring that can perform the shift operation after the shift operation and the adjustment target; Q represents the number of characters in the character substring that can perform the shift operation.

[0094] represents the difference between the number of overlapping characters after the shift and the number of shift operations. If the number of characters is greater than the number of shift operations, it means that the shift operation is positive, that is, it reduces the memory occupancy. Otherwise, it means that the shift operation increases the additional memory occupancy. represents the effect after the shift, that is, the increase in the overlapping rate. The greater the increase in the overlapping rate, the better the shift effect. The smaller R is, the fewer the number of shift operations performed by the character substring. The larger it is, the greater the memory occupancy that can be reduced after the shift, and the higher the adjustment efficiency of the character substring, and the less storage space occupied by the record consumption. The adjustment efficiency of the remaining character substrings can be obtained in the same way.

[0095] For example, the current character substring is ABCD, and the adjustment target is ACBD. The identical strings of these two are ABC and ACB respectively. Since the characters of these two corresponding identical strings are the same, but the characters in the same positions are not exactly the same, no matter how ACBD is shifted, it cannot become ABCD. Therefore, a compressed sequence can be obtained after the shift operation.

[0096] The information sharing module is used to compress the compressed sequence using run-length encoding and transmit the compressed data to the sharing party.

[0097] Through the above operations, a compressed sequence is obtained. Further, the compressed sequence is read column by column, and the read sequence is compressed using run-length encoding to achieve the purpose of compressing the compressed sequence, obtaining a compressed data packet, and transmitting the compressed data packet to the receiving end of the sharing personnel. After receiving the compressed data packet, the receiving end decompresses it according to the recorded change situation and outputs the decompressed project file information for the sharing personnel to view.

[0098] In summary, the present invention obtains the to-be-processed string converted from the file information, and obtains different types of repeated substrings according to the character frequency and the position interval distribution of the same characters. The uniformity of each type of repeated substring is obtained according to the distribution density and position of the same type of repeated substring. The reference substring is screened according to the repeated frequency of each type of repeated substring. The screening value of the reference substring is obtained according to the length and uniformity of the reference substring. The grouping length range of the character sequence is obtained according to the screening value and the distribution interval of the highest-frequency characters, and then the optimal grouping length is obtained. The to-be-processed string is divided into character substrings, and then the character substrings are combined, and the characters in the character substrings are shifted to obtain the finally adjusted and sorted compression sequence. After reading by columns, compression is performed to obtain a compressed data packet, and then it is sent to the sharing personnel. It can improve the sharing efficiency and reduce the storage space.

[0099] It should be noted that the above order of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. And the above description of specific embodiments of this specification is given. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0100] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.

[0101] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A compression system for engineering information data, characterized in that: The system includes: The repeated substring acquisition module is used to convert the engineering information into a character string to be processed; obtain the occurrence frequency of different characters in the character string to be processed, and obtain different types of repeated substrings according to the occurrence frequency and the interval between characters; A screening substring acquisition module is used to obtain the uniformity of each repeated substring according to the variance of the interval distance between repeated substrings in each repeated substring and the length of each repeated substring; screen the repeated substrings to obtain a reference substring; calculate the screening value of the reference substring according to the uniformity and length of the reference substring, and screen the reference substring according to the screening value to obtain a screening substring; An optimal group length acquisition module is used to obtain a group length range according to the minimum and maximum interval distances between every two screening substrings in each type of screening substring; to divide the character string to be processed using each group length within the group length range, and to obtain an optimal group length based on the division result; The compression sequence acquisition module is used to segment the character string to be processed using the optimal group length to obtain character substrings; take the first character substring appearing in each repeated substring as the target character substring; arrange the target character substrings in order to obtain a reference row; arrange the other character substrings in the columns corresponding to the target character substrings in the reference row according to the existence of repeated substrings in other character substrings, so as to obtain an initial compression sequence; and shift and adjust the characters in each character substring in the initial compression sequence to obtain a compression sequence; An information sharing module, used to compress the compression sequence using run-length coding and transmit the compressed data to a sharing party; The method of obtaining different types of repeated substrings according to the occurrence frequency and the interval between characters includes: Sort the characters according to the frequency of occurrence of each character in the character string to be processed to obtain a character sequence, in which there is only one character of each type in the character sequence; record the first character in the character sequence as the first character; obtain the interval distance between every two first characters in the character string to be processed, if the interval distance is greater than or equal to a preset value, select the two first characters with the smallest interval distance and appearing at the front of the character string to be processed, record them as first type characters, form a matching template with the first character of the first type characters and the characters between the two first characters, search for a character substring identical to the matching template in the character string to be processed, obtain a repeated substring of the same type, record it as a first type repeated substring; remove the first type repeated substring in the character string to be processed to obtain a first remaining character string to be processed; Find two first characters that have the same spacing distance as the first type of characters but different spacing characters and are located at the front of the first remaining character string to be processed, and record them as second type of characters; use the first character of the second type of characters and the characters between the two first characters of the second type of characters to form a second matching template, find a character substring that is the same as the matching template in the first remaining character string to be processed, and obtain a repeated substring of the same type, which is recorded as a second type of repeated substring; remove the second type of repeated substring in the first remaining character string to be processed, and obtain a second remaining character string to be processed; When all types of repeated substrings corresponding to the two first characters with the smallest interval distance are found, the same method is used to find all types of repeated substrings corresponding to the two first characters with the second smallest interval distance, until all types of repeated substrings corresponding to the first character are found; and so on, according to the order of the characters in the character sequence, different types of repeated substrings corresponding to the characters in the character sequence are found; The calculation formula of the uniformity is: , in, Indicates the uniformity of the J-th repeated substring; represents the variance of the interval distance between repeated substrings in the Jth repeated substring; represents the number of J-th repeated substrings; n represents the number of characters in the string to be processed; Indicates the length of the J-th repeated substring; Indicates rounding down; The step of calculating the screening value of the reference substring according to the uniformity and length of the reference substring, and screening the reference substring according to the screening value to obtain the screening substring comprises: Obtain the inverse of the length of the reference substring, and add it to the uniformity of the reference substring to obtain an addition result; normalize the addition result to obtain a screening value of the reference substring; set a screening threshold, and the reference substring with a screening value greater than the screening threshold is the screening substring; The step of arranging the target character substrings in order to obtain a reference row includes: The number of characters of the repeated substring in each target character substring is recorded as the repetition rate of the target character substring; the target character substrings are arranged in descending order of repetition rate to obtain a reference row; The step of shifting and adjusting the characters in each character substring in the initial compression sequence to obtain the compression sequence includes: Calculate the similarity between the current character substring and the character substring in the first row or the previous row of the column where the current character substring is located, and select the character substring with the largest similarity from the character substring in the first row or the previous row as the adjustment target; obtain the character string with the same characters in the current character substring and the adjustment target, and record it as the same character string; if the characters in the same position of the same character string corresponding to the current character substring and the same character string corresponding to the adjustment target are the same, the current character substring can be shifted, and if they are not the same, the shift operation cannot be performed; obtain the adjustment efficiency of the character substring that can be shifted according to the number of shift operations when the character substring that can be shifted is shifted and the number of characters in the character substring that can be shifted; if the adjustment efficiency is greater than or equal to the adjustment threshold, perform the shift operation on the character substring that can be shifted; The calculation formula of the adjustment efficiency is: , Among them, δ represents the adjustment efficiency of the character substring that can be shifted; R represents the number of shift operations when the character substring that can be shifted is shifted; max() represents the maximum value operation; Indicates the number of characters in the character substring that can be shifted that overlap with the adjustment target; It indicates the number of characters that overlap with the adjustment target after the character substring that can be shifted is shifted; Q indicates the number of characters in the character substring that can be shifted.

2. A system for compressing engineering information data according to claim 1, characterized in that: The step of screening repeated substrings to obtain reference substrings includes: The ratio of the number of each repeated substring to the number of all repeated substrings is the repetition frequency of each repeated substring; the repetition threshold is the inverse of the number of types of repeated substrings. If the repetition frequency of a type of repeated substring is greater than the repetition threshold, the repeated substring of this type is the reference substring.

3. The system for compressing engineering information data according to claim 1, characterized in that: The obtaining of the packet length range according to the minimum value and the maximum value of the interval distance between every two screening substrings in each type of screening substrings includes: The minimum interval distance between every two filter substrings in each type of filter substring is obtained respectively, and each minimum interval distance value is added to the first preset value to form a minimum interval distance sequence; the maximum interval distance between every two filter substrings in each type of filter substring is obtained respectively, and each maximum interval distance is added to the first preset value to form a maximum interval distance sequence; the least common multiple of each value in the minimum interval distance sequence is the lower limit value of the group length range, and the least common multiple of each value in the maximum interval distance sequence is the upper limit value of the group length range.

4. The system for compressing engineering information data according to claim 1, characterized in that: The method of dividing the character string to be processed by using each group length within the group length range and obtaining the optimal group length based on the division result includes: According to the preset step size, within the group length range, values ​​are taken starting from the lower limit value of the group length range until the upper limit value of the group length range, to obtain different group lengths; the average value of the ratio of the number of reference substrings of each type in the division result corresponding to a group length to the original number of reference substrings of each type is obtained, which is recorded as the division completeness rate corresponding to the group length, and the group length with the largest division completeness rate is selected, which is recorded as the optimal group length.

5. The system for compressing engineering information data according to claim 1, characterized in that: The obtaining of the initial compression sequence comprises: Place the character substrings containing the same repeated substring in the column where the corresponding target character substring is located according to their order in the character string to be processed. After completion, obtain the character substrings that have not been placed to form a set of unplaced substrings; Obtain the similarity between each character substring in the unplaced substring set and each target character substring, where the similarity is the number of identical letters between the character substring and the target character substring, and select each character substring in the unplaced substring set to place in the column where the target character substring with the highest similarity is located, wherein the order in the row direction during placement is determined by the order of each character substring in the unplaced substring set in the character string to be processed; The unplaced character substrings in the unplaced substring set are placed according to their order of appearance in the character string to be processed to obtain an initial compressed sequence.

Citation Information

Patent Citations

  • HDMI (High Definition Multimedia Interface) high-definition data optimization transmission method and system

    CN116723337A

  • Dictionary compression device and memory system

    US20230289293A1