Efficient Storage System for Engineering Project Data Based on Cloud Computing

By optimizing the character combination and position span value set of engineering project data, Huffman encoding solves the problem of the same character frequency in engineering project data, achieving efficient data storage.

CN119892111BActive Publication Date: 2025-07-04JIANGXI PROVINCIAL EXPRESSWAY INVESTMENT GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510360526.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-04
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

When the prior art uses Huffman encoding to compress engineering project data, due to the presence of a large number of repeated fields, the character frequency is the same, resulting in too many branches of the Huffman tree and poor compression effect, which in turn affects storage efficiency.

Method used

By obtaining the character combination length and position span values ​​in the project string, optimize the character combination and position span value set, select the optimal combination substring for Huffman encoding, reducing the probability of the same character combination frequency.

Benefits of technology

It improves the compression effect of Huffman encoding, realizes efficient storage of engineering project data, and improves storage efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119892111B_ABST
    Figure CN119892111B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data storage, and particularly relates to an efficient storage system for engineering project data based on cloud computing. The system includes a processor and a memory. The processor executes a computer program stored in the memory to implement the following steps: obtaining a set of combined substrings corresponding to each position span value, and obtaining a preference rate corresponding to each position span value according to the set of combined substrings corresponding to each position span value, selecting the position span value corresponding to the maximum preference rate as the target position span value, and finally performing Huffman coding on the combined substrings in the set of combined substrings corresponding to the target position span value according to the frequencies of different types of combined characters appearing in the set of combined substrings corresponding to the target position span value, and storing the data obtained after encoding. Moreover, the present invention can improve the compression effect of engineering project information, thereby realizing the high efficiency of engineering project data storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data storage, and particularly to an efficient storage system for engineering project data based on cloud computing. Background Art

[0002] Since engineering project data is relatively important for engineering project management, for example, by collecting and analyzing engineering project data, the project manager can monitor the progress of the project in real time to ensure that the project proceeds as planned. And because engineering project data has the characteristics of large data volume, diverse data types, and frequent data updates, Huffman coding is often used to encode and compress engineering project data currently, and then the compressed data is transmitted and stored for subsequent calls. However, since there may be a large number of duplicate fields in engineering project data, when there are a large number of duplicate fields in engineering project data, it will cause the phenomenon that some characters have the same frequency in the characters to be encoded after conversion. And this phenomenon will lead to more branches in the construction of the Huffman tree, thus resulting in a poor compression effect, and further unable to achieve the high efficiency of engineering project data storage, such as poor storage efficiency. Therefore, when using Huffman coding to compress engineering project data, how to improve the compression effect has become an urgent problem to be solved. Summary of the Invention

[0003] In order to solve the above problems, the present invention provides an efficient storage system for engineering project data based on cloud computing, and the specific technical solutions adopted are as follows:

[0004] An embodiment of the present invention provides an efficient storage system for engineering project data based on cloud computing, including a processor and a memory. The processor executes the computer program stored in the memory to implement the following steps:

[0005] Obtain an engineering project string to be encoded;

[0006] According to different character combination lengths and the frequencies of different types of characters in the engineering project string, combine the characters in the engineering project string to obtain a combined effect characterization value corresponding to each character combination length, and select the character combination length corresponding to the maximum combined effect characterization value as the target character combination length;

[0007] According to the frequencies of different types of characters in the engineering project string and the positions of each character in the engineering project string, obtain a set of position span values;

[0008] Based on the length of the target character combination, each position span value in the set of position span values, and the engineering project string, obtain the set of combined substrings corresponding to each position span value. Then, based on the set of combined substrings corresponding to each position span value, obtain the preference rate corresponding to each position span value. Select the position span value corresponding to the maximum preference rate as the target position span value. According to the frequencies of different types of combined sub-characters appearing in the set of combined substrings corresponding to the target position span value, perform Huffman coding on the combined substrings in the set of combined substrings corresponding to the target position span value, and store the data obtained after coding.

[0009] Advantageous effects: The present invention first obtains the engineering project string to be encoded, and then combines the characters in the engineering project string according to different character combination lengths and the frequencies of different types of characters in the engineering project string to obtain the combined effect characterization values corresponding to each character combination length. Select the character combination length corresponding to the maximum combined effect characterization value as the target character combination length. Then, according to the frequencies of different types of characters appearing in the engineering project string and the positions of each character in the engineering project string, obtain the set of position span values. Next, based on the target character combination length, each position span value in the set of position span values, and the engineering project string, obtain the set of combined substrings corresponding to each position span value. Then, based on the set of combined substrings corresponding to each position span value, obtain the preference rate corresponding to each position span value. Select the position span value corresponding to the maximum preference rate as the target position span value. Finally, according to the frequencies of different types of combined sub-characters appearing in the set of combined substrings corresponding to the target position span value, perform Huffman coding on the combined substrings in the set of combined substrings corresponding to the target position span value, and store the data obtained after coding. And the probability that the frequencies of different types of combined substrings in the set of combined substrings determined according to the preference rate are the same is relatively small in the present invention. Thus, the compression effect during subsequent compression using Huffman coding, that is, the compression ratio, can be improved. That is, this embodiment can improve the compression effect for engineering project information, thereby achieving the high-efficiency storage of engineering project data. Brief Description of the Drawings

[0010] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1Flow chart of an efficient storage method for engineering project data based on cloud computing according to the present invention. Detailed implementation manners

[0012] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art belong to the scope protected by the embodiments of the present invention.

[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art belonging to the technical field of the present invention.

[0014] This embodiment provides an efficient storage system for engineering project data based on cloud computing, including a processor and a memory. The processor executes the computer program stored in the memory to implement an efficient storage method for engineering project data based on cloud computing. As Figure 1 shown, the efficient storage method for engineering project data based on cloud computing includes the following steps:

[0015] Step S001: Obtain an engineering project string to be encoded.

[0016] Since the main purpose of this embodiment is to improve the compression effect of engineering project data, and in order to improve the compression effect in this embodiment, the subsequent compression effect is mainly improved by adjusting the preprocessing process before compression.

[0017] This embodiment first obtains the engineering project information to be compressed, then converts all Chinese parts in the obtained engineering project information into English, and records the completely converted data as the engineering project string to be encoded. When converting, it can be required to be represented in all capital letters or lowercase letters; and in this embodiment, the engineering project information text includes but is not limited to project pre-information, construction period information, and information after project completion, and the engineering project information text to be compressed can be read from the temporary storage file location.

[0018] Therefore, this embodiment obtains the engineering project string to be encoded through the above process, and in this embodiment, one English letter represents one character.

[0019] Step S002: According to different character combination lengths and the frequencies of different types of characters appearing in the engineering project string, combine the characters in the engineering project string to obtain combination effect characterization values corresponding to each character combination length, and select the character combination length corresponding to the largest combination effect characterization value as the target character combination length.

[0020] Since a large number of repeated fields are included in the engineering project information, after being translated into English, if single characters are used as the basic characters to construct the Huffman tree, the compression effect is not good because the occurrence frequencies of single characters are likely to be the same. Therefore, combined substrings are used as the basic characters, and then different lengths of codes are selected according to the different occurrence frequencies of different types of combined substrings. Similarly, the higher the occurrence frequency of a combined substring, the shorter the code used; and in this embodiment, it is expected that there are differences in the occurrence frequencies of different types of combined substrings in the future. Therefore, this embodiment needs to first analyze the combined effect characterization values corresponding to different character combination lengths. And the larger the combined effect characterization value corresponding to the character combination length, the more it can indicate that when combining characters with this character combination length, the frequency difference between different types of combined substrings is more obvious, that is, it is easier to generate combined substrings of the same type. Therefore, based on the above description, in the next step of this embodiment, the characters in the engineering project string will be combined according to different character combination lengths and the occurrence frequencies of different types of characters in the engineering project string, and the combined effect characterization values corresponding to each character combination length will be obtained, and the character combination length corresponding to the largest combined effect characterization value will be selected as the target character combination length.

[0021] And in this embodiment, the specific process of combining the characters in the engineering project string according to different character combination lengths and the occurrence frequencies of different types of characters in the engineering project string to obtain the combined effect characterization values corresponding to each character combination length is as follows:

[0022] First, all types of characters appearing in the project string are obtained, and the set constructed by all types of characters appearing in the project string is recorded as a character set. Then, the number of character types appearing in the project string is obtained, and recorded as a type quantity representation value. For example, if the project string is "timeismoney", the character types appearing in the project string are t, i, m, e, s, o, n and y respectively, and the obtained character set is {t, i, m, e, s, o, n, y}, and the type quantity representation value at this time is 8. Then, the frequency of occurrence of each character in the character set in the project string is counted, and recorded as the first frequency value of the corresponding character, and then all the characters in the character set are sorted in descending order according to the first frequency value to obtain a first character sequence. In general, the probability of occurrence of different types of combined substrings generally depends on the character with the lowest frequency. Therefore, when determining the combination effect representation value corresponding to different character combination lengths, it is necessary to determine it according to the result of combining from large to small frequencies, that is, it is necessary to determine it according to the result of combining from large to small frequencies, that is, it is necessary to determine it according to the first character sequence. The characters in the character sequence are combined according to their order; therefore, after obtaining the first character sequence, this embodiment will then obtain a character combination length sequence, and in specific applications, it is necessary to set the minimum character combination length and the maximum character combination length in the character combination length sequence according to actual conditions, but it is required that the character combination lengths in the character combination length sequence are all positive integers. For example, in this embodiment, the minimum character combination length in the character combination length sequence can be 2, and the maximum character combination length can be the upward rounded value of 10% of the total number of characters in the project character string. At this time, the character combination length sequence is composed of all integers between the minimum character combination length and the maximum character combination length; in addition, the character combination length refers to the number of characters in each combination obtained in the subsequent combination. If the character combination length is 2, the number of characters in each combination is 2; then, based on the first character sequence that has been obtained, this embodiment will take the acquisition process of the combination effect representation value corresponding to any character combination length in the character combination length sequence as an example to describe, that is, the acquisition process of the combination effect representation value corresponding to the character combination length is:

[0023] First, the value of the character combination length is recorded as A1, and A1 different types of characters are randomly selected in the project character string for permutation and combination. Then, after the permutation and combination is completed, the number of combinations obtained is counted and recorded as the permutation and combination value. Then, the reciprocal of the permutation and combination value is obtained and recorded as the first representation value corresponding to the character combination length, and the permutation and combination value can be obtained according to the permutation and combination calculation formula; then, according to the first character sequence and the value A1 of the character combination length, the combination quantity representation value corresponding to the character combination length is obtained, and the result obtained by subtracting the combination quantity representation value from the type quantity representation value is recorded as the feature difference value, and the reciprocal of the result obtained by adding Max(0,D1) to the preset first constant is recorded as the feature ratio value, Max() is the maximum value function, and D1 is the feature difference value; the result obtained by subtracting the feature ratio value from the preset first constant is recorded as the second representation value corresponding to the character combination length; finally, the average of the first representation value and the second representation value is obtained, and recorded as the combination effect representation value corresponding to the character combination length; and the specific calculation expression of the combination effect representation value corresponding to the character combination length is:

[0024]

[0025] Among them, P0 is the combination effect representation value corresponding to the character combination length, N1 is the number of permutations and combinations, N0 is the number of types, N2 is the number of combinations, and c1 is the preset first constant; and in specific applications, the implementer needs to set the value of the preset first constant according to actual conditions. For example, in this embodiment, the preset first constant can be set to 1; in addition, when N1 and The smaller the value, the larger the P0 value is, and the larger the P0 value is, the more obvious the frequency difference between different types of combined substrings is when characters are combined with the character combination length, that is, it is less likely to produce the same frequency phenomenon between different types of combined substrings.

[0026] In this embodiment, the specific process of obtaining the combination quantity representation value corresponding to the above-mentioned character combination length is as follows:

[0027] First, determine whether the number of character types in the first character sequence is not less than A1. If so, obtain the combined value corresponding to the first combined type, the second character sequence, and the second frequency value of each character in the second character sequence, and determine whether the number of character types in the second character sequence is not less than A1. If so, obtain the combined value corresponding to the second combined type, the third character sequence, and the third frequency value of each character in the third character sequence. Continue to determine whether the number of character types in the third character sequence is not less than A1. If so, obtain the combined value corresponding to the third combined type, the fourth character sequence, and the fourth frequency value of each character in the fourth character sequence. Continue to determine whether the number of character types in the fourth character sequence is not less than A1. If not, accumulate the combined values of all the combined types obtained, and use the accumulated result as the combined quantity characterization value corresponding to the length of the character combination. As another implementation, the result of adding the accumulated result of the combined values of all the combined types obtained to the number of special combinations can also be used as the combined quantity characterization value corresponding to the length of the character combination. The number of special combinations is , is the ceiling symbol, A0 is the sum of the fourth frequency values of all characters in the fourth character sequence; and the first combined type is the combined type formed by the first A1 characters at the front of the first character sequence, and the second combined type is the combined type formed by the first A1 characters at the front of the second character sequence. For example, if a character sequence is {t, i, m, e, o, n, y}, and A1 is 2 at this time, then the combined type formed by the first 2 characters in this character sequence is {ti}.

[0028] In addition, in this embodiment, the specific obtaining processes of the combined value corresponding to the first combined type, the second character sequence, the second frequency value of each character in the second character sequence, the combined value corresponding to the second combined type, the third character sequence, the third frequency value of each character in the third character sequence, the combined value corresponding to the third combined type, the fourth character sequence, and the fourth frequency value of each character in the fourth character sequence are as follows:

[0029] Record the first frequency value of the A1-th character in the first character sequence as the combined value corresponding to the first combined type; obtain the first remaining frequency values of each character in the first character sequence, and record the sequence after removing the characters with the first remaining frequency value of 0 in the first character sequence as the second character sequence. And the first remaining frequency value of the first A1 characters at the front of the first character sequence is the first quantity difference of the corresponding character. The first A1 characters at the front include the A1-th character, and the first remaining frequency value of the other characters behind the A1-th character in the first character sequence is the first frequency value of the corresponding character. The first quantity difference of the first A1 characters at the front of the first character sequence is the difference between the first frequency value of the corresponding character and the combined value corresponding to the first combined type. Record the first remaining frequency values of each character in the second character sequence as the second frequency value of the corresponding character, that is, the first quantity difference of the a0-th character in the first character sequence is the result of subtracting the combined value corresponding to the first combined type from the first frequency value of the corresponding character, and the value range of a0 is [1, A1].

[0030] Record the smallest second frequency value among the first A1 characters in the second character sequence as the combined value corresponding to the second combined type. Obtain the second remaining frequency values of each character in the second character sequence, and record the sequence after removing the characters with the second remaining frequency value of 0 in the second character sequence as the third character sequence. And the second remaining frequency value of the first A1 characters at the front of the second character sequence is the second quantity difference of the corresponding character. The second remaining frequency value of the other characters behind the A1-th character in the second character sequence is the second frequency value of the corresponding character. The second quantity difference of the first A1 characters at the front of the second character sequence is the difference between the second frequency value of the corresponding character and the combined value corresponding to the second combined type. Record the second remaining frequency values of each character in the third character sequence as the third frequency value of the corresponding character.

[0031] And the specific method for obtaining the combined value corresponding to the third combined type, the fourth character sequence, and the fourth frequency values of each character in the fourth character sequence is the same as the specific method for obtaining the combined value corresponding to the second combined type, the third character sequence, and the third frequency values of each character in the third character sequence above. Therefore, this embodiment will not be described in detail.

[0032] For the sake of easy understanding, this embodiment will illustrate the specific process of obtaining the characterization value of the combination quantity corresponding to the above character combination length: If the first character sequence at this time is {a, b, c, d}, A1 is 2, and the first frequency value of a is 7, the first frequency value of b is 6, the first frequency value of c is 5, and the first frequency value of d is 1, then the first combination type at this time is {ab}, and according to the first frequency values of a and b, it can be known that the maximum number of combinations belonging to the first combination type {ab} that can be obtained at this time is 6. After completing the combinations of the first combination type {ab}, the first character sequence composed of the remaining characters is {a, c, d}, and at this time the second frequency value of a is 1, the second frequency value of c is 5, and the second frequency value of d is 1, then the second combination type at this time is {ac}, and according to the second frequency values of a and c, it can be known that the maximum number of combinations belonging to the second combination type {ac} that can be obtained at this time is 1. After completing the combinations of the second combination type {ac}, the third character sequence composed of the remaining characters is {c, d}, and at this time the third frequency value of c is 4 and the third frequency value of d is 1, then the third combination type at this time is {cd}, and according to the third frequency values of c and d, it can be known that the maximum number of combinations belonging to the third combination type {cd} that can be obtained at this time is 1. After completing the combinations of the third combination type {cd}, only the character c remains, then a special combination appears. Since the number of c remaining at this time is 3, the number of combinations of the special combination that appears at this time is two, namely the combinations {cc} and {c}. Then, the combination quantity values obtained above are accumulated, and the finally obtained combination quantity characterization value is 6 + 1 + 1 + 2 = 10.

[0033] Therefore, through the above process, this embodiment obtains the characterization values of the combination effects corresponding to each character combination length, and then selects the character combination length corresponding to the largest combination effect characterization value as the target character combination length, and the target character combination length can indicate the number of characters in each combined substring finally obtained. For example, if the target character combination length is 2, then the number of characters in each combined substring finally obtained is 2.

[0034] Step S003: Obtain a set of position span values according to the frequencies of different types of characters appearing in the engineering project string and the positions of each character in the engineering project string.

[0035] In this embodiment, the position span value set will be obtained according to the frequencies of different types of characters appearing in the engineering project string and the positions of each character in the engineering project string; and the purpose of obtaining the position span value set is to make the probability that the different types of combined substrings finally obtained have the same frequency relatively small. Then, the specific process of obtaining the position span value set in this embodiment is as follows:

[0036] First, obtain the position marker values of each character in the engineering project string, and the position marker value of the b1-th character in the engineering project string is b1; then, in the engineering project string, extract all the characters that are the same as the first character in the first character sequence, and record the sequence constructed by all the extracted characters as the first sequence to be analyzed. Also extract all the characters that are the same as the second character in the first character sequence, and record the sequence constructed by all the extracted characters as the second sequence to be analyzed. The characters in the sequence to be analyzed are arranged in ascending order of the position marker values.

[0037] Immediately afterwards, according to the absolute value of the difference between the position marker values of the characters at the same position in the first sequence to be analyzed and the second sequence to be analyzed, obtain the marker difference sequence. The d1-th marker difference in the marker difference sequence is the absolute value of the difference between the position marker value of the d1-th character in the first sequence to be analyzed and the position marker value of the d1-th character in the second sequence to be analyzed. Moreover, the length of the marker difference sequence is consistent with the length of the shortest sequence to be analyzed between the first sequence to be analyzed and the second sequence to be analyzed.

[0038] Then, obtain the maximum marker difference in the marker difference sequence as the maximum position span value, take the preset first constant as the minimum position span value, and obtain the interval formed by the minimum position span value and the maximum position span value, and record it as the position span interval. After that, record the set composed of all integers in the position span interval as the position span value set.

[0039] Therefore, in this embodiment, the position span value set is obtained through the above process.

[0040] Step S004: According to the target character combination length, each position span value in the position span value set, and the engineering project string, obtain the combined substring sets corresponding to each position span value, and according to the combined substring sets corresponding to each position span value, obtain the preference rates corresponding to each position span value. Select the position span value corresponding to the maximum preference rate as the target position span value. According to the frequencies of different types of combined sub-characters appearing in the combined substring set corresponding to the target position span value, perform Huffman coding on the combined substrings in the combined substring set corresponding to the target position span value, and store the data obtained after coding.

[0041] After obtaining the set of position span values in this embodiment, the preferred rate of different position span values is obtained according to the set of combined substring obtained for different position span values. The larger the preferred rate, the smaller the probability that the occurrence frequencies of different types of combined substrings in the set of combined substrings obtained according to the corresponding position span value are the same, and the lower the complexity of reading the combined substrings according to the corresponding position span value. Therefore, based on the above analysis, it can be known that in the following, this embodiment will obtain the set of combined substrings corresponding to each position span value in the set of position span values according to the target character combination length, each position span value in the set of position span values, and the engineering project string, and obtain the preferred rate corresponding to each position span value according to the set of combined substrings corresponding to each position span value.

[0042] In this embodiment, the specific process of obtaining the set of combined substrings corresponding to each position span value in the set of position span values according to the target character combination length, each position span value in the set of position span values, and the engineering project string is as follows:

[0043] Record the value of the target character combination length as A2, and record the string formed by the first A2 characters in the first character sequence as the reference substring. For example, if the first character sequence is {a, b, c, d} and A2 is 3, then the reference substring is {abc}. Then, for the sake of easy understanding, the process of obtaining the set of combined substrings corresponding to the g0-th position span value in the position span interval is taken as an example for description. That is, the process of obtaining the set of combined substrings corresponding to the g0-th position span value is as follows:

[0044] First, obtain the product of the g0-th position span value and A2, and denote it as the segmentation length value corresponding to the g0-th position span value. Then, use the segmentation length value corresponding to the g0-th position span value to evenly and non-overlappingly segment the engineering project string, and denote all the obtained substrings as the substrings to be analyzed corresponding to the engineering project string under the g0-th position span value. After that, according to all the substrings to be analyzed corresponding to the engineering project string under the g0-th position span value and the reference substring, obtain the target substring corresponding to each substring to be analyzed, and denote the new string formed by the target substrings corresponding to all the substrings to be analyzed as the new string corresponding to the g0-th position span value. For example, if the engineering project string is {timeonyfirstserved}, and the segmentation length value corresponding to the g0-th position span value is 6, then all the substrings to be analyzed corresponding to the engineering project string under the g0-th position span value are {timeon}, {yfirst}, and {served}, and if the target substring corresponding to {timeon} is W, the target substring corresponding to {yfirst} is E, and the target substring corresponding to {served} is R, then the new string corresponding to the g0-th position span value is {WER} at this time.

[0045] Immediately afterwards, according to the g0-th position span value and the target character combination length, combine the characters in the engineering project string, and denote the set composed of all the obtained substrings as the first set. Then, according to the g0-th position span value and the target character combination length, combine the characters in the new string corresponding to the g0-th position span value, and denote the set composed of all the obtained substrings as the second set. Then, according to the first set and the second set, obtain the gain characterization value corresponding to the g0-th position span value, and determine whether the gain characterization value corresponding to the g0-th position span value is greater than the preset gain threshold. If so, it indicates that the encoding after combining the new string is better in terms of compression effect than the encoding after combining the engineering project string. Therefore, at this time, the second set is used as the set of combined substrings corresponding to the g0-th position span value. Otherwise, it indicates that the encoding after combining the engineering project string is better in terms of compression effect than the encoding after combining the new string. Therefore, at this time, the first set is used as the set of combined substrings corresponding to the g0-th position span value, and the substrings in the set of combined substrings are denoted as combined substrings. In addition, in specific applications, the implementer needs to set the preset gain threshold according to the actual situation. For example, in this embodiment, the preset gain threshold can be set to 0.

[0046] In this embodiment, the process of obtaining the target substring corresponding to the substring to be analyzed is as follows:

[0047] For any substring g1 corresponding to the engineering project string under the g0-th position span value: First, extract all the characters in the substring g1 to be analyzed that are the same as the characters in the reference substring, and denote the string constructed by all the extracted characters as the feature string, and the characters in the feature string are arranged in ascending order of the position marker values; then, according to the feature string, obtain all the positions where special characters are added in the substring g1 to be analyzed, and add special characters at all the positions where special characters are added. Denote the substring g1 after adding special characters as the target substring corresponding to the substring g1 to be analyzed; and this step is to group together the characters with relatively high occurrence frequencies as much as possible.

[0048] In this embodiment, the process of obtaining the positions where special characters are added in the substring g1 to be analyzed is as follows:

[0049] First, obtain all the local substrings corresponding to the feature string. If the string formed by the j1-th character to the j1 + A2 - 1-th character in the feature string is the same as the reference substring, then denote the string formed by the j1-th character to the j1 + A2 - 1-th character as the local substring, including the j1-th character and the j1 + A2 - 1-th character; then, in the substring g1 to be analyzed, obtain the characters with the same position marker values as the characters in the local substring, and denote them all as the characters to be analyzed; and for example, if the substring g1 to be analyzed is {timeonyfirst} and the reference substring is {ti}, then the feature string at this time is {tiit}, and the {ti} in the feature string is the same as the reference substring. Then, the 1st and 2nd characters in the substring g1 {timeonyfirst} to be analyzed are the characters to be analyzed.

[0050] After obtaining the character to be analyzed, the special character adding position in the substring g1 to be analyzed is obtained, and for the k1th character to be analyzed and the k1+1th character to be analyzed in the substring g1 to be analyzed, if the absolute value of the difference between the position mark value of the k1th character to be analyzed and the position mark value between the k1+1th character to be analyzed is less than the g0th position span value, the k1th character to be analyzed is not the last character in the local substring, and the k1+1th character to be analyzed is not the first character in the local substring, then any M0 positions between the k1th character to be analyzed and the k1+1th character to be analyzed are selected as special character adding positions, M0 is the absolute value of the difference between the g0th position span value and the difference representation value, and the difference representation value is the absolute value of the difference between the position mark value of the k1th character to be analyzed and the position mark value of the k1+1th character to be analyzed. For example, if the substring g1 to be analyzed is {timeonyfirst}, the span value of the g0th position is 4, the first and second characters in the substring g1 to be analyzed are the characters to be analyzed, and the string formed by the first two characters in the substring g1 to be analyzed is a local substring, and at this time the absolute value of the difference in position mark values ​​between the first character and the second character in the substring g1 to be analyzed is 1. Since the absolute value of the difference in position mark values ​​between the first character and the second character in the substring g1 to be analyzed is less than 4, the first character in the substring g1 to be analyzed is not the last character in the local substring, and the second character in the substring g1 to be analyzed is not the first character in the local substring, then at this time, 3 special character addition positions need to be randomly added between the first character and the second character in the substring g1 to be analyzed; in addition, it should be noted that the implementer needs to set special characters according to actual conditions. For example, in this embodiment, "#" can be set as a special character.

[0051] In this embodiment, the process of obtaining the first set and the second set is as follows:

[0052] The project string is recorded as the first string to be read, and starting from the first character on the first string to be read, the characters in the first string to be read are read to obtain the first read substring, and then the characters that have been read in the first string to be read are removed, and the string formed by the remaining characters after the removal is recorded as the second string to be read, and it is determined whether the number of characters in the second string to be read is greater than or not less than , if so, starting from the 1st character in the second string to be read, read the characters in the second string to be read to obtain the 2nd read substring. Then, remove the characters that have been read from the second string to be read, and denote the string formed by the remaining characters after removal as the third string to be read. Then, determine whether the number of characters in the third string to be read is not less than , if not, denote the third string to be read as the 3rd read substring, stop obtaining read substrings, and denote the set constructed by all the obtained read substrings as the first set. The absolute value of the difference between the position marking values of two adjacent characters in the 1st read substring and the 2nd read substring is the g0th position span value, that is, the absolute value of the difference between the position marking values of two adjacent characters in other read substrings except the last obtained read substring is the g0th position span value, and the lengths of other read substrings except the last obtained read substring are all the target character combination lengths. c2 is a preset second constant; in addition, in practical applications, the implementer needs to set the value of c2 according to the actual situation. For example, the value of c2 can be set to 2.

[0053] Immediately denote the new string corresponding to the g0th position span value as the first feature string. Since special characters are added to the new string, the position marking values obtained according to the positions of the characters in the engineering project string are no longer applicable at this time. Therefore, at this time, it is necessary to obtain the feature marking values of each character in the first feature string, and the feature marking value of the v1th character in the first feature string is v1; then, starting from the 1st character in the first feature string, read the characters in the first feature string to obtain the 1st feature substring. Then, remove the characters that have been read from the first feature string, and denote the string formed by the remaining characters after removal as the second feature string. Then, determine whether the number of characters in the second feature string is not less than , if so, starting from the 1st character in the second feature string, read the characters in the second feature string to obtain the 2nd feature substring. Then, remove the characters that have been read from the second feature string, and denote the string formed by the remaining characters after removal as the third feature string. Then, determine whether the number of characters in the third feature string is not less than , if not, record the third feature string as the 3rd feature substring, stop obtaining the feature substrings, and record the set constructed by all the obtained feature substrings as the second set. The absolute value of the difference between the position marker values of two adjacent characters in the 1st feature substring and the 2nd feature substring is the g0th position span value, that is, the absolute value of the difference between the position marker values of two adjacent characters in other feature substrings except the last obtained feature substring is the g0th position span value, and the lengths of other feature substrings except the last obtained reading substring are all the target character combination lengths.

[0054] For example, if the first string to be read is {timeonyfirst}, the g0th position span value is 2, and the target character combination length is 3, then the 1st reading substring is {tmo}, the second string to be read is {ienyfirst}, the 2nd reading substring is {ien}, the third string to be read is {yfirst}, the 3rd reading substring is {yis}, the fourth string to be read is {frt}, and the last 1 reading substring is {frt}.

[0055] In this embodiment, the process of obtaining the gain characterization value corresponding to the g0th position span value is as follows: count the total number of special characters added in all the substrings to be analyzed, and record it as the character addition quantity value; in the first set, obtain the number of reading substrings with an occurrence frequency of 1, and record it as the first quantity feature value, obtain the number of reading substrings with an occurrence frequency not equal to 1, and record it as the second quantity feature value; in the second set, obtain the number of feature substrings with an occurrence frequency of 1, and record it as the third quantity feature value, obtain the number of feature substrings with an occurrence frequency not equal to 1, and record it as the fourth quantity feature value; obtain the result of subtracting the first quantity feature value from the third quantity feature value, and record it as the first increment value, obtain the result of subtracting the second quantity feature value from the fourth quantity feature value, and record it as the second increment value, record the sum of the first increment value and the character addition quantity value as the fifth quantity feature value, and record Max(0, L1) as the gain characterization value corresponding to the g0th character span value, where L1 is the result of subtracting the fifth quantity feature value from the second increment value; and the larger the gain characterization value, the more it indicates that adding special characters can reduce the probability of the occurrence of substrings with an occurrence frequency of 1, thereby reducing the probability of different substrings having the same occurrence frequency, and two substrings of the same type mean that the lengths of these two substrings are not only the same, but also the characters at the same positions in these two substrings are the same. For example, the substring {yis} and the substring {yis} are substrings of the same type, while the substring {yvs} and the substring {yis} are substrings of different types.

[0056] Therefore, through the above process, each position span value in the position span value set can be obtained corresponding to the combined substring set. After obtaining the combined substring sets corresponding to each position span value, the preference rate corresponding to each position span value is obtained according to the combined substring sets corresponding to each position span value. For the convenience of understanding, the specific obtaining process of the preference rate corresponding to the g0-th position span value in the position span interval is taken as an example for description. That is, the specific obtaining process of the preference rate corresponding to the g0-th position span value is as follows:

[0057] First, denote the combined substring set corresponding to the g0-th character span value as the set to be analyzed, and denote the set constructed by all combined substring types that appear in the set to be analyzed as the first type set. Then, count the frequency of each combined substring type in the first type set that appears in the set to be analyzed, and denote it as the frequency value corresponding to the combined substring type. After that, denote the set constructed by all combined substring types with a frequency value of 1 in the first type set as the second type set, and denote the set constructed by all combined substring types in the first type set except the second type set as the third type set. Immediately afterwards, obtain the sum of the frequency values of all combined substring types in the second type set, and denote it as the first frequency sum value. Obtain the sum of the frequency values of all combined substring types in the third type set, and denote it as the second frequency sum value. And denote the normalized value of the result obtained by subtracting the second frequency sum value from the second frequency sum value as the first eigenvalue. Then, denote the set constructed by the frequency values of all combined substring types in the first type set as the frequency value set, and denote the reciprocal of the result obtained by adding the standard deviation of the frequency value set and a preset first constant as the second ratio. Denote the result obtained by subtracting the second ratio from the preset first constant as the second eigenvalue. After that, obtain the reciprocal of the total number of combined substring types in the second type set, and denote it as the third eigenvalue. Then, obtain the reciprocal of the total number of combined substring types in the first type set, and denote it as the fourth eigenvalue. Finally, take the mean of the first eigenvalue, the second eigenvalue, the third eigenvalue, and the fourth eigenvalue as the first preference value corresponding to the g0-th position span value. Denote the normalized value of the ratio of the target character combination length to the segmentation length value corresponding to the g0-th position span value as the second preference value corresponding to the g0-th position span value. Denote the mean of the first preference value and the second preference value as the preference rate corresponding to the g0-th position span value.

[0058] And the calculation expression of the preference rate corresponding to the g0-th position span value is:

[0059]

[0060] Wherein, is the preferred rate corresponding to the g0th position span value, L1 is the cumulative value of the first frequency, L2 is the cumulative value of the second frequency, S1 is the standard deviation of the frequency value set, Q2 is the total number of combined substring types in the second type set, Q1 is the total number of combined substring types in the first type set, Norm() is the normalization function, A2 is the value of the target character combination length, M2 is the segmentation length value corresponding to the g0th position span value, and , , is the first preferred value, is the second preferred value; and when and are larger, it indicates that the probability of the frequency values of the combined substrings of different types in the set to be analyzed appearing the same is smaller, and the complexity during grouped reading is lower, that is, it is convenient to obtain the combined substrings; in addition, when is larger, Q2 is smaller, S1 is larger, and Q1 is smaller, it indicates that the probability of the frequency values of the combined substrings of different types in the set to be analyzed appearing the same is smaller, is smaller, it indicates that the number of substrings to be analyzed corresponding to the engineering project string under the g0th position span value is smaller, then the complexity during combined substring acquisition is lower, and considering is to avoid the phenomenon of higher reading complexity when reading combined substrings.

[0061] Therefore, through the above process, this embodiment can obtain the preferred rate corresponding to each position span value, then select the position span value corresponding to the maximum preferred rate as the target position span value, and perform Huffman coding on the combined substrings in the combined substring set corresponding to the target position span value according to the frequencies of different types of combined characters that appear in the combined substring set corresponding to the target position span value, and store the data obtained after encoding. Moreover, the greater the frequency of the combined character that appears, the greater the length of the encoding. And in the case of knowing the frequencies of different types of combined characters that appear in the combined substring set corresponding to the target position span value, the process of Huffman coding is a well-known technology, so this embodiment will not be described again; in addition, it should be noted that if the compressed data has been restored, the special characters in the restored data need to be removed, and text conversion is performed on the data after removal.

[0062] So far, this embodiment has completed the compressed storage of the engineering project information.

[0063] In summary, in this embodiment, the engineering project string to be encoded is first obtained, and then, according to different character combination lengths and the frequencies of different types of characters in the engineering project string, the characters in the engineering project string are combined to obtain the combination effect characterization values corresponding to each character combination length, and the character combination length corresponding to the largest combination effect characterization value is selected as the target character combination length. Then, according to the frequencies of different types of characters in the engineering project string and the positions of each character in the engineering project string, a set of position span values is obtained. Then, according to the target character combination length, each position span value in the set of position span values, and the engineering project string, a set of combined substring corresponding to each position span value is obtained, and according to the set of combined substrings corresponding to each position span value, the preference rate corresponding to each position span value is obtained. The position span value corresponding to the largest preference rate is selected as the target position span value. Finally, according to the frequencies of different types of combined sub-characters in the set of combined substrings corresponding to the target position span value, Huffman coding is performed on the combined substrings in the set of combined substrings corresponding to the target position span value, and the data obtained after encoding is stored. And in this embodiment, the probability that the frequencies of different types of combined substrings in the set of combined substrings determined according to the preference rate are the same is small, so that the compression effect during subsequent compression using Huffman coding, that is, the compression rate, can be improved. That is, this embodiment can improve the compression effect on engineering project information, thereby realizing the high-efficiency storage of engineering project data.

[0064] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. An efficient storage system for engineering project data based on cloud computing, comprising a processor and a memory, characterized in that, The processor executes the computer program stored in the memory to implement the following steps: Obtain the engineering project string to be encoded; According to different character combination lengths and the frequencies of different types of characters appearing in the engineering project string, combine the characters in the engineering project string to obtain the combined effect characterization values corresponding to each character combination length, and select the character combination length corresponding to the maximum combined effect characterization value as the target character combination length; According to the frequencies of different types of characters appearing in the engineering project string and the positions of each character in the engineering project string, obtain a set of position span values; According to the target character combination length, each position span value in the set of position span values, and the engineering project string, obtain a set of combined substring corresponding to each position span value, and according to the set of combined substrings corresponding to each position span value, obtain the preference rate corresponding to each position span value, select the position span value corresponding to the maximum preference rate as the target position span value, perform Huffman coding on the combined substrings in the set of combined substrings corresponding to the target position span value according to the frequencies of different types of combined sub-characters appearing in the set of combined substrings corresponding to the target position span value, and store the data obtained after encoding; The method for obtaining the combined effect characterization value corresponding to each character combination length includes: denoting the set constructed by all types of characters appearing in the engineering project string as the character set, and denoting the number of types of characters appearing in the engineering project string as the type quantity characterization value; denoting the frequencies of each character in the character set appearing in the engineering project string as the first frequency values of the corresponding characters, sorting all the characters in the character set in descending order according to the first frequency values to obtain the first character sequence; for any character combination length: denoting the value of the character combination length as A1, randomly select A1 different types of characters in the engineering project string for combination, and denoting the reciprocal of the number of combinations obtained after combination as the first characterization value corresponding to the character combination length; according to the first character sequence and the value of the character combination length, obtain the combined quantity characterization value corresponding to the character combination length; denoting the result obtained by subtracting the combined quantity characterization value from the type quantity characterization value as the feature difference value, and denoting the reciprocal of the result obtained by adding Max(0, D1) and a preset first constant as the feature ratio value, where Max() is the maximum value function and D1 is the feature difference value; denoting the result obtained by subtracting the feature ratio value from the preset first constant as the second characterization value corresponding to the character combination length, and the first constant is 1; denoting the mean value of the first characterization value and the second characterization value as the combined effect characterization value corresponding to the character combination length; The method for obtaining the characterization value of the number of combinations corresponding to the length of the character combination includes: determining whether the number of character types in the first character sequence is not less than A1. If so, obtaining the number of combination values corresponding to the first combination type, the second character sequence, and the second frequency values of the respective characters in the second character sequence, and determining whether the number of character types in the second character sequence is not less than A1. If so, obtaining the number of combination values corresponding to the second combination type, the third character sequence, and the third frequency values of the respective characters in the third character sequence, and continuing to determine whether the number of character types in the third character sequence is not less than A1. If not, adding up the number of combination values of all the obtained combination types and using the sum result as the characterization value of the number of combinations corresponding to the length of the character combination. The first combination type is the combination type formed by the first A1 characters at the front of the first character sequence, and the second combination type is the combination type formed by the first A1 characters at the front of the second character sequence; The method for obtaining the combined value corresponding to the first combined type, the second character sequence, the second frequency value of each character in the second character sequence, the combined value corresponding to the second combined type, the third character sequence, and the third frequency value of each character in the third character sequence includes: recording the first frequency value of the A1-th character in the first character sequence as the combined value corresponding to the first combined type; obtaining the first remaining frequency value of each character in the first character sequence, and recording the sequence obtained by removing the characters with the first remaining frequency value of 0 in the first character sequence as the second character sequence. The first remaining frequency value of the first A1 characters in the first character sequence is the first quantity difference of the corresponding characters, and the first remaining frequency value of the other characters located after the A1-th character in the first character sequence is the first frequency value of the corresponding characters. The first quantity difference of the first A1 characters in the first character sequence is the difference between the first frequency value of the corresponding characters and the combined value corresponding to the first combined type; recording the first remaining frequency value of each character in the second character sequence as the second frequency value of the corresponding character; recording the smallest second frequency value among the first A1 characters in the second character sequence as the combined value corresponding to the second combined type; obtaining the second remaining frequency value of each character in the second character sequence, and recording the sequence obtained by removing the characters with the second remaining frequency value of 0 in the second character sequence as the third character sequence. The second remaining frequency value of the first A1 characters in the second character sequence is the second quantity difference of the corresponding characters, and the second remaining frequency value of the other characters located after the A1-th character in the second character sequence is the second frequency value of the corresponding characters. The second quantity difference of the first A1 characters in the second character sequence is the difference between the second frequency value of the corresponding characters and the combined value corresponding to the second combined type; recording the second remaining frequency value of each character in the third character sequence as the third frequency value of the corresponding character; The method for obtaining the set of position span values includes: obtaining the position marker values of each character in the engineering project string, where the position marker value of the b1-th character in the engineering project string is b1; denoting the sequence constructed by all characters in the engineering project string that are the same as the first character in the first character sequence as the first sequence to be analyzed, and denoting the sequence constructed by all characters in the engineering project string that are the same as the second character in the first character sequence as the second sequence to be analyzed, and the characters in the sequence to be analyzed are arranged in ascending order of the position marker values; obtaining a marker difference sequence according to the absolute value of the difference between the position marker values of the characters at the same position in the first sequence to be analyzed and the second sequence to be analyzed, where the d1-th marker difference in the marker difference sequence is the absolute value of the difference between the position marker values of the d1-th character in the first sequence to be analyzed and the d1-th character in the second sequence to be analyzed, and the length of the marker difference sequence is consistent with the length of the shortest sequence to be analyzed between the first sequence to be analyzed and the second sequence to be analyzed; taking the maximum marker difference in the marker difference sequence as the maximum position span value; taking a preset first constant as the minimum position span value, denoting the interval formed by the minimum position span value and the maximum position span value as the position span interval, and denoting the set composed of all integers in the position span interval as the set of position span values; The method for obtaining the preference rate corresponding to each position span value includes: for the g0-th position span value in the position span interval: denoting the set of combined substrings corresponding to the g0-th character span value as the set to be analyzed, and denoting the set constructed by all types of combined substrings that appear in the set to be analyzed as the first type set; counting the frequencies of each type of combined substring in the first type set that appear in the set to be analyzed, and denoting them as the frequency values corresponding to the combined substring types; denoting the set constructed by all types of combined substrings with a frequency value of 1 in the first type set as the second type set, and denoting the set constructed by all types of combined substrings in the first type set except the second type set as the third type set; denoting the sum of the frequency values of all types of combined substrings in the second type set as the first frequency sum value, denoting the sum of the frequency values of all types of combined substrings in the third type set as the second frequency sum value, and denoting the normalized value of the result obtained by subtracting the first frequency sum value from the second frequency sum value as the first eigenvalue; denoting the set constructed by the frequency values of all types of combined substrings in the first type set as the frequency value set, and denoting the normalized value of the standard deviation of the frequency value set as the second eigenvalue; denoting the reciprocal of the total number of types of combined substrings in the second type set as the third eigenvalue; denoting the reciprocal of the total number of types of combined substrings in the first type set as the fourth eigenvalue; taking the mean of the first eigenvalue, the second eigenvalue, the third eigenvalue, and the fourth eigenvalue as the first preference value corresponding to the g0-th position span value; denoting the normalized value of the ratio of the target character combination length to the segmentation length value corresponding to the g0-th position span value as the second preference value corresponding to the g0-th position span value, and taking the mean of the first preference value and the second preference value as the preference rate corresponding to the g0-th position span value.

2. The efficient storage system for engineering project data based on cloud computing according to claim 1, characterized in that, The method for obtaining the set of combined substrings corresponding to each position span value includes: Denoting the value of the target character combination length as A2, and denoting the string formed by the first A2 characters in the first character sequence as the reference substring; For the g0-th position span value in the position span interval: Denoting the result of multiplying the g0-th position span value by A2 as the segmentation length value corresponding to the g0-th position span value; uniformly and non-overlappingly segmenting the engineering project string using the segmentation length value, and denoting all the segmented substrings as the substrings to be analyzed corresponding to the engineering project string under the g0-th position span value; Based on all the substrings to be analyzed corresponding to the engineering project string under the g0-th position span value and the reference substring, obtaining the target substring corresponding to each substring to be analyzed, and denoting the new string formed by the target substrings corresponding to all the substrings to be analyzed as the new string corresponding to the g0-th position span value; Combine the characters in the engineering project string according to the g0-th position span value and the target character combination length to obtain a first set; combine the characters in the new string corresponding to the g0-th position span value according to the g0-th position span value and the target character combination length to obtain a second set; obtain the gain characterization value corresponding to the g0-th position span value according to the first set and the second set, and determine whether the gain characterization value is greater than a preset gain threshold. If so, use the second set as the combined substring set corresponding to the g0-th position span value; otherwise, use the first set as the combined substring set corresponding to the g0-th position span value.

3. The efficient storage system for engineering project data based on cloud computing according to claim 2, characterized in that, The method for obtaining the target substring corresponding to the substring to be analyzed includes: For any substring g1 to be analyzed corresponding to the engineering project string under the g0-th position span value: Denote the string constructed by all the characters in the substring g1 to be analyzed that are the same as the characters in the reference substring as the feature string; according to the feature string, obtain all the special character addition positions in the substring g1 to be analyzed, and add special characters at all the special character addition positions. Denote the substring g1 after the special characters are added as the target substring corresponding to the substring g1 to be analyzed. The method for obtaining the special character addition positions in the substring g1 to be analyzed includes: Obtain all the local substrings corresponding to the feature string. If the string formed by the j1-th character to the j1+A2-1-th character in the feature string is the same as the reference substring, denote the string formed by the j1-th character to the j1+A2-1-th character as the local substring; denote all the characters in the substring g1 to be analyzed with the same position marker value as the characters in the local substring as the characters to be analyzed; for the k1-th character to be analyzed and the k1+1-th character to be analyzed in the substring g1 to be analyzed, if the absolute value of the difference between the position marker value of the k1-th character to be analyzed and the position marker value of the k1+1-th character to be analyzed is less than the g0-th position span value, the k1-th character to be analyzed is not the last character in the local substring, and the k1+1-th character to be analyzed is not the first character in the local substring, then arbitrarily select M0 positions between the k1-th character to be analyzed and the k1+1-th character to be analyzed as the special character addition positions, where M0 is the absolute value of the difference between the g0-th position span value and the difference characterization value, and the difference characterization value is the absolute value of the difference between the position marker value of the k1-th character to be analyzed and the position marker value of the k1+1-th character to be analyzed.

4. The efficient storage system for engineering project data based on cloud computing according to claim 3, wherein The method for obtaining the first set and the second set includes: Record the engineering project string as the first string to be read. Starting from the first character in the first string to be read, read the characters in the first string to be read to obtain the first read substring. Then, remove the characters that have been read from the first string to be read. Denote the string formed by the remaining characters after removal as the second string to be read. Determine whether the number of characters in the second string to be read is not less than , if so, starting from the first character in the second string to be read, read the characters in the second string to be read to obtain the second read substring. Remove the characters that have been read from the second string to be read. Denote the string formed by the remaining characters after removal as the third string to be read. Then, determine whether the number of characters in the third string to be read is not less than , if not, denote the third string to be read as the third read substring, stop obtaining read substrings, and denote the set constructed by all the obtained read substrings as the first set. The absolute value of the difference between the position marker values of two adjacent characters in the first read substring and the second read substring is the g0th position span value. The lengths of both the first read substring and the second read substring are the target character combination lengths. The method for obtaining the second set is the same as the method for obtaining the first set, and c2 is a preset second constant.

5. The efficient storage system for engineering project data based on cloud computing according to claim 4, wherein The method for obtaining the gain characterization value corresponding to the g0-th position span value includes: Count the total number of special characters added in all substrings to be analyzed, and denote it as the character addition quantity value; In the first set, denote the number of substrings with a frequency of 1 as the first quantity feature value, and denote the number of substrings with a frequency not equal to 1 as the second quantity feature value; in the second set, denote the number of substrings with a frequency of 1 as the third quantity feature value, and denote the number of substrings with a frequency not equal to 1 as the fourth quantity feature value; denote the result obtained by subtracting the first quantity feature value from the third quantity feature value as the first increase value, denote the result obtained by subtracting the second quantity feature value from the fourth quantity feature value as the second increase value, denote the sum of the first increase value and the character addition quantity value as the fifth quantity feature value, and denote Max(0, L1) as the gain characterization value corresponding to the g0th character span value, where L1 is the result obtained by subtracting the fifth quantity feature value from the second increase value.

Citation Information

Patent Citations

  • Data compression method and server

    CN113765854A

  • Gas alarm system data storage method based on Internet of Things platform

    CN116015312A