Distributed Clinical Trial Data Optimization Storage Method
By independently encoding and hierarchical storage of clinical trial data, the problem that LZW encoding cannot decode fields separately is solved, and the cost of data query and efficiency are reduced.
Patent Information
- Application Number
- CN202510525578.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The existing LZW encoding methods cannot independently decode individual fields in clinical trial data, resulting in high query cost and inefficient efficiency.
The fields of clinical trial data are converted into a character sequence, and the encoded objects belonging to multiple fields are not encoded during the compression process. The identification of the end of the field encoding is set, the encoding results and identification are stored layer by layer, and the decoding path is constructed to realize the individual decoding of the fields.
Through independent encoding and layered storage of fields, the amount of data is reduced and the query efficiency and query speed of clinical trial data is improved.
Smart Images

Figure CN120048409B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical information processing, and particularly to a method for optimizing the storage of distributed clinical trial data. Background Art
[0002] With the rapid development of modern medical research, the scale and complexity of clinical trial data have increased exponentially. Therefore, distributed storage is usually adopted for clinical trial data to meet the growing data requirements of clinical trials.
[0003] However, the amount of data in each data node of distributed storage is still extremely large. To improve the storage efficiency of data, the data information in each data node is usually stored after being compressed. For example, a patent document with the publication number "CN116631550B" and the name "A Method for Data Management and Logical Verification of Clinical Trials and Its Medical System" discloses that: an original dictionary for each category is respectively constructed according to the distribution of different strings in each category, and an initial dictionary for each category is obtained according to the repetition degree of different strings in the original dictionary of each category; according to the semantic information of the string combinations formed by the strings in the initial dictionary of each category and the suffix characters in the clinical trial data of the category, the number of suffix characters is iteratively judged and the dictionary is updated, and the clinical trial data of each category is compressed according to the updated dictionary.
[0004] The above method is an improvement of the LZW coding. Although it can improve the compression efficiency of clinical trial data, due to the characteristic of the LZW coding that the dictionary is updated while compressing, when querying a certain field in the clinical trial data, it is impossible to decode only a single field, but all the compressed data in the data node needs to be completely decompressed, resulting in a high query cost and low query efficiency for clinical trial data. Summary of the Invention
[0005] To solve the problem that the LZW coding cannot decode only a single field, resulting in a high query cost and low query efficiency for clinical trial data, the present invention provides a method for optimizing the storage of distributed clinical trial data, including:
[0006] Converting all fields in a piece of data information of clinical trial data into a character sequence to be compressed; during the process of compressing the character sequence by using the LZW coding, in response to the character included in the coding object belonging to at least two fields, not encoding the coding object, and forming a new coding object with all the characters belonging to the first field among the at least two fields in the coding object; otherwise, encoding the coding object; setting an identifier for indicating the end of field encoding for each field, and storing the encoding results and identifiers corresponding to all fields in layers in sequence to obtain compressed data, where each layer of the compressed data contains at most one encoding result or identifier of each field.
[0007] The present invention does not encode the encoding object whose contained characters belong to at least two fields, but splits the encoding object into new encoding objects whose contained characters belong to only one field, avoiding the common encoding of characters in different fields into one encoding result, which may lead to the inability to decode the fields separately, and providing a basis for subsequent realization of separate field decoding; the present invention stores the encoding results and identifiers corresponding to all fields in layers, ensuring that the encoding results corresponding to the fields can be located during decoding, realizing separate field decoding, reducing the amount of decompressed data, accelerating the decompression speed, reducing the query cost of clinical trial data, and improving the query efficiency of clinical trial data.
[0008] Preferably, the conversion of all fields in a data message of clinical trial data into a character sequence to be compressed includes: encoding each field in the data message into binary data, dividing each group of k bits of the binary data corresponding to each field into a group, and converting each group of binary data into a decimal number; regarding a decimal number as a character, then each field corresponds to at least one character, and the characters corresponding to all fields are combined into a character sequence to be compressed in the order of the fields, where k is a preset length.
[0009] The present invention converts the data message containing multiple data types into a character sequence containing only one data type, increasing the repeatability of the data, thereby improving the compression efficiency of the data message.
[0010] Preferably, the hierarchical storage of the encoding results and identifiers corresponding to all fields in order to obtain compressed data includes: taking each encoding result and identifier of the field as a compression element of the field; storing the i-th compression element of each field as the i-th layer in the order of the fields, and when the i-th compression element of a certain field does not exist, the field does not participate in the storage of the i-th layer; where i represents the layer number; when all compression elements of all fields are stored, all the obtained layers constitute the compressed data.
[0011] The number of encoding results corresponding to each field is different. The present invention stores the encoding results and identifiers corresponding to all fields in layers, ensuring that the encoding results corresponding to the fields can be located during decoding, thereby realizing separate field decoding.
[0012] Preferably, it further includes: in response to the storage system receiving query information, determining, according to the identifiers in each layer of the compressed data, each encoding result corresponding to the query field to which the query information belongs and the decoding path of each encoding result, where the decoding path only includes other encoding results that must be decoded when decoding the encoding result, decoding each encoding result corresponding to the query field according to the decoding path to obtain decoded information; in response to the query information being consistent with the decoded information, decoding the compressed data to obtain a query result.
[0013] The decoding path constructed in the present invention only includes other encoding results that must be decoded when decoding the encoding result, avoiding the need to decode all encoding results when decompressing the encoding result corresponding to the field, reducing the amount of decompressed data, and at the same time accelerating the decompression speed, thereby reducing the query cost of clinical trial data and improving the query efficiency of clinical trial data.
[0014] Preferably, the determining each encoding result corresponding to the query field to which the query information belongs includes: locating the compressed elements corresponding to the query field in each layer: ; where represents the position serial number of the compressed element of the query field in the i-th layer; represents the position serial number of the encoding result of the query field in the (i - 1)-th layer; represents the number of identifiers that appear before the -th element in the (i - 1)-th layer; h represents the position serial number of the query field in the data information; when the -th element in the i-th layer is not an identifier, the -th element in the i-th layer is the encoding result of the query field in the i-th layer, and at this time, continue to obtain the compressed element of the query field in the next layer; when the -th element in the i-th layer is an identifier and the -th element in the i-th layer is the identifier of the query field, at this time, do not obtain the compressed element of the query field in the next layer.
[0015] Preferably, setting a identifier for each field to indicate the end of field encoding includes: setting a unified first identifier for each field to indicate the end of field encoding; the encoding of the encoding object further includes: every time an encoding object is encoded, updating the dictionary of the LZW encoding according to the encoding object.
[0016] Preferably, the method for obtaining the decoding path is: sequentially numbering all the encoding results in the compressed data in the order from top to bottom and from left to right, and whenever the first identifier is encountered, jump to the first layer of the compressed data and continue numbering from the first unnumbered encoding result in the first layer; taking the encoding result to be decoded as the first target encoding result , in response to the j-th target encoding result Greater than , add the th coding result to the path sequence, and use the th coding result as the (j + 1)-th target coding result ; otherwise, add the j-th target coding result to the path sequence; where represents the length of the initial dictionary; reverse the path sequence and use it as the decoding path of the coding result to be decoded.
[0017] Preferably, a flag for indicating the end of field coding is set for each field, including: during the process of compressing a character sequence using LZW coding, in response to the characters included in the coding object belonging to at least two fields, setting a second flag for indicating not to update the dictionary for the first field among the at least two fields; after the field compression ends, in response to the non-existence of the second flag in the field, setting a third flag for indicating to update the dictionary for the field; the second flag and the third flag can be used to indicate the end of field coding; the encoding of the coding object further includes: for each coding object encoded, in response to the string formed by the coding object and its next character in the character sequence existing in the dictionary, not updating the dictionary; otherwise, updating the dictionary according to the coding object.
[0018] When the present invention encodes a coding object, it selectively updates the dictionary, avoiding the appearance of duplicate strings in the dictionary, reducing the length of the dictionary, and thus improving the compression efficiency of data information. At the same time, the present invention sets a second flag and a third flag to indicate whether the dictionary has been updated, ensuring that during subsequent decoding, the decoding path corresponding to the coding result of the field can be accurately obtained.
[0019] Preferably, the method for obtaining the decoding path is: sequentially put the elements in the compressed data into the data sequence in the order from top to bottom and from left to right. During the putting process, whenever a flag is put into the data sequence, jump to the first layer of the compressed data; use the coding result to be decoded as the first target coding result ; in response to the j-th target coding result not being greater than , add to the path sequence; otherwise, in response to the non-existence of the second flag before the th coding result in the data sequence, add the th coding result to the path sequence, and use the th coding result in the data sequence as the (j + 1)-th target coding result ; in response to the existence of the second flag before the th coding result, add the th coding result in the data sequence to the path sequence, and add the The j+1 target coding result is the obtained coding result ; represents the length of the initial dictionary; represents the number of the second identifier before the th coding result; Reverse the path sequence to be the decoding path of the coding result to be decoded.
[0020] Preferably, decoding the respective coding results corresponding to the query field to obtain decoding information, including: using the decoding path of the coding result to be decoded as the first path, using the decoding path of each element in the first path as the second path, representing the first element in the second path with D, and using the element with the serial number D in the initial dictionary as the decoding element of the second path; concatenating the decoding elements of all the second paths into the decoding string of the coding result to be decoded; concatenating the decoding strings corresponding to all the coding results of the query field together to obtain the decoding result of the query field; converting the decoding result into field information to obtain the decoding information.
[0021] The present invention has the following technical effects: The present invention avoids jointly encoding the characters in different fields into one coding result, providing a basis for realizing separate decoding of fields;
[0022] Further, through hierarchical storage, the present invention ensures that the coding results corresponding to the fields can be located during decoding, thereby realizing separate decoding of fields;
[0023] Further, by constructing the decoding path, the present invention reduces the amount of decompressed data, speeds up the decompression speed, reduces the query cost of clinical trial data, and improves the query efficiency of clinical trial data. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flowchart of the method in the distributed clinical trial data optimization storage method according to an embodiment of the present invention;
[0025] Figure 2 is a schematic diagram of the order of numbering the coding results in the compressed data of Table 5;
[0026] Figure 3 is a schematic diagram of the order of putting the elements in the compressed data of Table 6 into the data sequence. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] The objective of the present invention is to optimize the storage of clinical trial data, and to separately decode a certain field in the stored compressed data, so as to improve the retrieval efficiency of clinical trial data. To achieve the objective, the following three points need to be done: First, independently encode the fields, such as step S2 of the present invention; Second, store in layers, such as step S3 of the present invention; Third, construct a decoding path for the encoding result corresponding to the field, such as step S4 of the present invention.
[0029] Among them, independently encoding the fields is to avoid jointly encoding the characters in different fields into one encoding result during the compression process, which may lead to the inability to separately decode the fields; storing in layers is to be able to locate the encoding result corresponding to the field during decoding; constructing the decoding path is to avoid decoding all the encoding results before the encoding result corresponding to the field.
[0030] The following details the distributed clinical trial data optimized storage method disclosed in the embodiments of the present invention, referring to Figure 1 , including steps S1 - S5:
[0031] S1: Convert each data information of the clinical trial data into a character sequence to be compressed.
[0032] The clinical trial data contains multiple data information. Some data information includes the trial start time, the trial end time, the information of the test products, such as the generic name, specification, source, batch number, expiration date, and storage conditions of the test drugs or medical devices, etc. Some data information includes the basic information of the subjects, such as gender, age, weight, height, baseline data, disease information, etc. Some data information includes the trial process, such as the usage, dosage, time of the test drugs or medical devices, etc. Some data information includes the trial results, such as the changes of each index of the subjects compared with the baseline data, the occurrence time, severity, and detailed description of adverse events, etc.
[0033] The data volume of the clinical trial data is huge and needs to be compressed and stored. Currently, the commonly used compression algorithms include entropy encoding such as Huffman coding and arithmetic coding, and dictionary-based compression algorithms such as LZW coding and LZ77 coding. The compression efficiency of these compression algorithms depends on the repeatability of the data. Since a data information contains multiple data types, such as the trial start time and the trial end time are dates, the gender of the subject is a character, the age is an integer, and the weight and height are floating-point numbers. The repeatability of the data in each data information is very low, and it is difficult to achieve good compression efficiency using the above compression algorithms. Therefore, the present invention preprocesses the data information to increase the repeatability of the data in the data information, so as to improve the compression efficiency of the data information.
[0034] Specifically, for each data information in the clinical trial data, each field in the data information encoding is encoded into binary data using GB2312 encoding. Each group of k bits of the binary data corresponding to each field is divided into a group, and each group of binary data is converted into a decimal number. Regarding a decimal number as a character, each field corresponds to at least one character. The characters corresponding to all fields are formed into a character sequence to be compressed in the order of the fields. Here, k is a preset length. It should be noted that the encoding length of GB2312 for each letter or number is 8 bits, and the encoding length for each Chinese character is 16 bits. Therefore, the length of the binary data corresponding to each field is a multiple of 8. To ensure that the binary data corresponding to each field can be completely divided, k needs to be selected as a factor of 8 other than 1, such as 2, 4, 8, etc., and the implementer can set it according to the actual implementation situation.
[0035] For example: The data information {Amlodipine Hydrochloride Tablets, 5mg / tablet, a certain pharmaceutical company, 20240730, July 2026, store in a cool and dry place, protected from light} contains the fields of the generic name, specification, source, batch number, expiration date, and storage conditions of the test drug. The result of encoding the field "Amlodipine Hydrochloride Tablets" using GB2312 is 1101000111001110110010111110000110110000101100011100001011001000101101011101100011000110101111011100011010101100. If k is set to 8, the corresponding characters for "Amlodipine Hydrochloride Tablets" are {209, 206, 203, 225, 176, 177, 194, 200, 181, 216, 198, 189, 198, 172}. Similarly, the corresponding characters for "5mg / tablet" are {53, 109, 103, 47, 198, 172}, the corresponding characters for "a certain pharmaceutical company" are {196, 179, 214, 198, 210, 169, 211, 208, 207, 222, 185, 171, 203, 190}, the corresponding characters for "20240730" are {50, 48, 50, 52, 48, 55, 51, 48}, the corresponding characters for "July 2026" are {50, 48, 50, 54, 196, 234, 48, 55, 212, 194}, and the corresponding characters for "store in a cool and dry place, protected from light" are {210, 245, 193, 185, 184, 201, 212, 239, 180, 166, 163, 172, 177, 220, 185, 226, 177, 163, 180, 230}. Then the character sequence to be compressed is {209, 206, 203, 225, 176, 177, 194, 200, 181, 216, 198, 189, 198, 172, 53, 109, 103, 47, 198, 172, 196, 179, 214, 198, 210, 169, 211, 208, 207, 222, 185, 171, 203, 190, 50, 48, 50, 52, 48, 55, 51, 48, 50, 48, 50, 54, 196, 234, 48, 55, 212, 194, 210, 245, 193, 185, 184, 201, 212, 239, 180, 166, 163, 172, 177, 220, 185, 226, 177, 163, 180, 230}.
[0036] It should be noted that the present invention is described only by taking GB2312 encoding as an example. Implementers can select encoding methods according to the actual implementation situation, such as GBK encoding, UTF-8 encoding, etc.
[0037] In addition, implementers can also use other methods to increase the repeatability of data in the data information. For example, first encode each field in the data information into binary data, and then use Base64 encoding to encode the binary data corresponding to each field into multiple characters, and form a character sequence to be compressed from the characters corresponding to all fields.
[0038] Thus, the data information in the clinical trial data is converted into a character sequence to be compressed.
[0039] S2: Compress the character sequence to be compressed to obtain multiple encoding results corresponding to each field in the data information.
[0040] It should be noted that the LZW encoding is a dictionary-based compression algorithm, but the dictionary used for compression is not stored in the LZW encoding. During the decompression process, the dictionary is constructed while decompressing. When querying clinical trial data, it is necessary to decode a certain field in the data information. Due to the characteristic that the dictionary is constructed while decompressing, it is necessary to decode the compressed data corresponding to the data information from the beginning, resulting in low efficiency. At the same time, the characters corresponding to two or more fields in the data information may be encoded into one encoding result, making it impossible to decode a single field separately. Therefore, the present invention improves the LZW encoding method so that a certain field in the data information can be decoded separately later, improving the query efficiency of clinical trial data.
[0041] In one embodiment, use LZW encoding to compress the character sequence to be compressed. During the compression process, judge the current encoding object:
[0042] If the characters included in the encoding object belong to two or more fields, do not encode the encoding object, and form a new encoding object from all the characters of the first field among the two or more fields to which the encoding object belongs; if the characters included in the encoding object belong to only one field, use the serial number corresponding to the same character or string as the encoding object in the dictionary to encode the encoding object to obtain an encoding result. At this time, add the string formed by the encoding object and the next character of it in the character sequence to the end of the dictionary to update the dictionary.
[0043] For example, when the data information is {Amlodipine Hydrochloride Tablets, 5 mg / tablet, a certain pharmaceutical company, 20240730, July 2026, store in a cool and dry place, protected from light}, the corresponding character sequence to be compressed is {209, 206, 203, 225, 176, 177, 194, 200, 181, 216, 198, 189, 198, 172, 53, 109, 103, 47, 198, 172, 196, 179, 214, 198, 210, 169, 211, 208, 207, 222, 185, 171, 203, 190, 50, 48, 50, 52, 48, 55, 51, 48, 50, 50, 54, 196, 234, 48, 55, 212, 194, 210, 245, 193, 185, 184, 201, 212, 239, 180, 166, 163, 172, 177, 220, 185, 226, 177, 163, 180, 230}. When the coding object is the 42nd to 43rd characters in the character sequence to be compressed {48, 50}, where {48} belongs to the character corresponding to the batch number field "20240730" and {50} belongs to the character corresponding to the expiration date field "July 2026", then the coding object {48, 50} is not encoded, and {48} is used as the new coding object. The new coding object {48} only belongs to the character corresponding to the batch number field "20240730". At this time, the new coding object {48} is encoded, and the string {48, 50} formed by {48} and its next character {50} in the character sequence to be compressed is added to the end of the dictionary to update the dictionary.
[0044] Set a first identifier for each field to indicate the end of encoding for that field. To distinguish the encoding result from the first identifier, the set first identifier cannot be the same as the sequence numbers of all elements in the dictionary. The first identifier can be set to 0, or a number greater than L, where L represents the length of the finally obtained dictionary.
[0045] For example, when the data information is {Amlodipine Hydrochloride Tablets, 5 mg / tablet, a certain pharmaceutical company, 20240730, July 2026, store in a cool and dry place, protected from light}, the corresponding character sequence to be compressed is {209, 206, 203, 225, 176, 177, 194, 200, 181, 216, 198, 189, 198, 172, 53, 109, 103, 47, 198, 172, 196, 179, 214, 198, 210, 169, 211, 208, 207, 222, 185, 171, 203, 190, 50, 48, 50, 52, 48, 55, 51, 48, 50, 50, 54, 196, 234, 48, 55, 212, 194, 210, 245, 193, 185, 184, 201, 212, 239, 180, 166, 163, 172, 177, 220, 185, 226, 177, 163, 180, 230}. The first identifier is set to 0. The character sequence to be compressed is compressed by the method in this embodiment, and the obtained dictionary is shown in Table 1. The encoded results corresponding to each field and the first identifier are shown in Table 2.
[0046] Table 1
[0047]
[0048] Table 2
[0049]
[0050] It should be noted that inside a computer, data is stored in binary form. Therefore, after the LZW encoding ends, the compression result needs to be converted to binary for storage. In LZW encoding, when the length of the dictionary increases, it may cause the number of bits for converting each encoded result to binary to increase, thus affecting the compression efficiency. For example, when the length of the dictionary is less than 64, each encoded result is less than 64, and each encoded result can be represented by 6 - bit binary. When the length of the dictionary is greater than 64 and less than 128, each encoded result may be greater than 64, and 7 - bit binary is needed to represent each encoded result. In the above embodiment, the string formed by the new encoding object and its next character in the character sequence already exists in the dictionary. After encoding the new encoding object, the string formed by the new encoding object and its next character in the character sequence is added to the end of the dictionary, resulting in the existence of the same elements in the dictionary, increasing the length of the dictionary and thus affecting the compression efficiency. Therefore, in another embodiment of the present invention, during the process of compressing the character sequence to be compressed using LZW encoding, the string formed by the encoding object and its next character in the character sequence is selectively added to the dictionary to make the length of the dictionary as short as possible, thereby improving the compression efficiency.
[0051] In another embodiment, the character sequence to be compressed is compressed using LZW coding. During the compression process, the current coding object is judged:
[0052] If the characters included in the coding object belong to two or more fields, the coding object is not encoded. All the characters of the first field among the two or more fields included in the coding object form a new coding object. At this time, the serial number corresponding to the same character or string as the new coding object in the dictionary is used to encode the new coding object to obtain a coding result. At this time, the dictionary is not updated, and a second identifier is set for the field corresponding to the new coding object. The second identifier is used to indicate that the dictionary has not been updated.
[0053] If the characters included in the coding object belong to only one field, the serial number corresponding to the same character or string as the coding object in the dictionary is used to encode the coding object. At this time, the dictionary is updated, and the update method is the same as the method in LZW coding, that is, the string formed by the coding object and its next character in the character sequence is added to the end of the dictionary to realize the update of the dictionary.
[0054] After all the characters corresponding to a certain field are compressed, if the field does not have a second identifier, a third identifier is set for the field. The third identifier is used to indicate that the dictionary has been updated during the compression of the characters included in the field.
[0055] The second identifier and the third identifier can both indicate the end of the encoding of the field. In order to distinguish the coding result, the second identifier and the third identifier, the set second identifier and third identifier cannot be the same, and cannot be the same as the serial numbers of all elements in the dictionary. The second identifier and the third identifier can be set to 0, or a number greater than L, where L represents the length of the finally obtained dictionary.
[0056] For example, when the data information is {Amlodipine Hydrochloride Tablets, 5mg / tablet, a certain pharmaceutical company, 20240730, July 2026, store in a cool and dry place, protected from light}, the corresponding character sequence to be compressed is {209, 206, 203, 225, 176, 177, 194, 200, 181, 216, 198, 189, 198, 172, 53, 109, 103, 47, 198, 172, 196, 179, 214, 198, 210, 169, 211, 208, 207, 222, 185, 171, 203, 190, 50, 48, 50, 52, 48, 55, 51, 48, 50, 48, 50, 54, 196, 234, 48, 55, 212, 194, 210, 245, 193, 185, 184, 201, 212, 239, 180, 166, 163, 172, 177, 220, 185, 226, 177, 163, 180, 230}. The second identifier is set to L + 1, and the third identifier is set to 0, where L represents the length of the finally obtained dictionary. During the process of compressing the character sequence to be compressed by the method in this embodiment, the constructed dictionary is shown in Table 3, and the corresponding encoding results and identifiers of each field are shown in Table 4.
[0057] Table 3
[0058]
[0059] Table 4
[0060]
[0061] Thus, the compression of the character sequence to be compressed is achieved, and multiple encoding results and identifiers corresponding to each field in the data information are obtained.
[0062] S3: Hierarchically store the multiple encoding results and identifiers corresponding to each field in the data information to obtain compressed data.
[0063] It should be noted that in traditional compression methods, the obtained coding results are stored in sequence to form a sequence. Since the number of coding results corresponding to each field is different, for example, the coding results corresponding to the general name field "Amlodipine Hydrochloride Tablets" are 14, and the coding objects corresponding to the specification field "5mg / tablet" are 5. When decoding, the positions of the coding results corresponding to each field cannot be known in the compressed data, resulting in the inability to decode a field separately. Therefore, the present invention stores the coding results and identifiers of all fields in layers, so that when decoding, the positions of the coding results corresponding to each field can be accurately located in the compressed data, so that each field can be decoded separately, avoiding decoding the entire compressed data when querying clinical trial data, thereby improving the query efficiency.
[0064] Hierarchical storage means that according to the order of the coding results of the fields, the coding results are sequentially placed in each layer for storage; that is to say, the first coding result of each field is placed in the first layer, the second coding result of each field is placed in the second layer, the third coding result of each field is placed in the third layer, and so on. The following is an example for easy understanding.
[0065] Specifically, after the compression of the character sequence to be compressed is completed, each field corresponds to at least one coding result and one identifier, and all the coding results and identifiers of the field are used as a compression element of the field. The i-th compression element of each field is stored as the i-th layer information in the order of the fields. When the i-th compression element of a certain field does not exist, the field does not participate in the construction of the i-th layer information.
[0066] When all the compression elements of all fields are stored, all the layer information obtained constitutes the compressed data of the data information. It should be noted that the identifier can indicate the end of the coding of the field. In the compressed data, the layer where the identifier is located can also reflect the number of coding results corresponding to a field. For example, if the identifier of a certain field is located in the 3rd layer of the compressed data, it means that the field corresponds to two coding results.
[0067] For example: The coding results and identifiers corresponding to each field in Table 2 are stored in layers, and the obtained compressed data is shown in Table 5. Each row in Table 5 represents one layer of information.
[0068] Table 5
[0069]
[0070] The coding results and identifiers corresponding to each field in Table 4 are stored in layers, and the obtained compressed data is shown in Table 6. Each row in Table 6 represents one layer of information.
[0071] Table 6
[0072]
[0073] In the present invention, the compressed data of all data information is stored distributively, so that the compressed data is stored on multiple data nodes, thereby ensuring the reliability and availability of the data. It should be noted that the hierarchical storage in the present invention refers to hierarchically storing the multiple encoding results and identifiers corresponding to each field in a piece of data information on one data node.
[0074] S4: Determine the respective encoding results corresponding to the query fields to which the query information belongs according to the identifiers in each layer of the compressed data.
[0075] When a staff member needs to query the information of a certain clinical trial data, the query information is input on the query system, and the field to which the query information belongs is called the query field. The query system synchronizes the query information to the data nodes containing the query field, and the data nodes only decode the query field in each compressed data. When decoding, it is first necessary to locate the encoding results corresponding to the query field in each layer of the compressed data.
[0076] It should be noted that the positions of the encoding results corresponding to the query field in each layer of the compressed data are determined by the rules of hierarchical storage. That is to say, the implementer can infer the positions of the encoding results corresponding to the query field in each layer of the compressed data according to the rules of hierarchical storage. A reasoning method is given below.
[0077] Specifically, for any piece of compressed data, obtain the encoding results corresponding to the query field in the information of each layer of the compressed data:
[0078] ;
[0079] where represents the position serial number of the compressed element of the query field in the i-th layer, that is, in the information of the i-th layer, the compressed element corresponding to the query field is the -th element in the information of the i-th layer; represents the position serial number of the encoding result of the query field in the (i - 1)-th layer, that is, in the information of the (i - 1)-th layer, the encoding result corresponding to the query field is the -th element in the information of the (i - 1)-th layer; represents the number of identifiers that appear before the -th element in the (i - 1)-th layer; represents the position serial number of the query field in the data information, that is, the -th field in the data information is the query field.
[0080] When the -th element in the i-th layer is not an identifier, the The element is the encoding result of the query field at the i-th layer. At this time, continue to obtain the compressed element of the query field in the next layer; when the element at the is an identifier, the element at the
[0081] in the i-th layer is the identifier of the query field. At this time, do not obtain the compressed element of the query field in the next layer.
[0082] It should be noted that when step S2 uses the first identifier, the identifier in this step refers to the first identifier. When step S2 uses the second identifier and the third identifier, the identifier in this step includes the second identifier and the third identifier.
[0083] S5: Obtain the decoding paths of the encoding results corresponding to the query field, decode the encoding results according to the decoding paths to obtain decoding information, and obtain the query results according to the decoding information.
[0084] It should be noted that the numerical value of the encoding result represents the serial number of its corresponding string in the dictionary. Since the dictionary has not been constructed, it cannot be directly decoded. In LZW encoding, the serial number of the string in the dictionary represents the order in which the string is added to the dictionary. The previous encoding object when the string is added to the dictionary is the string composed of the first S-1 characters of the string, where S represents the length of the string, and the first character of the next encoding object is the last character of the string. The string corresponding to the encoding result can be decoded according to the previous encoding object and the next encoding object when the string corresponding to the encoding result is added to the dictionary. According to the serial number of the string in the dictionary and the second identifier, the encoding result corresponding to the previous encoding object when the string is added to the dictionary and the encoding result corresponding to the next encoding object can be obtained. Therefore, for the encoding results corresponding to the query field, the present invention decodes the encoding results by constructing the decoding paths of the encoding results, and the decoding paths include other encoding results that must be decoded to decode the encoding results.
[0085] In one embodiment, when step S2 uses the first identifier, the method for obtaining the decoding path is as follows:
[0086] Number all the encoding results in the compressed data in the order from top to bottom and from left to right. Whenever an identifier is encountered, jump to the first layer of the compressed data, and start from the first unnumbered encoding result in the first layer, and continue to number the unnumbered encoding results in the compressed data in the order from top to bottom and from left to right. Stop when all the encoding results have been numbered. The number of the encoding result represents the encoding order of the encoding object corresponding to the encoding result in the process of compressing the character sequence to be encoded. For example, the schematic diagram of the numbering order of the encoding results in the compressed data in Table 5 is shown inFigure 2 , Figure 2 The medium gray one is the first identifier, and no numbering is required for the first identifier.
[0087] Construct an empty path sequence. Denote the currently decoded coding result as the first target coding result, and represent the first target coding result with When is less than or equal to , add to the path sequence, where represents the length of the initial dictionary; when is greater than , add the coding result with the number to the path sequence, and use the coding result with the number as the second target coding result, represented by ;
[0088] When the second target coding result is less than or equal to , add to the path sequence; when the second target coding result is greater than , add the coding result with the number to the path sequence, and use the coding result with the number as the third target coding result, represented by ;
[0089] When the third target coding result is less than or equal to , add to the path sequence; when the third target coding result is greater than , add the coding result with the number to the path sequence, and use the coding result with the number as the fourth target coding result, represented by ;
[0090] And so on, until a target coding result is added to the path sequence, then stop the iteration.
[0091] Reverse the order of the finally obtained path sequence, and use the result of the reverse order as the decoding path of the currently decoded coding result.
[0092] For example: The sequence formed by all the encoded objects in the compressed data of Table 5 in ascending order of their numbers is {35, 32, 31, 43, 16, 17, 26, 29, 20, 40, 28, 23, 28, 15, 6, 10, 9, 1, 61, 27, 18, 39, 28, 36, 13, 37, 34, 33, 42, 22, 14, 31, 24, 3, 2, 3, 5, 2, 8, 4, 2, 82, 3, 7, 27, 46, 86, 38, 26, 36, 48, 25, 22, 21, 30, 38, 47, 19, 12, 11, 15, 17, 41, 22, 44, 17, 11, 19, 45}. The initial dictionary can be seen in Table 7, and the length of the initial dictionary is 48. The numbers of the encoded objects corresponding to the validity period field are from 42 to 49. Among them, the encoded object with the number 42 is {82}. Add the encoded object {2} with the number 82 - 48 + 1 = 35 to the path sequence. Since the encoded object {3} with the number 82 - 48 = 34 is less than 48, add the encoded object {3} with the number 82 - 48 = 34 to the path sequence. Then the final path sequence is {2, 3}, and the decoding path is {3, 2}.
[0093] Table 7
[0094]
[0095] In another embodiment, when step S2 uses the second identifier and the third identifier, the method for obtaining the decoding path is as follows:
[0096] Construct an empty data sequence, and put all the elements in the compressed data into the data sequence in the order from top to bottom and from left to right. Whenever an identifier is encountered, after putting the identifier into the data sequence, jump to the first layer of the compressed data, and start from the first encoded result that has not been put into the data sequence in the first layer, and continue to put the elements in the compressed data that have not been put into the data sequence into the data sequence in the order from top to bottom and from left to right. Stop until all the elements in the compressed data have been put into the data sequence. For example, the schematic diagram of the order of putting the elements in the compressed data of Table 6 into the data sequence can be seen in Figure 3 , Figure 3The second identifier and the third identifier are in medium gray. The second identifier and the third identifier also need to be put into the data sequence. The finally obtained data sequence is {35, 32, 31, 43, 16, 17, 26, 29, 20, 40, 28, 23, 28, 15, 0, 6, 10, 9, 1, 61, 0, 27, 18, 39, 28, 36, 13, 37, 34, 33, 42, 22, 14, 31, 24, 0, 3, 2, 3, 5, 2, 8, 4, 2, 116, 82, 3, 7, 27, 46, 86, 38, 26, 0, 36, 48, 25, 22, 21, 30, 38, 47, 19, 12, 11, 15, 17, 41, 22, 44, 17, 11, 19, 45, 0}.
[0097] Construct an empty path sequence. Denote the currently decoded encoding result as the first target encoding result, and represent the first target encoding result with When is less than or equal to , add to the path sequence, where represents the length of the initial dictionary; when is greater than , obtain the th encoding result in the data sequence. If there is no second identifier before the th encoding result in the data sequence, add the th encoding result in the data sequence to the path sequence, and use the th encoding result in the data sequence as the second target encoding result, which is represented by ; if there is a second identifier before the th encoding result in the data sequence, add the th encoding result in the data sequence to the path sequence, and use the th encoding result in the data sequence as the second target encoding result, which is represented by , where represents the number of second identifiers before the th encoding result in the data sequence;
[0098] When the second target encoding result is less than or equal to , add to the path sequence; when the second target encoding result is greater than , if there is no second identifier before the th encoding result in the data sequence, add the th encoding result in the data sequence to the path sequence, and use the The third target coding result is represented by ; if there is a second identifier before the th coding result in the data sequence, then the th coding result in the data sequence is added to the path sequence, and the th coding result in the data sequence is used as the third target coding result, represented by , where represents the number of second identifiers before the
[0099] th coding result in the data sequence; When the third target coding result is less than or equal to , is added to the path sequence; when the third target coding result is greater than , if there is no second identifier before the th coding result in the data sequence, then the th coding result in the data sequence is added to the path sequence, and the th coding result in the data sequence is used as the fourth target coding result, represented by ; if there is a second identifier before the th coding result in the data sequence, then the th coding result in the data sequence is added to the path sequence, and the th coding result in the data sequence is used as the fourth target coding result, represented by , where represents the number of second identifiers before the
[0100] And so on, until a target coding result is added to the path sequence and the iteration stops.
[0101] The finally obtained path sequence is reversed, and the reversed result is used as the decoding path of the currently to-be-decoded coding result.
[0102] For example: The data sequence formed by all elements in the compressed data of Table 6 is {35, 32, 31, 43, 16, 17, 26, 29, 20, 40, 28, 23, 28, 15, 0, 6, 10, 9, 1, 61, 0, 27, 18, 39, 28, 36, 13, 37, 34, 33, 42, 22, 14, 31, 24, 0, 3, 2, 3, 5, 2, 8, 4, 2, 116, 82, 3, 7, 27, 46, 86, 38, 26, 0, 36, 48, 25, 22, 21, 30, 38, 47, 19, 12, 11, 15, 17, 41, 22, 44, 17, 11, 19, 45, 0}. The initial dictionary is shown in Table 7. The length of the initial dictionary is 48. The validity period field corresponds to the 42nd to 49th coded objects in the data sequence. Among them, the second identifier and the third identifier in the data sequence do not participate in the counting. The 47th coded object is {86}. Since there is no second identifier 116 before the 39th coded object {8} in the data sequence, that is, 86 - 48 + 1 = 39, the 39th coded object {8} in the data sequence is added to the path sequence. For the 38th coded object {2} in the data sequence, that is, 86 - 48 = 38, since the 38th coded object {2} in the data sequence is less than 48, the 38th coded object {2} in the data sequence is added to the path sequence. Then the final path sequence is {8, 2}, and the decoding path is {2, 8}.
[0103] Thus, the decoding path is obtained.
[0104] Decode the coding result according to the decoding path, specifically:
[0105] Take the decoding path of the coding result to be decoded as the first path, take the decoding path of each element in the first path as the second path, represent the first element in the second path with D, and take the element with serial number D in the initial dictionary as the decoding element of the second path. Concatenate the decoding elements of all second paths into the decoding string of the coding result to be decoded.
[0106] For example: In the compressed data of Table 5, the decoding paths of each element in the decoding path {3, 2} of the coded object {82} with serial number 42 are {3} and {2} respectively. The initial dictionary is shown in Table 7. Then the elements with serial numbers 3 and 2 in the initial dictionary are {50} and {48} respectively. Then the decoding string is {50, 48}.
[0107] In the compressed data of Table 6, the decoding paths of each element in the decoding path {2, 8} of the 47th coded object {86} are {2} and {8} respectively. The initial dictionary is shown in Table 7. Then the elements with serial numbers 2 and 8 in the initial dictionary are {48} and {55} respectively. Then the decoding string is {48, 55}.
[0108] In one embodiment, the decoded strings corresponding to all the encoding results of the query field are concatenated together to obtain the decoded result of the query field. Each character in the decoded result is actually a decimal number. Each decimal number in the decoded result is respectively converted into a binary number. All the obtained binary numbers are concatenated together, and the concatenated binary numbers are decoded using GB2312 encoding to obtain the decoded information. When the decoded result is the same as the query information, it is considered that the data information corresponding to the compressed data is the information to be queried. At this time, the compressed data is completely decompressed. Otherwise, it is considered that the data information corresponding to the compressed data is not the information to be queried. At this time, the compressed data is not decompressed, and the decompression judgment is performed on the encoding result corresponding to the query field in the next compressed data until the information to be queried is obtained and the iteration stops.
[0109] In another embodiment, the method in step S1 is adopted to encode the query information into multiple characters. When decoding the query field of each compressed data, when a decoded string of an encoding result is obtained, it is immediately determined whether the decoded string is consistent with the string at the corresponding position among the multiple characters corresponding to the query information. When they are consistent, the decoding of the next encoding result is performed. Otherwise, it is considered that the data information corresponding to the compressed data is not the content to be queried, and the decoding of the subsequent encoding results is not performed. This can further reduce the amount of decoded data and improve the retrieval efficiency.
[0110] The above are all the preferred embodiments of the present invention. The protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.
Claims
1. A distributed clinical trial data optimization storage method, characterized in that, Including: Converting all fields in a piece of data information of clinical trial data into a character sequence to be compressed; including: Encoding each field in the data information into binary data, dividing the binary data corresponding to each field into groups of every k bits, and converting each group of binary data into a decimal number; regarding a decimal number as a character, then each field corresponds to at least one character, and the characters corresponding to all fields are formed into a character sequence to be compressed in the order of the fields, where k is a preset length; During the process of compressing the character sequence using LZW coding, in response to the character included in the coding object belonging to at least two fields, not coding the coding object, and forming a new coding object with all the characters belonging to the first field among the at least two fields in the coding object; otherwise, coding the coding object; Setting a flag for each field to indicate the end of field coding, and storing the coding results and flags corresponding to all fields hierarchically in order to obtain compressed data, including: Regarding each coding result and flag of a field as a compression element of the field; storing the i-th compression element of each field as the i-th layer in the order of the fields, and when there is no i-th compression element for a certain field, that field does not participate in the storage of the i-th layer; where i represents the layer number; When all the compression elements of all fields are stored, all the layers obtained constitute the compressed data; Each layer of the compressed data contains at most one coding result or flag of each field; In response to the storage system receiving a query message, determining the coding results and the decoding paths of the coding results corresponding to the query field to which the query message belongs according to the flags in each layer of the compressed data, where the decoding path only includes other coding results that must be decoded when decoding the coding result, decoding the coding results corresponding to the query field according to the decoding path to obtain decoded information; in response to the query message being consistent with the decoded information, decoding the compressed data to obtain a query result; The determining of the coding results corresponding to the query field to which the query message belongs includes: Locate the compressed elements corresponding to the query fields at each layer: ; Among them, represents the position serial number of the compressed element of the query field in the i-th layer; represents the position serial number of the encoding result of the query field in the (i - 1)-th layer; represents the number of identifiers that appeared before the th element in the (i - 1)-th layer; h represents the position serial number of the query field in the data information; When the th element of the i-th layer is not an identifier, the th element of the i-th layer is the encoding result of the query field at the i-th layer. At this time, continue to obtain the compressed element of the query field in the next layer; when the th element of the i-th layer is an identifier, and the th element in the i-th layer is the identifier of the query field. At this time, do not obtain the compressed element of the query field in the next layer.
2. The distributed clinical trial data optimization storage method according to claim 1, wherein, The setting of a flag for each field to indicate the end of field coding includes: Setting a unified first flag for each field to indicate the end of field coding; The coding of the coding object further includes: every time a coding object is coded, updating the dictionary of the LZW coding according to the coding object.
3. The distributed clinical trial data optimized storage method according to claim 2, characterized in that The method for obtaining the decoding path is: Numbering all the coding results in the compressed data in sequence from top to bottom and from left to right, and whenever the first flag is encountered, jumping to the first layer of the compressed data and starting to number from the first unnumbered coding result in the first layer; Take the encoded result to be decoded as the first target encoded result , in response to the j-th target encoded result being greater than C, add the -th encoded result to the path sequence, and take the -th encoded result as the (j + 1)-th target encoded result ; otherwise, add the j-th target encoded result to the path sequence; where C represents the length of the initial dictionary Reversing the order of the path sequence as the decoding path of the coding result to be decoded.
4. The distributed clinical trial data optimized storage method according to claim 1, wherein The setting of a flag for each field to indicate the end of field coding includes: In the process of compressing a character sequence using LZW coding, in response to a character included in the coding object belonging to at least two fields, a second identifier for indicating not to update the dictionary is set for the first field among the at least two fields; after the field compression ends, in response to the field not having the second identifier, a third identifier for indicating to update the dictionary is set for the field; the second identifier and the third identifier can be used to indicate the end of field coding; Encoding the coding object further includes: for each coding object encoded, in response to the string formed by the coding object and its next character in the character sequence existing in the dictionary, not updating the dictionary; otherwise, updating the dictionary according to the coding object.
5. The distributed clinical trial data optimized storage method according to claim 4, wherein, The method for obtaining the decoding path is as follows: the elements in the compressed data are sequentially placed into the data sequence in the order from top to bottom and from left to right. During the placement process, whenever an identifier is placed into the data sequence, jump to the first layer of the compressed data; use the encoding result to be decoded as the first target encoding result ; In response to the j-th target coding result not greater than C, add to the path sequence; otherwise, in response to the non-existence of a second identifier before the -th coding result in the data sequence, add the -th coding result to the path sequence, and use the -th coding result in the data sequence as the (j + 1)-th target coding result ; In response to the presence of a second identifier before the th coding result, add the th coding result in the data sequence to the path sequence, and use the th coding result in the data sequence as the (j + 1)-th target coding result ; C represents the length of the initial dictionary; represents the number of second identifiers before the th coding result; Reverse the order of the path sequence to be the decoding path of the encoded result to be decoded.
6. The distributed clinical trial data optimized storage method according to claim 3 or 5, wherein Decoding the respective encoded results corresponding to the query field to obtain decoding information, including: Taking the decoding path of the encoded result to be decoded as the first path, taking the decoding path of each element in the first path as the second path, representing the first element in the second path with D, and taking the element with the serial number D in the initial dictionary as the decoding element of the second path; concatenating the decoding elements of all the second paths into the decoding string of the encoded result to be decoded; Concatenating the decoding strings corresponding to all the encoded results of the query field together to obtain the decoding result of the query field; converting the decoding result into field information to obtain the decoding information.
Citation Information
Patent Citations
A data management and logical verification method for clinical trials and its medical system
CN116631550B
Character-type communication message compression method adopting inter-frame coding
CN102811114A
Standardized clinical big data center system
CN110875095A