Distributed clinical test data optimization storage method

By converting all fields of clinical trial data into a character sequence to be compressed, and splitting the multi-field encoding object and hierarchically storing the encoding results during the LZW encoding process, the problem that LZW encoding cannot decode the fields separately is solved, and efficient clinical trial data query is achieved.

CN120048409AActive Publication Date: 2025-05-27YIDIXI PHARM TECH (JIAXING) CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510525578.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-27
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

LZW encoding cannot decode only a single field, resulting in high cost of querying clinical trial data and low query efficiency.

Method used

By converting all fields of clinical trial data into a sequence of characters to be compressed, and during the LZW encoding process, the encoding object containing multiple fields is not encoded, but is split into the encoding object of each field, the identification of the end of the field encoding is set, and the encoding results and identification are stored layer by layer to achieve separate decoding of the fields.

Benefits of technology

The field is decoding separately, which reduces the amount of compressed data, speeds up compressed data, reduces the query cost of clinical trial data, and improves query efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048409A_ABST
    Figure CN120048409A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical information processing, in particular to a distributed clinical trial data optimization storage method. The method comprises the following steps: converting all fields in one piece of data information of clinical test data into a character sequence to be compressed; in the process of compressing the character sequence, when characters contained in a coding object belong to at least two fields, the character of the first field in the coding object forms a new coding object, an identifier is set for each field, and coding results of all the fields and the identifiers are stored in a layered manner to obtain compressed data. And determining an encoding result corresponding to the query field and a decoding path of the encoding result according to the identifier in the compressed data, decoding the encoding result according to the decoding path to obtain decoding information, and decoding the compressed data to obtain a query result when the query information is consistent with the decoding information. According to the method, the independent decoding of the fields can be realized, and the query efficiency of clinical test data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical information processing, and particularly to a method for optimizing the storage of distributed clinical trial data. Background Art

[0002] With the rapid development of modern medical research, the scale and complexity of clinical trial data have increased exponentially. Therefore, distributed storage is usually adopted for clinical trial data to meet the growing data requirements of clinical trials.

[0003] However, the amount of data in each data node of distributed storage is still extremely large. To improve the storage efficiency of data, the data information in each data node is usually stored after being compressed. For example, a patent document with the publication number "CN116631550B" and the title "A Method for Data Management and Logical Verification of Clinical Trials and Its Medical System" discloses: constructing an original dictionary for each category according to the distribution of different strings in each category, and obtaining an initial dictionary for each category according to the repetition degree of different strings in the original dictionary of each category; according to the semantic information of the string combinations formed by the strings in the initial dictionary of each category and the suffix characters in the clinical trial data of the category, iteratively judging the number of suffix characters and updating the dictionary, and completing the compression of the clinical trial data of each category according to the updated dictionary.

[0004] The above method is an improvement on the LZW coding. Although it can improve the compression efficiency of clinical trial data, due to the characteristic of the LZW coding that the dictionary is updated while compressing, when querying a certain field in the clinical trial data, it is impossible to decode only a single field, but all the compressed data in the data node needs to be completely decompressed, resulting in a large query cost and low query efficiency for clinical trial data. Summary of the Invention

[0005] To solve the problem that the LZW coding cannot decode only a single field, resulting in a large query cost and low query efficiency for clinical trial data, the present invention provides a method for optimizing the storage of distributed clinical trial data, including: Converting all fields in a piece of data information of clinical trial data into a character sequence to be compressed; during the process of compressing the character sequence using the LZW coding, in response to the characters included in the coding object belonging to at least two fields, not encoding the coding object, and forming a new coding object with all the characters belonging to the first field among the at least two fields in the coding object; otherwise, encoding the coding object; setting an identifier for each field to indicate the end of field encoding, and storing the encoding results and identifiers corresponding to all fields in layers in sequence to obtain compressed data, where each layer of the compressed data contains at most one encoding result or identifier of each field.

[0006] The present invention does not encode the encoded object whose contained characters belong to at least two fields, but splits the encoded object into new encoded objects whose contained characters belong to only one field, avoiding jointly encoding the characters in different fields into one encoding result, which may lead to the inability to decode the fields separately, and providing a basis for subsequent realization of separate field decoding; the present invention stores the encoding results and identifiers corresponding to all fields in layers, ensuring that the encoding results corresponding to the fields can be located during decoding, realizing separate field decoding, reducing the amount of decompressed data, accelerating the decompression speed, reducing the query cost of clinical trial data, and improving the query efficiency of clinical trial data.

[0007] Preferably, the conversion of all fields in a piece of data information of clinical trial data into a character sequence to be compressed includes: encoding each field in the data information into binary data, dividing every k bits of the binary data corresponding to each field into a group, and converting each group of binary data into a decimal number; regarding a decimal number as a character, then each field corresponds to at least one character, and the characters corresponding to all fields are formed into a character sequence to be compressed in the order of the fields, where k is a preset length.

[0008] The present invention converts the data information containing multiple data types into a character sequence containing only one data type, increasing the repeatability of the data, thereby improving the compression efficiency of the data information.

[0009] Preferably, the hierarchical storage of the encoding results and identifiers corresponding to all fields in order to obtain compressed data includes: taking each encoding result and identifier of a field as a compression element of the field; storing the i-th compression element of each field as the i-th layer in the order of the fields, and when the i-th compression element of a certain field does not exist, the field does not participate in the storage of the i-th layer; where i represents the layer number; when all the compression elements of all fields are stored, all the obtained layers constitute the compressed data.

[0010] The number of encoding results corresponding to each field is different. The present invention stores the encoding results and identifiers corresponding to all fields in layers, ensuring that the encoding results corresponding to the fields can be located during decoding, thereby realizing separate field decoding.

[0011] Preferably, it further includes: in response to the storage system receiving a query message, determining the encoding results and the decoding paths of the encoding results corresponding to the query field to which the query message belongs according to the identifiers in each layer of the compressed data, where the decoding path only includes other encoding results that must be decoded when decoding the encoding result, decoding the encoding results corresponding to the query field according to the decoding path to obtain decoded information; in response to the query message being consistent with the decoded information, decoding the compressed data to obtain a query result.

[0012] The decoding path constructed by the present invention only includes other coding results that must be decoded when decoding the coding result, avoiding the need to decode all coding results when decompressing the coding result corresponding to the field, reducing the amount of data to be decompressed, and at the same time accelerating the decompression speed, thereby reducing the query cost of clinical trial data and improving the query efficiency of clinical trial data.

[0013] Preferably, determining each coding result corresponding to the query field to which the query information belongs includes: locating the compressed elements corresponding to the query field in each layer: ; wherein, represents the position serial number of the compressed element of the query field in the i-th layer; represents the position serial number of the coding result of the query field in the (i - 1)-th layer; represents the number of identifiers that appear before the -th element in the (i - 1)-th layer; h represents the position serial number of the query field in the data information; when the -th element in the i-th layer is not an identifier, the -th element in the i-th layer is the coding result of the query field in the i-th layer, and at this time, continue to obtain the compressed element of the query field in the next layer; when the -th element in the i-th layer is an identifier, and the -th element in the i-th layer is the identifier of the query field, at this time, do not obtain the compressed element of the query field in the next layer.

[0014] Preferably, setting a flag for each field to indicate the end of field coding includes: setting a unified first flag for each field to indicate the end of field coding; encoding the coding object further includes: every time a coding object is encoded, update the dictionary of the LZW coding according to the coding object.

[0015] Preferably, the method for obtaining the decoding path is: number all the coding results in the compressed data in sequence from top to bottom and from left to right. Whenever the first flag is encountered, jump to the first layer of the compressed data, and start numbering from the first unnumbered coding result in the first layer; use the coding result to be decoded as the first target coding result , in response to the j-th target coding result being greater than , add the -th coding result to the path sequence, and use the -th coding result as the (j + 1)-th target coding result ; otherwise, add the j-th target coding result to the path sequence; wherein, Represents the length of the initial dictionary; reverse the path sequence and use it as the decoding path of the encoded result to be decoded.

[0016] Preferably, a flag for indicating the end of field encoding is set for each field, including: during the process of compressing a character sequence using LZW encoding, in response to the characters included in the encoding object belonging to at least two fields, setting a second flag for indicating not to update the dictionary for the first field among the at least two fields; after the field compression ends, in response to the non-existence of the second flag in the field, setting a third flag for indicating to update the dictionary for the field; the second flag and the third flag can be used to indicate the end of field encoding; the encoding of the encoding object further includes: for each encoded object, in response to the string formed by the encoded object and its next character in the character sequence existing in the dictionary, not updating the dictionary; otherwise, updating the dictionary according to the encoded object.

[0017] When the present invention encodes an encoding object, it selectively updates the dictionary, avoiding the appearance of duplicate strings in the dictionary, reducing the length of the dictionary, and thus improving the compression efficiency of data information. At the same time, the present invention sets the second flag and the third flag to indicate whether the dictionary has been updated, ensuring that during subsequent decoding, the decoding path corresponding to the encoded result of the field can be accurately obtained.

[0018] Preferably, the method for obtaining the decoding path is: sequentially put the elements in the compressed data into the data sequence in the order from top to bottom and from left to right. During the putting process, whenever a flag is put into the data sequence, jump to the first layer of the compressed data; use the encoded result to be decoded as the first target encoded result ; in response to the j-th target encoded result not being greater than , add to the path sequence; otherwise, in response to the non-existence of the second flag before the -th encoded result in the data sequence, add the -th encoded result to the path sequence, and use the -th encoded result in the data sequence as the (j + 1)-th target encoded result ; in response to the existence of the second flag before the -th encoded result, add the -th encoded result in the data sequence to the path sequence, and use the -th encoded result in the data sequence as the (j + 1)-th target encoded result ; Represents the length of the initial dictionary; Represents the number of second flags before the -th encoded result; reverse the path sequence and use it as the decoding path of the encoded result to be decoded.

[0019] Preferably, decoding each encoding result corresponding to the query field to obtain decoding information includes: taking the decoding path of the encoding result to be decoded as the first path, taking the decoding path of each element in the first path as the second path, representing the first element in the second path by D, and taking the element with the serial number D in the initial dictionary as the decoding element of the second path; concatenating the decoding elements of all the second paths into the decoding string of the encoding result to be decoded; concatenating the decoding strings corresponding to all the encoding results of the query field together to obtain the decoding result of the query field; and converting the decoding result into field information to obtain the decoding information.

[0020] The present invention has the following technical effects: The present invention avoids jointly encoding the characters in different fields into one encoding result, providing a basis for realizing separate decoding of fields; Furthermore, through hierarchical storage, the present invention ensures that the encoding results corresponding to the fields can be located during decoding, thereby realizing separate decoding of fields; Furthermore, by constructing the decoding path, the present invention reduces the amount of decompressed data, speeds up the decompression speed, reduces the query cost of clinical trial data, and improves the query efficiency of clinical trial data. Description of the Drawings

[0021] Figure 1 is the flowchart of the method in the distributed clinical trial data optimization storage method according to the embodiment of the present invention; Figure 2 is the schematic diagram of the order of numbering the encoding results in the compressed data of Table 5; Figure 3 is the schematic diagram of the order of putting the elements in the compressed data of Table 6 into the data sequence. Detailed Embodiments

[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0023] The objective of the present invention is to optimize the storage of clinical trial data and separately decode a certain field in the stored compressed data, thereby improving the retrieval efficiency of clinical trial data. To achieve the objective, the following three points need to be done: First, independently encode the fields, such as step S2 of the present invention; second, hierarchical storage, such as step S3 of the present invention; third, construct the decoding path of the encoding result corresponding to the field, such as step S4 of the present invention.

[0024] Among them, independent encoding of fields is to avoid jointly encoding the characters in different fields into one encoding result during the compression process, which may lead to the inability to decode the fields separately; hierarchical storage is to be able to locate the encoding results corresponding to the fields during decoding; constructing a decoding path is to avoid decoding all the encoding results before the encoding results corresponding to the fields.

[0025] The following details the distributed clinical trial data optimization storage method disclosed in the embodiments of the present invention. Refer to Figure 1 , including steps S1 - S5: S1: Convert each data information of the clinical trial data into a character sequence to be compressed.

[0026] The clinical trial data contains multiple data information. Some data information includes the trial start time, the trial end time, the information of the test product, such as the generic name, specification, source, batch number, expiration date, and storage conditions of the test drug or medical device, etc. Some data information includes the basic information of the subjects, such as gender, age, weight, height, baseline data, disease information, etc. Some data information includes the trial process, such as the usage, dosage, time, etc. of the test drug or medical device. Some data information includes the trial results, such as the changes in each index of the subjects compared with the baseline data, the occurrence time, severity, and detailed description of adverse events, etc.

[0027] The amount of clinical trial data is huge and needs to be compressed and stored. Currently, common compression algorithms include entropy coding such as Huffman coding and arithmetic coding, and dictionary - based compression algorithms such as LZW coding and LZ77 coding. The compression efficiency of these compression algorithms depends on the repeatability of the data. Since a data information contains multiple data types, such as the trial start time and the trial end time being dates, the gender of the subject being a character, the age being an integer, and the weight and height being floating - point numbers. The repeatability of the data in each data information is very low, and it is difficult to achieve good compression efficiency using the above - mentioned compression algorithms. Therefore, the present invention pre - processes the data information to increase the repeatability of the data in the data information, thereby improving the compression efficiency of the data information.

[0028] Specifically, for each data information in the clinical trial data, each field in the data information encoding is encoded into binary data using GB2312 encoding. Each group of k bits of the binary data corresponding to each field is divided into a group, and each group of binary data is converted into a decimal number. Regarding a decimal number as a character, each field corresponds to at least one character. The characters corresponding to all fields are formed into a character sequence to be compressed in the order of the fields. Here, k is a preset length. It should be noted that the encoding length of GB2312 for each letter or number is 8 bits, and the encoding length for each Chinese character is 16 bits. Therefore, the length of the binary data corresponding to each field is a multiple of 8. To ensure that the binary data corresponding to each field can be completely divided, k needs to be selected as a factor of 8 other than 1, such as 2, 4, 8, etc., and the implementer can set it according to the actual implementation situation.

[0029] For example: The data information {Amlodipine Hydrochloride Tablets, 5mg / tablet, a certain pharmaceutical company, 20240730, July 2026, store in a cool and dry place, protected from light} contains the fields of the generic name, specification, source, batch number, expiration date, and storage conditions of the test drug. The result of encoding the field "Amlodipine Hydrochloride Tablets" using GB2312 is 1101000111001110110010111110000110110000101100011100001011001000101101011101100011000110101111011100011010101100. If k is set to 8, the corresponding characters for "Amlodipine Hydrochloride Tablets" are {209, 206, 203, 225, 176, 177, 194, 200, 181, 216, 198, 189, 198, 172}. Similarly, the corresponding characters for "5mg / tablet" are {53, 109, 103, 47, 198, 172}, the corresponding characters for "a certain pharmaceutical company" are {196, 179, 214, 198, 210, 169, 211, 208, 207, 222, 185, 171, 203, 190}, the corresponding characters for "20240730" are {50, 48, 50, 52, 48, 55, 51, 48}, the corresponding characters for "July 2026" are {50, 48, 50, 54, 196, 234, 48, 55, 212, 194}, and the corresponding characters for "store in a cool and dry place, protected from light" are {210, 245, 193, 185, 184, 201, 212, 239, 180, 166, 163, 172, 177, 220, 185, 226, 177, 163, 180, 230}. Then the character sequence to be compressed is {209, 206, 203, 225, 176, 177, 194, 200, 181, 216, 198, 189, 198, 172, 53, 109, 103, 47, 198, 172, 196, 179, 214, 198, 210, 169, 211, 208, 207, 222, 185, 171, 203, 190, 50, 48, 50, 52, 48, 55, 51, 48, 50, 48, 50, 54, 196, 234, 48, 55, 212, 194, 210, 245, 193, 185, 184, 201, 212, 239, 180, 166, 163, 172, 177, 220, 185, 226, 177, 163, 180, 230}.

[0030] It should be noted that the present invention is described only by taking the GB2312 encoding as an example. Implementers can select the encoding method according to the actual implementation situation, such as GBK encoding, UTF-8 encoding, etc.

[0031] In addition, implementers can also use other methods to increase the repeatability of data in the data information. For example, first encode each field in the data information into binary data, and then use Base64 encoding to encode the binary data corresponding to each field into multiple characters, and form a character sequence to be compressed with the characters corresponding to all fields.

[0032] At this point, the data information in the clinical trial data has been converted into a character sequence to be compressed.

[0033] S2: Compress the character sequence to be compressed to obtain multiple encoding results corresponding to each field in the data information.

[0034] It should be noted that the LZW encoding is a dictionary-based compression algorithm, but the dictionary used for compression by the LZW encoding is not stored. During the decompression process, the dictionary is constructed while decompressing. When querying clinical trial data, it is necessary to decode a certain field in the data information. Due to the characteristic that the dictionary is constructed while decompressing, it is necessary to decode the compressed data corresponding to the data information from the beginning, and the efficiency is low. At the same time, the characters corresponding to two or more fields in the data information may be encoded into one encoding result, making it impossible to decode a single field separately. Therefore, the present invention improves the LZW encoding method so that a certain field in the data information can be decoded separately later, improving the query efficiency of clinical trial data.

[0035] In one embodiment, use LZW encoding to compress the character sequence to be compressed. During the compression process, judge the current encoding object: If the characters included in the encoding object belong to two or more fields, the encoding object is not encoded, and all the characters of the first field among the two or more fields to which the encoding object belongs are formed into a new encoding object; if the characters included in the encoding object belong to only one field, use the serial number corresponding to the same character or string as the encoding object in the dictionary to encode the encoding object to obtain an encoding result. At this time, the string formed by the encoding object and its next character in the character sequence is added to the end of the dictionary to update the dictionary.

[0036] For example, when the data information is {Amlodipine Hydrochloride Tablets, 5mg / tablet, a certain pharmaceutical company, 20240730, July 2026, store in a cool and dry place, protected from light}, the corresponding character sequence to be compressed is {209, 206, 203, 225, 176, 177, 194, 200, 181, 216, 198, 189, 198, 172, 53, 109, 103, 47, 198, 172, 196, 179, 214, 198, 210, 169, 211, 208, 207, 222, 185, 171, 203, 190, 50, 48, 50, 52, 48, 55, 51, 48, 50, 50, 54, 196, 234, 48, 55, 212, 194, 210, 245, 193, 185, 184, 201, 212, 239, 180, 166, 163, 172, 177, 220, 185, 226, 177, 163, 180, 230}. When the coding object is the 42nd to 43rd characters in the character sequence to be compressed {48, 50}, where {48} belongs to the character corresponding to the batch number field "20240730" and {50} belongs to the character corresponding to the expiration date field "July 2026", then the coding object {48, 50} is not encoded, and {48} is taken as the new coding object. The new coding object {48} only belongs to the character corresponding to the batch number field "20240730". At this time, the new coding object {48} is encoded, and the string {48, 50} formed by {48} and its next character {50} in the character sequence to be compressed is added to the end of the dictionary to update the dictionary.

[0037] Set a first identifier for each field to indicate the end of encoding for that field. To distinguish the encoding result from the first identifier, the set first identifier cannot be the same as the sequence numbers of all elements in the dictionary. The first identifier can be set to 0, or a number greater than L, where L represents the length of the finally obtained dictionary.

[0038] For example, when the data information is {Amlodipine Hydrochloride Tablets, 5mg / tablet, a certain pharmaceutical company, 20240730, July 2026, store in a cool and dry place, protected from light}, the corresponding character sequence to be compressed is {209, 206, 203, 225, 176, 177, 194, 200, 181, 216, 198, 189, 198, 172, 53, 109, 103, 47, 198, 172, 196, 179, 214, 198, 210, 169, 211, 208, 207, 222, 185, 171, 203, 190, 50, 48, 50, 52, 48, 55, 51, 48, 50, 50, 54, 196, 234, 48, 55, 212, 194, 210, 245, 193, 185, 184, 201, 212, 239, 180, 166, 163, 172, 177, 220, 185, 226, 177, 163, 180, 230}. The first identifier is set to 0. By using the method in this embodiment to compress the character sequence to be compressed, the obtained dictionary is shown in Table 1, and the encoding results corresponding to each field and the first identifier are shown in Table 2.

[0039] Table 1

[0040] Table 2

[0041] It should be noted that inside a computer, data is stored in binary form. Therefore, after the LZW encoding ends, the compression result needs to be converted to binary for storage. In LZW encoding, when the length of the dictionary increases, it may cause the number of bits for converting each encoding result to binary to increase, thus affecting the compression efficiency. For example, when the length of the dictionary is less than 64, each encoding result is less than 64, and each encoding result can be represented by 6 - bit binary. When the length of the dictionary is greater than 64 and less than 128, each encoding result may be greater than 64, and 7 - bit binary is required to represent each encoding result. In the above embodiment, the string formed by the new encoding object and its next character in the character sequence already exists in the dictionary. After encoding the new encoding object, the string formed by the new encoding object and its next character in the character sequence is added to the end of the dictionary, resulting in the existence of the same elements in the dictionary, increasing the length of the dictionary and thus affecting the compression efficiency. Therefore, in another embodiment of the present invention, during the process of compressing the character sequence to be compressed using LZW encoding, the string formed by the encoding object and its next character in the character sequence is selectively added to the dictionary, making the length of the dictionary as short as possible, thereby improving the compression efficiency.

[0042] In another embodiment, the character sequence to be compressed is compressed using LZW coding. During the compression process, the current coding object is judged: If the characters included in the coding object belong to two or more fields, the coding object is not coded. All the characters of the first field among the two or more fields included in the coding object form a new coding object. At this time, the new coding object is coded using the serial number corresponding to the same character or string as the new coding object in the dictionary, and the coding result is obtained. At this time, the dictionary is not updated, and a second identifier is set for the field corresponding to the new coding object. The second identifier is used to indicate that the dictionary has not been updated.

[0043] If the characters included in the coding object belong to only one field, the coding object is coded using the serial number corresponding to the same character or string as the coding object in the dictionary. At this time, the dictionary is updated, and the update method is the same as the method in LZW coding, that is, the string formed by the coding object and its next character in the character sequence is added to the end of the dictionary to implement the update of the dictionary.

[0044] After all the characters corresponding to a certain field are compressed, if the second identifier does not exist for this field, a third identifier is set for this field. The third identifier is used to indicate that the dictionary has been updated during the compression of the characters included in this field.

[0045] The second identifier and the third identifier can both indicate the end of the coding of the field. In order to distinguish the coding result, the second identifier and the third identifier, the set second identifier and third identifier cannot be the same, and cannot be the same as the serial numbers of all elements in the dictionary. The second identifier and the third identifier can be set to 0, or a number greater than L, where L represents the length of the finally obtained dictionary.

[0046] For example, when the data information is {Amlodipine Hydrochloride Tablets, 5mg / tablet, a certain pharmaceutical company, 20240730, July 2026, store in a cool and dry place, protected from light}, the corresponding character sequence to be compressed is {209, 206, 203, 225, 176, 177, 194, 200, 181, 216, 198, 189, 198, 172, 53, 109, 103, 47, 198, 172, 196, 179, 214, 198, 210, 169, 211, 208, 207, 222, 185, 171, 203, 190, 50, 48, 50, 52, 48, 55, 51, 48, 50, 48, 50, 54, 196, 234, 48, 55, 212, 194, 210, 245, 193, 185, 184, 201, 212, 239, 180, 166, 163, 172, 177, 220, 185, 226, 177, 163, 180, 230}. The second identifier is set to L + 1, and the third identifier is set to 0, where L represents the length of the finally obtained dictionary. During the process of compressing the character sequence to be compressed by the method in this embodiment, the constructed dictionary is shown in Table 3, and the corresponding encoding results and identifiers for each field are shown in Table 4.

[0047] Table 3

[0048] Table 4

[0049] Thus, the compression of the character sequence to be compressed is achieved, and multiple encoding results and identifiers corresponding to each field in the data information are obtained.

[0050] S3: Hierarchically store the multiple encoding results and identifiers corresponding to each field in the data information to obtain compressed data.

[0051] It should be noted that in traditional compression methods, the obtained encoding results are stored in sequence to form a sequence. Since the number of encoding results corresponding to each field is different, for example, the encoding results corresponding to the general name field "Amlodipine Hydrochloride Tablets" are 14, and the encoding objects corresponding to the specification field "5mg / tablet" are 5. When decoding, the position of the encoding result corresponding to each field cannot be known in the compressed data, resulting in the inability to decode a single field separately. Therefore, the present invention hierarchically stores the encoding results and identifiers of all fields, so that when decoding, the position of the encoding result corresponding to each field can be accurately located in the compressed data, thereby enabling each field to be decoded separately, avoiding decoding the entire compressed data when querying clinical trial data, and thus improving the query efficiency.

[0052] Hierarchical storage means storing the encoding results in each layer in the order of the encoding results of the fields; that is, the first encoding result of each field is placed in the first layer, the second encoding result of each field is placed in the second layer, the third encoding result of each field is placed in the third layer, and so on. The following is an example for easy understanding.

[0053] Specifically, after the compression of the character sequence to be compressed is completed, each field corresponds to at least one encoding result and an identifier, and all the encoding results and identifiers of the field are used as a compression element of the field respectively. The i-th compression element of each field is stored as the i-th layer information in the order of the fields. When the i-th compression element of a certain field does not exist, the field does not participate in the construction of the i-th layer information.

[0054] When all the compression elements of all fields are stored, all the layer information obtained constitutes the compressed data of the data information. It should be noted that the identifier can indicate the end of the field encoding. In the compressed data, the layer where the identifier is located can also reflect the number of encoding results corresponding to a field. For example, if the identifier of a certain field is located in the 3rd layer of the compressed data, it means that the field corresponds to two encoding results.

[0055] For example: Hierarchical storage is performed on the encoding results and identifiers corresponding to each field in Table 2, and the obtained compressed data is shown in Table 5. Each row in Table 5 represents one layer of information.

[0056] Table 5

[0057] Hierarchical storage is performed on the encoding results and identifiers corresponding to each field in Table 4, and the obtained compressed data is shown in Table 6. Each row in Table 6 represents one layer of information.

[0058] Table 6

[0059] In the present invention, the compressed data of all data information is stored distributively, so that the compressed data is stored on multiple data nodes, thereby ensuring the reliability and availability of the data. It should be noted that the hierarchical storage in the present invention means hierarchically storing multiple encoding results and identifiers corresponding to each field in a data information on a single data node.

[0060] S4: Determine each encoding result corresponding to the query field to which the query information belongs according to the identifiers in each layer of the compressed data.

[0061] When a staff member needs to query information about a certain clinical trial data, the query information is input on the query system, and the field to which the query information belongs is called the query field. The query system synchronizes the query information to the data node containing the query field, and the data node only decodes the query field in each compressed data. When decoding, it is first necessary to locate the encoding results corresponding to the query field in each layer of the compressed data.

[0062] It should be noted that the position of the encoding result corresponding to the query field in each layer of the compressed data is determined by the rule of hierarchical storage. That is to say, the implementer can infer the position of the encoding result corresponding to the query field in each layer of the compressed data according to the rule of hierarchical storage. The following gives a reasoning method.

[0063] Specifically, for any piece of compressed data, obtain the encoding results corresponding to the query field in the information of each layer of the compressed data: ; Among them, represents the position serial number of the compressed element of the query field in the i-th layer, that is, in the information of the i-th layer, the compressed element corresponding to the query field is the th element in the information of the i-th layer; represents the position serial number of the encoding result of the query field in the (i - 1)-th layer, that is, in the information of the (i - 1)-th layer, the encoding result corresponding to the query field is the th element in the information of the (i - 1)-th layer; represents the number of identifiers that appear before the th element in the (i - 1)-th layer; represents the position serial number of the query field in the data information, that is, the th field in the data information is the query field.

[0064] When the th element in the i-th layer is not an identifier, the th element in the i-th layer is the encoding result of the query field in the i-th layer. At this time, continue to obtain the compressed element of the query field in the next layer; when the th element in the i-th layer is an identifier, the th element in the i-th layer is the identifier of the query field. At this time, do not obtain the compressed element of the query field in the next layer.

[0065] It should be noted that when the first identifier is used in step S2, the identifier in this step refers to the first identifier. When the second identifier and the third identifier are used in step S2, the identifier in this step includes the second identifier and the third identifier.

[0066] So far, the encoding results corresponding to the query field have been obtained.

[0067] S5: Obtain the decoding paths of the respective encoding results corresponding to the query fields, decode the encoding results according to the decoding paths to obtain decoding information, and obtain query results according to the decoding information.

[0068] It should be noted that the numerical value of the encoding result represents the serial number of its corresponding string in the dictionary. Since the dictionary has not been constructed, it cannot be directly decoded. In LZW encoding, the serial number of a string in the dictionary represents the order in which the string is added to the dictionary. The previous encoding object when the string is added to the dictionary is the string composed of the first S - 1 characters of the string, where S represents the length of the string, and the first character of the next encoding object is the last character of the string. The string corresponding to the encoding result can be decoded according to the previous encoding object and the next encoding object when the string corresponding to the encoding result is added to the dictionary. According to the serial number of the string in the dictionary and the second identifier, the encoding result corresponding to the previous encoding object when the string is added to the dictionary and the encoding result corresponding to the next encoding object can be obtained. Therefore, for the respective encoding results corresponding to the query fields, the present invention decodes the encoding results by constructing the decoding paths of the encoding results, and the decoding paths include other encoding results that must be decoded for decoding the encoding results.

[0069] In one embodiment, when step S2 uses the first identifier, the method for obtaining the decoding path is as follows: Number all the encoding results in the compressed data in sequence from top to bottom and from left to right. Whenever a identifier is encountered, jump to the first layer of the compressed data, and start from the first unnumbered encoding result in the first layer, and continue to number the unnumbered encoding results in the compressed data in sequence from top to bottom and from left to right until all the encoding results have been numbered. Stop when all the encoding results have been numbered. The number of the encoding result represents the encoding order of the encoding object corresponding to the encoding result during the compression of the character sequence to be encoded. For example, for the schematic diagram of the numbering order of the encoding results in the compressed data in Table 5, see Figure 2 , Figure 2 The gray ones in it are the first identifiers, and the first identifiers do not need to be numbered.

[0070] Construct an empty path sequence, denote the currently to-be-decoded encoding result as the first target encoding result, and represent the first target encoding result with . When is less than or equal to , add to the path sequence, where represents the length of the initial dictionary; when is greater than , add the encoding result with the number to the path sequence, and add the encoding result with the number The encoding result is used as the second target encoding result, and is used to represent it; When the second target encoding result is less than or equal to , is added to the path sequence; when the second target encoding result is greater than , the encoding result with the number is added to the path sequence, and the encoding result with the number is used as the third target encoding result, and is used to represent it; When the third target encoding result is less than or equal to , is added to the path sequence; when the third target encoding result is greater than , the encoding result with the number is added to the path sequence, and the encoding result with the number is used as the fourth target encoding result, and is used to represent it; And so on, until there is a target encoding result added to the path sequence, the iteration stops.

[0071] The finally obtained path sequence is reversed, and the reversed result is used as the decoding path of the currently to-be-decoded encoding result.

[0072] For example: The sequence formed by all the encoding objects in the compressed data of Table 5 in ascending order of the numbers is {35, 32, 31, 43, 16, 17, 26, 29, 20, 40, 28, 23, 28, 15, 6, 10, 9, 1, 61, 27, 18, 39, 28, 36, 13, 37, 34, 33, 42, 22, 14, 31, 24, 3, 2, 3, 5, 2, 8, 4, 2, 82, 3, 7, 27, 46, 86, 38, 26, 36, 48, 25, 22, 21, 30, 38, 47, 19, 12, 11, 15, 17, 41, 22, 44, 17, 11, 19, 45}. The initial dictionary is shown in Table 7, and the length of the initial dictionary is 48. The numbers of the encoding objects corresponding to the validity period field are from 42 to 49. Among them, the encoding object with the number 42 is {82}. The encoding object {2} with the number 82 - 48 + 1 = 35 is added to the path sequence. Since the encoding object {3} with the number 82 - 48 = 34 is less than 48, the encoding object {3} with the number 82 - 48 = 34 is added to the path sequence. Then the final path sequence is {2, 3}, and the decoding path is {3, 2}.

[0073] Table 7

[0074] In another embodiment, when the second identifier and the third identifier are adopted in step S2, the method for obtaining the decoding path is as follows: Construct an empty data sequence, and put all the elements in the compressed data into the data sequence in the order from top to bottom and from left to right. Whenever an identifier is encountered, after putting the identifier into the data sequence, jump to the first layer of the compressed data, and start from the first encoded result that has not been put into the data sequence in the first layer, and continue to put the elements in the compressed data that have not been put into the data sequence into the data sequence in the order from top to bottom and from left to right. Stop until all the elements in the compressed data have been put into the data sequence. For example, for the schematic diagram of the order of putting the elements in the compressed data in Table 6, see Figure 3 , Figure 3 The gray ones in are the second identifier and the third identifier, and the second identifier and the third identifier also need to be put into the data sequence. The finally obtained data sequence is {35, 32, 31, 43, 16, 17, 26, 29, 20, 40, 28, 23, 28, 15, 0, 6, 10, 9, 1, 61, 0, 27, 18, 39, 28, 36, 13, 37, 34, 33, 42, 22, 14, 31, 24, 0, 3, 2, 3, 5, 2, 8, 4, 2, 116, 82, 3, 7, 27, 46, 86, 38, 26, 0, 36, 48, 25, 22, 21, 30, 38, 47, 19, 12, 11, 15, 17, 41, 22, 44, 17, 11, 19, 45, 0}.

[0075] Construct an empty path sequence, denote the currently decoded encoded result as the first target encoded result, and represent the first target encoded result with . When is less than or equal to , add to the path sequence, where represents the length of the initial dictionary; when is greater than , obtain the th encoded result in the data sequence. If there is no second identifier before the th encoded result in the data sequence, add the th encoded result in the data sequence to the path sequence, and use the th encoded result in the data sequence as the second target encoded result, and represent it with ; if there is a second identifier before the th encoded result in the data sequence, add the th encoded result in the data sequence to the path sequence, and use the The encoding result is used as the second target encoding result and is represented by , where represents the number of second identifiers before the th encoding result in the data sequence; When the second target encoding result is less than or equal to , is added to the path sequence; when the second target encoding result is greater than , if there is no second identifier before the th encoding result in the data sequence, the th encoding result in the data sequence is added to the path sequence, and the th encoding result in the data sequence is used as the third target encoding result and is represented by ; if there is a second identifier before the th encoding result in the data sequence, the th encoding result in the data sequence is added to the path sequence, and the th encoding result in the data sequence is used as the third target encoding result and is represented by , where represents the number of second identifiers before the th encoding result in the data sequence; When the third target encoding result is less than or equal to , is added to the path sequence; when the third target encoding result is greater than , if there is no second identifier before the th encoding result in the data sequence, the th encoding result in the data sequence is added to the path sequence, and the th encoding result in the data sequence is used as the fourth target encoding result and is represented by ; if there is a second identifier before the th encoding result in the data sequence, the th encoding result in the data sequence is added to the path sequence, and the th encoding result in the data sequence is used as the fourth target encoding result and is represented by , where represents the number of second identifiers before the th encoding result in the data sequence; And so on, until a target encoding result is added to the path sequence and the iteration stops.

[0076] Reverse the obtained path sequence, and use the result of the reverse arrangement as the decoding path of the currently to-be-decoded encoded result.

[0077] For example: The data sequence formed by all elements in the compressed data in Table 6 is {35, 32, 31, 43, 16, 17, 26, 29, 20, 40, 28, 23, 28, 15, 0, 6, 10, 9, 1, 61, 0, 27, 18, 39, 28, 36, 13, 37, 34, 33, 42, 22, 14, 31, 24, 0, 3, 2, 3, 5, 2, 8, 4, 2, 116, 82, 3, 7, 27, 46, 86, 38, 26, 0, 36, 48, 25, 22, 21, 30, 38, 47, 19, 12, 11, 15, 17, 41, 22, 44, 17, 11, 19, 45, 0}. The initial dictionary is shown in Table 7. The length of the initial dictionary is 48. The valid period field corresponds to the 42nd to 49th encoded objects in the data sequence. Among them, the second identifier and the third identifier in the data sequence do not participate in the counting. The 47th encoded object is {86}. Since the second identifier 116 does not exist before the 39th encoded object {8} in the data sequence (86 - 48 + 1 = 39), the 39th encoded object {8} in the data sequence is added to the path sequence. For the 38th encoded object {2} in the data sequence (86 - 48 = 38), since the 38th encoded object {2} in the data sequence is less than 48, the 38th encoded object {2} in the data sequence is added to the path sequence. Then the final path sequence is {8, 2}, and the decoding path is {2, 8}.

[0078] So far, the decoding path has been obtained.

[0079] Decode the encoded result according to the decoding path. Specifically: Use the decoding path of the to-be-decoded encoded result as the first path, use the decoding path of each element in the first path as the second path, represent the first element in the second path with D, and use the element with the serial number D in the initial dictionary as the decoding element of the second path. Concatenate the decoding elements of all the second paths into the decoding string of the to-be-decoded encoded result.

[0080] For example: In the compressed data in Table 5, the decoding paths of each element in the decoding path {3, 2} of the encoded object {82} with the number 42 are {3} and {2} respectively. The initial dictionary is shown in Table 7. Then the elements with the serial numbers 3 and 2 in the initial dictionary are {50} and {48} respectively. Then the decoding string is {50, 48}.

[0081] In the compressed data of Table 6, for the decoding path {2, 8} of the 47th encoded object {86}, the decoding paths of each element are {2} and {8} respectively. Referring to Table 7 for the initial dictionary, the elements with serial numbers 2 and 8 in the initial dictionary are {48} and {55} respectively, so the decoded string is {48, 55}.

[0082] In one embodiment, the decoded strings corresponding to all the encoding results of the query field are concatenated together to obtain the decoding result of the query field. Each character in the decoding result is actually a decimal number. Each decimal number in the decoding result is separately converted into a binary number, and all the obtained binary numbers are concatenated together, and the concatenated binary number is decoded using GB2312 encoding to obtain the decoded information. When the decoding result is the same as the query information, it is considered that the data information corresponding to the compressed data is the information to be queried. At this time, the compressed data is completely decompressed. Otherwise, it is considered that the data information corresponding to the compressed data is not the information to be queried. At this time, the compressed data is not decompressed, and the decoding judgment is performed on the encoding result corresponding to the query field in the next compressed data until the information to be queried is obtained and the iteration stops.

[0083] In another embodiment, using the method in step S1, the query information is encoded into multiple characters. When decoding the query field of each compressed data, when the decoded string of an encoding result is obtained, it is immediately judged whether the string at the corresponding position in the decoded string is consistent with the strings at the corresponding positions in the multiple characters corresponding to the query information. When they are consistent, the decoding of the next encoding result is performed. Otherwise, it is considered that the data information corresponding to the compressed data is not the content to be queried, and the decoding of the subsequent encoding results is not performed. This can further reduce the amount of data to be decoded and improve the retrieval efficiency.

[0084] The above are all the preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A distributed clinical trial data optimization storage method, characterized in that: include: Convert all fields in a piece of data information of clinical trial data into a character sequence to be compressed; In the process of compressing a character sequence using LZW encoding, in response to characters included in the encoding object belonging to at least two fields, the encoding object is not encoded, and all characters in the encoding object belonging to a first field of the at least two fields form a new encoding object; otherwise, the encoding object is encoded; An identifier is set for each field to indicate the end of field encoding, and the encoding results and identifiers corresponding to all fields are stored in layers in order to obtain compressed data, wherein each layer of the compressed data contains at most one encoding result or identifier of each field.

2. The distributed clinical trial data optimization storage method according to claim 1, characterized in that: The step of converting all fields in a piece of data information of the clinical trial data into a character sequence to be compressed includes: Encode each field in the data information into binary data, divide the binary data corresponding to each field into a group of k bits, and convert each group of binary data into a decimal number; regard a decimal number as a character, then each field corresponds to at least one character, and the characters corresponding to all fields are arranged in the order of the fields to form a character sequence to be compressed, where k is a preset length.

3. The distributed clinical trial data optimization storage method according to claim 1, characterized in that: The encoding results and identifiers corresponding to all fields are stored in layers in order to obtain compressed data, including: Each encoding result and identifier of the field is used as a compression element of the field; the i-th compression element of each field is stored as the i-th layer according to the order of the fields. When a field does not have the i-th compression element, the field does not participate in the storage of the i-th layer; i represents the layer number; When all compressed elements of all fields are stored, all the resulting layers constitute the compressed data.

4. The distributed clinical trial data optimization storage method according to claim 1, characterized in that: Also includes: In response to the storage system receiving the query information, the encoding results corresponding to the query field to which the query information belongs and the decoding path of each encoding result are determined according to the identifiers in each layer of the compressed data, wherein the decoding path only includes other encoding results that must be decoded when decoding the encoding result, and the encoding results corresponding to the query field are decoded according to the decoding path to obtain decoding information; in response to the query information being consistent with the decoding information, the compressed data is decoded to obtain the query result.

5. The distributed clinical trial data optimization storage method according to claim 4, characterized in that: The determining of each encoding result corresponding to the query field to which the query information belongs includes: Locate the compressed elements corresponding to the query field at each layer: ; in, Indicates the position number of the compressed element of the query field in the i-th layer; Indicates the position number of the encoding result of the query field at the i-1 layer; Indicates the i-1th layer The number of identifiers that appear before the element; h represents the position number of the query field in the data information; When the i-th layer When the element is not a label, the i-th layer The first element is the encoding result of the query field at the i-th layer, and then continue to obtain the compressed elements of the query field at the next layer; when the first element of the i-th layer The element is the identifier, and the The element is the identifier of the query field. At this time, the compressed element of the query field in the next layer is not obtained.

6. The distributed clinical trial data optimization storage method according to claim 4, characterized in that: The step of setting a flag for each field to indicate the end of field encoding includes: A unified first identifier is set for each field to indicate the end of field encoding; The encoding of the encoding object further includes: each time an encoding object is encoded, updating the LZW encoding dictionary according to the encoding object.

7. The distributed clinical trial data optimization storage method according to claim 6, characterized in that: The method for obtaining the decoding path is: Number all the encoding results in the compressed data in order from top to bottom and from left to right. Whenever the first identifier is encountered, jump to the first layer of the compressed data and continue numbering from the first unnumbered encoding result in the first layer; The encoding result to be decoded is used as the first target encoding result , in response to the j-th target encoding result is greater than C, The encoding results are added to the path sequence. The encoding result is used as the j+1th target encoding result ; Otherwise, encode the jth target Add the path sequence; where C represents the length of the initial dictionary; Arrange the path sequence in reverse order as the decoding path of the encoding result to be decoded.

8. The distributed clinical trial data optimization storage method according to claim 4, characterized in that: The step of setting a flag for each field to indicate the end of field encoding includes: In the process of compressing a character sequence using LZW encoding, in response to characters included in the encoding object belonging to at least two fields, setting a second flag for indicating that a dictionary is not updated for a first field of the at least two fields; after the field compression is completed, in response to the field not having the second flag, setting a third flag for indicating that a dictionary is updated for the field; the second flag and the third flag can be used to indicate the end of field encoding; The encoding of the encoding object further includes: each time an encoding object is encoded, in response to a character string consisting of the encoding object and its next character in the character sequence already existing in the dictionary, not updating the dictionary; otherwise, updating the dictionary according to the encoding object.

9. The distributed clinical trial data optimization storage method according to claim 8, characterized in that: The method for obtaining the decoding path is as follows: the elements in the compressed data are placed in the data sequence in order from top to bottom and from left to right. During the placing process, each time the identifier is placed in the data sequence, jump to the first layer of the compressed data; the encoding result to be decoded is used as the first target encoding result. ; In response to the jth target encoding result Not greater than C, Add to the path sequence; otherwise, in response to the first There is no second identifier before the encoding result. The encoding results are added to the path sequence, and the The encoding result is taken as the j+1th target encoding result In response to the There is a second identifier before the encoding result, The encoding results are added to the path sequence, and the first The encoding result is taken as the j+1th target encoding result ; C represents the length of the initial dictionary; Indicates the The number of second identifiers before the encoding result; Arrange the path sequence in reverse order as the decoding path of the encoding result to be decoded.

10. The distributed clinical trial data optimization storage method according to claim 7 or 9, characterized in that: The decoding of each encoding result corresponding to the query field to obtain decoding information includes: The decoding path of the encoding result to be decoded is taken as the first path, the decoding path of each element in the first path is taken as the second path, the first element in the second path is represented by D, and the element with the sequence number D in the initial dictionary is taken as the decoding element of the second path; all the decoding elements of the second path are concatenated into a decoding string of the encoding result to be decoded; The decoded strings corresponding to all the encoding results of the query field are concatenated together to obtain the decoding result of the query field; the decoding result is converted into field information to obtain decoding information.

Citation Information

Patent Citations

  • A data management and logical verification method for clinical trials and its medical system

    CN116631550B

  • Character-type communication message compression method adopting inter-frame coding

    CN102811114A

  • Standardized clinical big data center system

    CN110875095A

  • Character string storage method and device, and electronic equipment

    CN112232025A

  • Encoding method and device, decoding method and device and computer readable storage medium

    CN112711935A