Data processing method, device, computer equipment and storage medium
By obtaining the distribution range of medical data in the original data set and determining the query method, and looking up the statistical results of the data to be queried in the data repository, the problem of low performance in large-scale medical data search in the existing technology is solved, and the overall data processing performance and human-computer interaction efficiency of the system are improved.
Patent Information
- Application Number
- CN202111591994.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-12-23
AI Technical Summary
The existing technology has poor performance when searching for large-scale medical data, resulting in reduced overall data processing performance of the system and inefficient human-computer interaction.
By responding to the statistical result query instruction of the data to be queried, the distribution range to which the data to be queried belongs in the original data set, and the query method is determined based on the distribution range, and the appropriate query method is used to find the statistical results of the data to be queried in the data store.
It improves the performance when searching data, improves the overall data processing performance of the system, and thus improves the human-computer interaction efficiency.
Smart Images

Figure CN114356983B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular, to a data processing method, apparatus, computer device, and storage medium. Background Art
[0002] With the development of information technology, the information data generated by major enterprises is increasing, the storage volume of data is increasing day by day, and the query requirements of users for data in the database are becoming more and more complex.
[0003] Taking medical data as an example, a large amount of medical data is collected in many medical scenarios. In this case, when searching for data, the related technology usually performs a one-by-one search by traversing and matching. However, due to the large data scale, the performance of data search in the related technology is low, which reduces the overall data processing performance of the system, resulting in low human-computer interaction efficiency. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a data processing method, apparatus, computer device, and storage medium, which can improve the performance when searching for data, improve the overall data processing performance of the system, and thus improve the human-computer interaction efficiency.
[0005] In a first aspect, this application provides a data processing method, which includes:
[0006] In response to a statistical result query instruction for the data to be queried, obtain the distribution range of the data to be queried in the original data set;
[0007] Determine the query method according to the distribution range of the data to be queried in the original data set;
[0008] Search for the statistical result of the data to be queried in the data repository through the query method.
[0009] In one embodiment, obtaining the distribution range of the data to be queried in the original data set includes:
[0010] According to the location information of the data to be queried in the original data set, determine the distribution range of the data to be queried in the original data set from the respective distribution ranges in the original data set;
[0011] Wherein, the respective distribution ranges in the original data set are obtained by dividing the original data set according to the distribution characteristic types of the data in the original data set.
[0012] In one embodiment, determining the query method according to the distribution range of the data to be queried in the original data set includes:
[0013] Obtain the distribution characteristic type of the data in the obtained distribution range;
[0014] Determine the storage method of the statistical results of the data in the distribution range to which it belongs in the data repository according to the distribution characteristic type;
[0015] Determine the query method corresponding to the storage method.
[0016] In one embodiment, determining the storage method of the statistical results of the data in the distribution range to which it belongs in the data repository according to the distribution characteristic type includes:
[0017] If the distribution characteristic type is dense distribution, the storage method of the statistical results of the data in the distribution range to which it belongs in the data repository is the array storage method;
[0018] If the distribution characteristic type is sparse distribution, the storage method of the statistical results of the data in the distribution range to which it belongs in the data repository is the linked list storage method.
[0019] In one embodiment, determining the query method corresponding to the storage method includes:
[0020] If the storage method is the array storage method, determine that the query method corresponding to the storage method is the array query method;
[0021] If the storage method is the linked list storage method, determine that the query method corresponding to the storage method is the linked list query method.
[0022] In one embodiment, finding the statistical results of the data to be queried in the data repository through the query method includes:
[0023] If the query method is the array query method, determine the first identifier corresponding to the data to be queried; wherein, the first identifier is used to identify the element of the array in the data repository that stores the statistical results of the data to be queried;
[0024] According to the first identifier and the starting address of the array, determine the address of the element corresponding to the first identifier in the data repository;
[0025] According to the content stored in the address, determine the statistical results of the data to be queried.
[0026] In one embodiment, finding the statistical results of the data to be queried in the data repository through the query method includes:
[0027] If the query method is the linked list query method, determine the second identifier corresponding to the data to be queried; wherein, the second identifier is used to identify the node of the linked list in the data repository that stores the statistical results of the data to be queried;
[0028] Starting from the head node of the linked list, compare the second identifier corresponding to the data to be queried with the identifiers of each node in the linked list in the order of the linked list connections, until a node matching the second identifier is found;
[0029] Determine the statistical result of the data to be queried according to the content stored in the node.
[0030] In a second aspect, an embodiment of the present application provides a data processing device, which includes:
[0031] An acquisition module, configured to acquire the distribution range of the data to be queried in the original data set in response to a statistical result query instruction of the data to be queried;
[0032] A determination module, configured to determine a query method according to the distribution range of the data to be queried in the original data set;
[0033] A search module, configured to search for the statistical result of the data to be queried in the data repository by the query method.
[0034] In a third aspect, an embodiment of the present application provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of any method provided in the first aspect embodiment are implemented.
[0035] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any method provided in the first aspect embodiment are implemented.
[0036] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any method provided in the first aspect embodiment are implemented.
[0037] A data processing method, device, computer device, and storage medium provided by an embodiment of the present application, in response to a statistical result query instruction of the data to be queried, acquire the distribution range of the data to be queried in the original data set, determine a query method according to the distribution range of the data to be queried in the original data set, and then search for the statistical result of the data to be queried in the data repository by the query method. After the computer device receives the query instruction, it will first determine the distribution range to which the data to be queried belongs, and based on this belonging distribution range, a query method will be determined. In this way, performing the query operation in a query method suitable for its belonging distribution range can improve the performance when searching for data, improve the overall data processing performance of the system, and thus improve the human-computer interaction efficiency. Description of the Drawings
[0038] Figure 1A block diagram of the internal structure of a computer device provided for an embodiment;
[0039] Figure 2 A schematic flowchart of a data processing method provided for an embodiment;
[0040] Figure 3 A schematic diagram of the data distribution in an original dataset provided for an embodiment;
[0041] Figure 4 A schematic flowchart of a data processing method provided for another embodiment;
[0042] Figure 5 A schematic flowchart of a data processing method provided for another embodiment;
[0043] Figure 6 A schematic diagram of an array method provided for an embodiment;
[0044] Figure 7 A schematic flowchart of a data processing method provided for another embodiment;
[0045] Figure 8 A schematic diagram of a linked list method provided for an embodiment;
[0046] Figure 9 A schematic diagram of the distribution of heartbeat data provided for an embodiment;
[0047] Figure 10 A schematic flowchart of a data processing method provided for another embodiment;
[0048] Figure 11 A schematic diagram of the structure of a data processing device provided for an embodiment. Detailed implementation manners
[0049] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0050] The data processing method provided by the embodiments of the present application can be applied to any type of computer device, including but not limited to various personal computers, laptop computers, smart terminals, tablet computers, wearable devices, etc. The embodiments of the present application do not limit the type of computer device.
[0051] Such as Figure 1As shown in the figure, a schematic diagram of the internal structure of a computer device is provided. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device is used to store data during the data processing process. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a data processing method.
[0052] Embodiments of the present application provide a data processing method, apparatus, computer device, and storage medium, which can improve the performance when searching for data, improve the overall data processing performance of the system, and thus improve the human-computer interaction efficiency. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application.
[0053] In one embodiment, a data processing method is provided for application to Figure 1 Taking the application environment shown in the figure as an example, this embodiment involves the specific process of determining the distribution range of the data to be queried in the original dataset according to the statistical result query instruction of the data to be queried, determining the query method based on this distribution range, and using this query method to search for the statistical result of the data to be queried in the data repository, as Figure 2 shown in the figure, this embodiment includes the following method steps:
[0054] S101, in response to the statistical result query instruction of the data to be queried, obtain the distribution range of the data to be queried in the original dataset.
[0055] The data to be queried generally refers to data in any field and of any type. However, it should be understood that in the data query process targeted in the present application, when applied, the data to be queried corresponds to the data already stored in the data repository to be queried. For example, a large amount of heartbeat data has been stored in the data repository, and then based on this, in order to count the cardiac cycle, it is necessary to query and count the heartbeat data. In this scenario, the data to be queried refers to the heartbeat data to be queried.
[0056] Among them, the statistical result query instruction of the data to be queried can be triggered by the user on the interface of the computer device, can also be automatically sent by other devices according to the program, or can also be automatically triggered by the computer device during the response to other instructions. The embodiments of the present application do not limit the triggering timing, triggering method, etc. of the query instruction.
[0057] In practical applications, in many application scenarios, the data collected by the acquisition device has different distribution characteristics. It is relatively dense in some ranges and relatively sparse in other ranges. Based on this, different distribution ranges can be divided according to the distribution characteristics of the data in the pre-collected data set, that is, the original data set. For example, the original data set can be divided into distribution range 1, distribution range 2, distribution range 3,..., distribution range N in the order from dense to sparse according to the density. Then for the computer device, when it receives the statistical result query instruction of the data to be queried, it first determines the distribution range to which the data to be queried belongs in the original data set.
[0058] S102. Determine the query method according to the distribution range to which the data to be queried belongs in the original data set.
[0059] Among them, the query methods corresponding to the data in different distribution ranges are different.
[0060] Multiple query methods are preset in advance, such as the array method, the linked list method, the conditional query method, the instruction query method, and so on. For different distribution ranges, according to the advantages and disadvantages of each query method, a suitable query method is set for each distribution range. For example, distribution range 1 corresponds to the array method, and distribution range 2 corresponds to the linked list method.
[0061] Then, after the computer device determines the distribution range to which the data to be queried belongs in the original data set in the above steps, according to the distribution range to which it belongs, it directly determines the query method corresponding to this distribution range as the query method.
[0062] S103. Search for the statistical result of the data to be queried in the data storage library through the query method.
[0063] The computer device uses this query method to search for the statistical result of the data to be queried in the data storage library. Among them, the statistical information of the data set is stored in the data storage library, and this statistical result can represent the statistical result of the data to be queried. For example, at a certain moment, the existence or non-existence of the data to be queried; or the number of times the data to be queried appears; or the specific value of the data to be queried, and so on. The embodiments of the present application do not limit the statistical dimensions of the data set in the data storage library.
[0064] The data processing method improved in this embodiment, in response to a statistical result query instruction for data to be queried, obtains the distribution range to which the data to be queried belongs in the original dataset, determines a query method according to the distribution range to which the data to be queried belongs in the original dataset, and then searches for the statistical result of the data to be queried in the data repository through the query method. After the computer device receives the query instruction, it will first determine the distribution range to which the data to be queried belongs, and based on this belonging distribution range, a query method will be determined. In this way, performing the query operation in a query method suitable for its belonging distribution range can improve the performance when searching for data, improve the overall data processing performance of the system, and thus improve the human-computer interaction efficiency.
[0065] Based on the above embodiment, through an embodiment, the process of obtaining the distribution range of the data to be queried in the original dataset is described in detail. In one embodiment, S101 above includes the following steps: determining the distribution range to which the data to be queried belongs in the original dataset from each distribution range in the original dataset according to the position information of the data to be queried in the original dataset; wherein, each distribution range in the original dataset is obtained by dividing the original dataset according to the distribution characteristic type of the data in the original dataset.
[0066] The distribution characteristic types of the data in different distribution ranges are different.
[0067] First, analyze the distribution characteristic type of the data in the original dataset. For example, by analyzing the original dataset as a whole from dimensions such as central tendency, dispersion degree, and distribution shape, the distribution characteristic type of the data in the original dataset can be determined.
[0068] In one embodiment, the list method and the graphing method can be used to analyze the coordinate positions of each data in the original dataset. For example, the list method can be to express the data in a list form to analyze the rules of the data at different positions. Another example is that the graphing method can be to express the change relationship of the coordinate positions of each data in a prominent visual way, etc. Based on these data relationships, the distribution characteristic type of the data in the original dataset can be analyzed.
[0069] In another embodiment, analyzing the distribution characteristic type of the original dataset can be to use a preset neural network model for analysis. For example, a neural network model is pre-trained, and the original dataset is input into this neural network model, and the output result is the distribution characteristic type of the data in this original dataset.
[0070] After determining the distribution characteristic type of the data in the original dataset, the original dataset is divided into multiple distribution ranges according to this distribution characteristic type, and the distribution characteristic types of the data in the divided distribution ranges are different.
[0071] Exemplarily, such as Figure 3As shown, taking heartbeat data as an example, assume Figure 3 is the original data set of heartbeat data. The distribution of the points shown thereon represents the distribution of the original data set of heartbeat data. From Figure 3 it can be seen that the original heartbeat data set is divided into four distribution ranges. Among them, the data distribution characteristic type in distribution range 1 is dense distribution, the data distribution characteristic type in distribution range 2 is sub-dense distribution, the data distribution characteristic type in distribution range 3 is sub-sparse distribution, and the data distribution characteristic type in distribution range 4 is sparse distribution.
[0072] Among them, dense means that the statistical result quantity of the data is large, and the data distribution is relatively dense. Sparse means that the statistical result quantity of the data is small, and the distribution is relatively sparse. Naturally, the above-mentioned dense distribution, sub-dense distribution, sub-sparse distribution, and sparse distribution refer to the levels divided in the order of the statistical result quantity of the data from more to less, that is, the statistical result quantity and the density degree of the data in the dense distribution are both greater than those of the data in the sub-dense distribution; and the statistical result quantity of the data in the sub-sparse distribution is greater than that of the data in the sparse distribution, and the sparse degree of the sub-sparse distribution is less than the dense degree of the sparse distribution.
[0073] It should be noted that the process of dividing the distribution range of the original data set can be completed in advance, that is, in practical applications, before performing the step of determining the distribution range to which the data to be queried belongs in the original data set from each distribution range of the original data set according to the position information of the data to be queried in the original data set, multiple distribution ranges of the original data set have been determined in advance.
[0074] The position information of the data to be queried in the original data set represents the specific position of the data to be queried in the original data set. Based on this position information and the multiple distribution ranges that have been determined for the original data set, it can be determined which distribution range the data to be queried belongs to, so as to determine the distribution range to which the data to be queried belongs in the original data set.
[0075] For example, if the specific position of the data to be queried in Figure 3 is the position where point A is located, then the distribution range to which the data to be queried belongs in the original data set is distribution range 4.
[0076] In this embodiment, according to the position information of the data to be queried in the original dataset, the distribution range to which the data to be queried belongs in the original dataset is determined from each distribution range in the original dataset; wherein, each distribution range in the original dataset is obtained by dividing the original dataset according to the distribution characteristic type of the data in the original dataset; the distribution characteristic types of the data in different distribution ranges are different. In this way, by analyzing the distribution characteristic types of the data, the distribution ranges of the original dataset can be effectively and accurately divided. In this way, the distribution range to which the data to be queried belongs can be determined according to the position information of the data to be queried, improving the accuracy and efficiency of determining the distribution range to which the data to be queried belongs.
[0077] Based on any of the above embodiments, after determining the distribution range corresponding to the data to be queried in the original dataset, the query method for querying the statistical result of the data to be queried can be determined according to the distribution characteristic type of the data in this distribution range. In one embodiment, as Figure 4 shown, the above S102 includes:
[0078] S301, obtain the distribution characteristic type of the data in the distribution range to which it belongs.
[0079] Based on the above embodiments, it can be known that the distribution characteristic types of the data in different distribution ranges divided in the original dataset are different, that is, the distribution characteristic type of the data in each distribution range has been determined. Then, based on this, after determining the distribution range to which the data to be queried belongs in the original dataset, the distribution characteristic type of the data in this distribution range to which it belongs can be directly obtained. For example, if the distribution range to which the data to be queried belongs is the distribution range 1 in the above Figure 3 , and the distribution characteristic type of the data in the distribution range 1 is dense distribution, it can be determined that the distribution characteristic type of the data in the distribution range to which the data to be queried belongs is dense distribution.
[0080] Optionally, when the computer device obtains the distribution characteristic type of the data in the distribution range to which the data to be queried belongs in the original dataset, it can directly determine the distribution characteristic type of the data in the distribution range to which it belongs according to the correspondence relationship between the distribution range and the distribution characteristic type pre-stored in the memory, so as to improve the efficiency and accuracy of determining the distribution characteristic type of the data in the distribution range to which it belongs.
[0081] S302, according to the distribution characteristic type, determine the storage method of the statistical result of the data in the distribution range to which it belongs in the data repository.
[0082] Among them, the storage methods of the statistical results of the data in different distribution ranges in the data repository are different.
[0083] Based on the above distribution characteristic types, the computer device further determines the storage method of the statistical results of the data in the data repository within the determined distribution range.
[0084] In the embodiments of the present application, what is queried is the statistical result of the data. Correspondingly, what is stored in the database is also the statistical result of the data. Among them, the statistical result represents the result after summarizing and analyzing the data in a specific dimension. For example, the statistical result can be the number of occurrences, the existence or not, the proportion in the total, etc. The embodiments of the present application do not limit the statistical result.
[0085] Among them, the storage method refers to the storage method of the statistical results of the data in each distribution range in the original data set. In actual applications, in some scenarios, the statistical results of the data in each distribution range can be stored according to the data distribution characteristics of each distribution range.
[0086] In one embodiment, a realizable way to determine the storage method of the statistical results of the data in the data repository within the determined distribution range according to the distribution characteristic type includes: if the distribution characteristic type is dense distribution, the storage method of the statistical results of the data in the data repository is the array storage method; if the distribution characteristic type is sparse distribution, the storage method of the statistical results of the data in the data repository is the linked list storage method.
[0087] Differentiating the storage method by the type of distribution characteristic, if the distribution characteristic type of the statistical results of the data in the determined distribution range is dense distribution, then determine that the storage method of the statistical results of the data in the data repository is the array storage method. However, if the distribution characteristic type of the statistical results of the data in the determined distribution range is sparse distribution, then determine that the storage method of the statistical results of the data in the data repository is the linked list storage method.
[0088] S303. Determine the query method corresponding to the storage method.
[0089] Different storage methods correspond to the query methods for the statistical results of the data stored by using this storage method.
[0090] In one embodiment, determining the query method corresponding to the storage method includes the following two scenarios:
[0091] Scenario A: If the storage method is the array storage method, determine that the query method corresponding to this storage method is the array query method.
[0092] In scenario A, the array storage method is used as an example for illustration. Among them, the array storage method refers to a method of allocating a continuous and same-sized space in memory to store the statistical results of data. For the statistical results of data stored in the array storage method, the query method corresponding to the distribution range to which the data belongs is the data query method.
[0093] For example, if the range to which the data to be queried belongs is Figure 3 the distribution range 1 in, and the data distribution characteristic type in this distribution range is dense distribution. After using the array method to statistically store the statistical results of the data in distribution range 1, when looking up the statistical results of the corresponding data, the array query method is used for lookup. In this way, even if the data scale in distribution range 1 is large, the address where the statistical results of the data to be queried are located can be quickly queried, and then the statistical results of the data to be queried are looked up according to this address.
[0094] Scenario B: If the storage method is the linked list storage method, it is determined that the query method corresponding to this storage method is the linked list query method.
[0095] In scenario B, the linked list storage method is used as an example for illustration. Among them, the linked list storage method refers to a method of statistically storing data in the form of a linked list. For the statistical results of data stored in the linked list storage method, the query method corresponding to the distribution range to which the data belongs is the linked list query method.
[0096] For example, if the range to which the data to be queried belongs is Figure 3 the distribution range 4 in, and the data distribution characteristic in this distribution range is sparse distribution and there will be no overlapping situation. After using the linked list method to statistically store the statistical results of the data in distribution range 4, when looking up, the linked list query method is also used for lookup. When looking up in the linked list, it starts from the starting element of the linked list and sequentially looks for the corresponding data backward until the corresponding data is found. In this way, because the data in distribution range 4 is sparsely distributed over a wide range, when traversing according to the elements in the linked list, the number of traversals can be greatly reduced, thereby improving the query efficiency and avoiding resource waste.
[0097] In this embodiment, for the intervals where the data distribution characteristics are relatively dense and there are many duplicate data, the array statistics method is used. For the intervals where the data distribution characteristics are relatively sparse and the data is not common, the linked list statistics method is used. In this way, in actual statistics, since there is a lot of data in the duplicate data intervals, using array statistics can improve performance. For the uncommon intervals, the linked list statistics method can reduce the number of traversals and save system resources.
[0098] Furthermore, based on any of the above embodiments, taking the array query method and the linked list query method as examples, the specific query methods when using these two query methods are specifically described respectively.
[0099] In one embodiment, as Figure 5 shown, the statistical result of the data to be queried is found in the data repository by querying, including:
[0100] S401, if the query method is an array query method, determine the first identifier corresponding to the data to be queried; wherein, the first identifier is used to identify the element of the array storing the statistical result of the data to be queried in the data repository.
[0101] S402, determine the address of the element corresponding to the first identifier in the data repository according to the first identifier and the starting address of the array.
[0102] S403, determine the statistical result of the data to be queried according to the content stored in the address.
[0103] This embodiment is directed to the process of querying using the array query method, that is, when the query method is the array query method, first determine the first identifier corresponding to the data to be queried.
[0104] Among them, the first identifier of the data to be queried can be the element of the array storing the statistical result of the data to be queried in the data repository. For example, the first identifier can be the one-dimensional data generated according to the two-dimensional coordinates of the data to be queried. For example, taking the heartbeat data as an example, its corresponding coordinate system is a 2000*2000 coordinate system, and the two-dimensional coordinates of the heartbeat data are (1200, 1400), which represents that the time interval of the previous beat in the heartbeat is 1200 milliseconds, and the time interval of the next beat is 1400 milliseconds. Based on this, the method of converting the two-dimensional coordinates (1200, 1400) of the heartbeat data into one-dimensional data is: 1200*2000 + 1400 = 2401400, that is, the first identifier of the heartbeat data is 2401400.
[0105] Based on the first identifier, determine the address of the element corresponding to the first identifier in the data repository according to the first identifier of the data to be queried and the starting address of the array.
[0106] Specifically, this embodiment is directed to the process of querying using the array query method, and its storage is also in the form of a one-dimensional array. It should be emphasized that the arrays in the embodiments of this application all refer to one-dimensional arrays. Then, in the one-dimensional array, according to the first identifier of the data to be queried and the starting address of the array, the offset from the first identifier to the starting address of the array can be determined, and this offset is equivalent to the distance from the address of the element corresponding to the first identifier in the data repository to the starting address of the array. Then, based on this offset, the address of the element corresponding to the first identifier in the data repository can be determined.
[0107] The content stored in the address is the statistical result of the data to be queried.
[0108] Specifically, when storing in an array, the value stored in the address of each element of the array can be set according to the actual situation. For example, please refer to Figure 6 as shown in Figure 6 the values in the array in
[0109] include 1 and 0, where 1 indicates that there is corresponding data to be queried at this address, and 0 indicates that there is no corresponding data to be queried at this address, that is, the statistical result is whether the data to be queried exists or not.
[0110] Continue to refer to Figure 6 as shown in Figure 6 Assume that the first identifier of the data to be queried is 99999997, and the statistical result is based on whether 99999997 exists in the array. Then, when querying the statistical result of the data to be queried, it is necessary to determine the offset as 99999997 according to 99999997 and the starting address of the array. Then, offset the starting address of the array by 99999997, and only need to check whether the value corresponding to the 99999997th element is 1. If so, it is determined that the 99999997 to be queried exists in the array.
[0111] In this embodiment, when querying data using the array query method, only need to determine the address of the element corresponding to the first identifier in the data storage library according to the first identifier corresponding to the data to be queried and the starting address of the array, and then directly search at the corresponding address. In this way, since there are many data in the repeated data intervals, directly searching at the corresponding address can improve the query performance and make the query efficiency higher.
[0112] In another embodiment, as Figure 7 shown in
[0113] S501, if the query method is the linked list query method, determine the second identifier corresponding to the data to be queried; wherein, the second identifier is used to identify the node of the linked list storing the statistical result of the data to be queried in the data storage library.
[0114] S502: Starting from the head node of the linked list, in the order of the links of the linked list, compare the second identifier corresponding to the data to be queried with the identifier of each node in the linked list one by one until a node matching the second identifier is found; determine the statistical result of the data to be queried according to the content stored in the node.
[0115] This embodiment is directed to the process of querying using the linked list query method. That is, when the query method is the linked list query method, it is necessary to start from the starting element of the linked list and search one by one in the order of the links of the linked list.
[0116] Specifically, the search process is to compare the second identifier corresponding to the data to be queried with the identifier of each node in the linked list, comparing them one by one until a node matching the second identifier is found, and determine the statistical result of the data to be queried according to the content stored in the node.
[0117] Among them, the content stored in the nodes of the linked list can be set according to the actual situation. For example, similar to the above array, it can only store 1 or 0 to indicate whether the statistical result exists or not; or store a positive integer greater than or equal to 0 to indicate the number of times the data to be queried appears.
[0118] Or, as Figure 8 shown, the second identifier corresponding to the data to be queried is the one-dimensional data after the two-dimensional coordinates of the data to be queried are converted, that is, the second identifier is 45847. Then, taking whether the statistical result 45847 exists in the linked list as an example, during the query, it is necessary to traverse the entire linked list and perform matching one by one until the matching is successful to determine the statistical result of the data to be queried.
[0119] It should be noted that Figure 8 Figure shows the schematic diagram of the second identifier corresponding to the data to be queried stored in the linked list node. However, in actual applications, whether it is the linked list method or the array method, it can be one of the relevant statistical information such as the first identifier / second identifier of the data to be queried, the data itself, the flag value indicating whether it exists, the number of occurrences, etc., or any combination of several of them stored in the elements of the linked list. In this way, when the data to be stored is queried, other statistical information can also be obtained together.
[0120] In this embodiment, when querying data using the linked list method, it is necessary to search backward one by one from the head node of the linked list until the corresponding data is found. For uncommon intervals, the data scale is relatively sparse and there will be no overlap, which can reduce the number of traversals and thus achieve the effect of saving system resources.
[0121] In practical applications, in many application scenarios, most of the data collected by the acquisition device is concentrated in certain intervals. As a result, most intervals in the array statistics will have no values, leading to a waste of space. Moreover, due to the large range of the data itself, the entire array occupies too much space, wasting system resources.
[0122] For example, taking heartbeat data as an example, in order to count the RR intervals of heartbeats, it is necessary to count ultra-large-scale heartbeat data. Please refer to Figure 9 as shown in Figure 9 Fig. is a schematic diagram of a Lorenz scatter plot. Among them, the Lorenz scatter plot uses adjacent RR intervals (the time required from the start of one ventricular depolarization to the start of the next ventricular depolarization) as the horizontal and vertical coordinates, and depicts the RR interval scatter set of long-term dynamic electrocardiogram in a plane rectangular coordinate system. A scatter point in the Lorenz diagram represents two adjacent RR intervals.
[0123] However, due to the characteristics of the heartbeat data itself, the vast majority of the heartbeat data is within a certain range. Then, according to the data processing method provided by the embodiments of the present application, the area where the heartbeat data is mainly concentrated can be counted in the form of an array, while other less common areas (such as noise points, etc.) can be counted in the form of a linked list. In this way, since there is a large amount of heartbeat data in the concentrated area, using an array to count can improve the query performance. For less common areas, the traversal times can be reduced and system resources can be saved through the method of linked list statistics.
[0124] As Figure 10 shown, the embodiments of the present application also provide a data processing method, including the following steps:
[0125] S1, in response to a statistical result query instruction of the data to be queried, analyze the distribution characteristic type of the original data set.
[0126] S2, divide the original data set according to the distribution characteristic type to obtain at least one distribution range; the distribution characteristic types of the data in different distribution ranges are different.
[0127] S3, according to the position information of the data to be queried, determine the distribution range to which the data to be queried belongs in the original data set from each distribution range.
[0128] S4, obtain the distribution characteristic type of the data in the belonging distribution range; if the distribution characteristic type is dense distribution, the storage method of the statistical result of the data in the belonging distribution range in the data storage library is an array storage method; if the distribution characteristic type is sparse distribution, the storage method of the statistical result of the data in the belonging distribution range in the data storage library is a linked list storage method.
[0129] S5. If the storage method is the array storage method, determine that the query method corresponding to the storage method is the array query method; if the storage method is the linked list storage method, determine that the query method corresponding to the storage method is the linked list query method.
[0130] S6. If the query method is the array query method, determine the address of the element corresponding to the first identifier in the data storage repository according to the first identifier of the data to be queried and the starting address of the array.
[0131] S7. Determine the statistical result of the data to be queried according to the content stored in the address.
[0132] S8. If the query method is the linked list query method, determine the second identifier corresponding to the data to be queried, and starting from the head node of the linked list, compare the second identifier corresponding to the data to be queried with the identifier of each node in the linked list in the linked order until a node matching the second identifier is found.
[0133] S9. Determine the statistical result of the data to be queried according to the content stored in the node.
[0134] For the specific limitations of the data processing method provided in this embodiment, reference can be made to the limitations of each step in the data processing method in the above text, which will not be elaborated here.
[0135] It should be understood that although each step in the flowchart attached in the above embodiment is displayed in sequence according to the indication of the arrow, these steps do not necessarily execute in the order indicated by the arrow. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the figures attached in the above embodiment may include multiple steps or multiple stages. These steps or stages do not necessarily execute at the same moment, but can execute at different moments, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0136] In one embodiment, as Figure 11 shown, the embodiment of the present application further provides a data processing device, and the device includes: an acquisition module 10, a determination module 11, and a search module 12;
[0137] The acquisition module 10 is configured to obtain the distribution range to which the data to be queried belongs in the original data set in response to a statistical result query instruction of the data to be queried;
[0138] The determination module 11 is configured to determine a query method according to the distribution range to which the data to be queried belongs in the original data set;
[0139] A search module 12 for searching for the statistical result of the data to be queried in a data repository by querying.
[0140] In one embodiment, the above-mentioned acquisition module 10 is further configured to determine, according to the position information of the data to be queried in the original dataset, the distribution range to which the data to be queried belongs in the original dataset from each distribution range in the original dataset; each distribution range in the original dataset is obtained by dividing the original dataset according to the distribution characteristic type of the data in the original dataset.
[0141] In one embodiment, the above-mentioned determination module 11 includes:
[0142] A characteristic type acquisition unit for acquiring the distribution characteristic type of the data in the distribution range to which it belongs;
[0143] A storage method determination unit for determining the storage method of the statistical result of the data in the distribution range to which it belongs in the data repository according to the distribution characteristic type;
[0144] A query method determination unit for determining the query method corresponding to the storage method.
[0145] In one embodiment, the above-mentioned storage method determination unit is further configured to, if the distribution characteristic type is dense distribution, the storage method of the statistical result of the data in the distribution range to which it belongs in the data repository is an array storage method;
[0146] If the distribution characteristic type is sparse distribution, the storage method of the statistical result of the data in the distribution range to which it belongs in the data repository is a linked list storage method.
[0147] In one embodiment, the above-mentioned query method determination unit is further configured to, if the storage method is an array storage method, determine that the query method corresponding to the storage method is an array query method; if the storage method is a linked list storage method, determine that the query method corresponding to the storage method is a linked list query method.
[0148] In one embodiment, the above-mentioned search module 12 includes:
[0149] An identifier determination unit for determining the first identifier corresponding to the data to be queried if the query method is an array query method; wherein, the first identifier is used to identify the element of the array storing the statistical result of the data to be queried in the data repository;
[0150] An address determination unit for determining the address of the element corresponding to the first identifier in the data repository according to the first identifier and the start address of the array;
[0151] A first statistical result determination unit for determining the statistical result of the data to be queried according to the content stored in the address.
[0152] In one embodiment, the above-mentioned search module 12 includes:
[0153] A search unit, configured to determine a second identifier corresponding to the data to be queried if the query method is a linked list query method; wherein, the second identifier is used to identify a node of the linked list storing the statistical result of the data to be queried in the data repository;
[0154] A second statistical result determination unit, configured to start from the head node of the linked list, and sequentially compare the second identifier corresponding to the data to be queried with the identifiers of each node in the linked list according to the link order of the linked list until a node matching the second identifier is found; determine the statistical result of the data to be queried according to the content stored in the node.
[0155] For the specific limitations of the data processing device, reference can be made to the limitations of each step in the data processing method described above, which will not be elaborated here. Each module in the above data processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the target device in hardware form or independent of the target device, or stored in the memory of the target device in software form, so that the target device can call and execute the operations corresponding to the above modules.
[0156] In one embodiment, a computer device is provided. The computer device can be the above data processing device. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a data processing method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0157] Those skilled in the art can understand that the above structural description of the computer device is only a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0158] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0159] In response to a statistical result query instruction of data to be queried, obtain the distribution range of the data to be queried in the original dataset;
[0160] Determine a query method according to the distribution range of the data to be queried in the original dataset;
[0161] Find the statistical result of the data to be queried in the data repository by the query method.
[0162] When the computer device provided in the above embodiment implements the above steps, its implementation principle and technical effect are similar to the principle of the method steps executed by the above data processing method, and will not be elaborated here.
[0163] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps executed by the processor in the above radiotherapy plan planning system are implemented:
[0164] In response to a statistical result query instruction of data to be queried, obtain the distribution range of the data to be queried in the original dataset;
[0165] Determine a query method according to the distribution range of the data to be queried in the original dataset;
[0166] Find the statistical result of the data to be queried in the data repository by the query method.
[0167] When the computer-readable storage medium provided in the above embodiment implements the above steps, its implementation principle and technical effect are similar to the principle of the method steps executed by the above data processing method, and will not be elaborated here.
[0168] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0169] In response to a statistical result query instruction of data to be queried, obtain the distribution range of the data to be queried in the original dataset;
[0170] Determine a query method according to the distribution range of the data to be queried in the original dataset;
[0171] Find the statistical result of the data to be queried in the data repository by the query method.
[0172] When implementing the above steps, the implementation principle and technical effects of a computer program product provided by the above embodiments are similar to the method steps executed by the above data processing method, and will not be elaborated here.
[0173] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0174] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0175] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A data processing method, characterized in that, The method includes: In response to a statistical result query instruction for the data to be queried, obtaining the distribution range to which the data to be queried belongs in the original dataset; Determining a query method according to the distribution range to which the data to be queried belongs in the original dataset; Searching for the statistical result of the data to be queried in the data repository through the query method; Wherein, the determining a query method according to the distribution range to which the data to be queried belongs in the original dataset includes: Obtaining the distribution characteristic type of the data in the belonging distribution range; Determining a storage method of the statistical result of the data in the belonging distribution range in the data repository according to the distribution characteristic type; the storage method is an array storage method or a linked list storage method; Determining the query method corresponding to the storage method.
2. The data processing method according to claim 1, wherein The obtaining the distribution range to which the data to be queried belongs in the original dataset includes: According to the position information of the data to be queried in the original dataset, determining the distribution range to which the data to be queried belongs in the original dataset from each distribution range in the original dataset; Wherein, each distribution range in the original dataset is obtained by dividing the original dataset according to the distribution characteristic type of the data in the original dataset.
3. The data processing method according to claim 1, wherein The determining the storage method of the statistical result of the data in the belonging distribution range in the data repository according to the distribution characteristic type includes: If the distribution characteristic type is dense distribution, the storage method of the statistical result of the data in the belonging distribution range in the data repository is an array storage method; If the distribution characteristic type is sparse distribution, the storage method of the statistical result of the data in the belonging distribution range in the data repository is a linked list storage method.
4. The data processing method according to claim 3, wherein The determining the query method corresponding to the storage method includes: If the storage method is an array storage method, determining that the query method corresponding to the storage method is an array query method; If the storage method is a linked list storage method, determining that the query method corresponding to the storage method is a linked list query method.
5. The data processing method according to claim 1 or 2, characterized in that, The searching for the statistical result of the data to be queried in the data repository through the query method includes: If the query method is an array query method, determining a first identifier corresponding to the data to be queried; wherein, the first identifier is used to identify an element of the array storing the statistical result of the data to be queried in the data repository; Determining the address of the element corresponding to the first identifier in the data repository according to the first identifier and the start address of the array; Determining the statistical result of the data to be queried according to the content stored in the address.
6. The data processing method according to claim 1 or 2, characterized in that The searching for the statistical result of the data to be queried in the data repository through the query method includes: If the query method is a linked list query method, determining a second identifier corresponding to the data to be queried; wherein, the second identifier is used to identify a node of the linked list storing the statistical result of the data to be queried in the data repository; Starting from the head node of the linked list, compare the second identifier corresponding to the data to be queried with the identifiers of each node in the linked list in sequence according to the link order of the linked list until a node matching the second identifier is found; Determine the statistical result of the data to be queried according to the content stored in the node.
7. A data processing device, characterized in that, The device includes: An acquisition module, configured to acquire the distribution range to which the data to be queried belongs in the original data set in response to a statistical result query instruction of the data to be queried; A determination module, configured to determine a query method according to the distribution range to which the data to be queried belongs in the original data set; A search module, configured to search for the statistical result of the data to be queried in the data storage library by the query method; Obtain the distribution characteristic type of the data in the obtained distribution range; According to the distribution characteristic type, determine a storage method of the statistical result of the data in the obtained distribution range in the data storage library; the storage method is an array storage method or a linked list storage method; Determine the query method corresponding to the storage method.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Data processing method and device, electronic equipment and storage medium
CN113609313A