Big data storage method and system applied to laboratory

By extracting data features and classifying models, laboratory data storage is dynamically adjusted, solving the problems of high storage pressure and long response time in traditional storage methods, and achieving rapid data access and efficient operation of the storage system.

CN120669916APending Publication Date: 2025-09-19河北省畜牧总站(河北省奶源工作总站)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510778053.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Traditional laboratory data storage methods result in high pressure on storage devices and long response times, and cannot effectively meet the demand for fast data access during scientific research.

Method used

Through data feature extraction and machine learning model classification, laboratory data is dynamically migrated to a suitable storage area, and the different storage characteristics of the hot data layer, warm data layer, and cold data layer are utilized to achieve dynamic adjustment and optimized storage of data.

Benefits of technology

It reduces the waiting time for data access, improves the performance and efficiency of the laboratory data storage system, and meets the demand for fast data access in the scientific research process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669916A_ABST
    Figure CN120669916A_ABST
Patent Text Reader

Abstract

The invention provides a big data storage method and system applied to a laboratory, and belongs to the technical field of data storage, the method is applied to a controller of a storage device, and the method comprises the following steps: in response to receiving first data, performing feature extraction on the first data to obtain a plurality of target features; the first data is to-be-stored laboratory data; inputting the first data and the plurality of target features into a target classification model to obtain a classification result of the first data; storing the first data to a corresponding storage area in the storage device based on the classification result; in response to the fact that the access feature of the second data in the first time period does not meet the access feature of the storage area where the second data is located, migrating the second data to the corresponding storage area to be stored based on the access feature of the second data; the second data is data stored in the storage device. According to the big data storage method and system applied to the laboratory, the laboratory data can be accurately classified and stored, and the overall response time is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of data storage technology, and more specifically, relates to a method and system for storing big data applied to a laboratory. Background Art

[0002] As scientific research equipment becomes increasingly digital, the amount of data generated in laboratories is growing exponentially. Traditional laboratory data storage methods often use a "hot and cold" two-tier architecture, with categorized storage based primarily on timestamps or file types and fixed data storage areas. This results in high storage pressure and long response times for existing laboratory storage devices.

[0003] Therefore, a big data storage method for laboratory use is needed. Summary of the Invention

[0004] The purpose of this application is to provide a big data storage method and system for laboratory use, so as to accurately classify and store laboratory data and reduce the overall response time.

[0005] In a first aspect of an embodiment of the present application, a method for storing big data in a laboratory is provided, which is applied to a controller of a storage device, including: In response to receiving first data, performing feature extraction on the first data to obtain a plurality of target features; the first data is laboratory data to be stored; Inputting the first data and the plurality of target features into a target classification model to obtain a classification result of the first data; storing the first data in a corresponding storage area in a storage device based on the classification result; In response to the access characteristics of the second data in the first time period not satisfying the access characteristics of the storage area where the second data is located, the second data is migrated to its corresponding storage area for storage based on the access characteristics of the second data; the second data is the data stored in the storage device.

[0006] A second aspect of the embodiments of the present application provides a big data storage system for a laboratory, comprising: a feature extraction module, configured to extract features from the first data in response to receiving the first data, to obtain a plurality of target features; the first data being laboratory data to be stored; a classification module, configured to input the first data and a plurality of target features into a target classification model to obtain a classification result of the first data; A storage module, configured to store the first data in a corresponding storage area in a storage device based on the classification result; A storage adjustment module is used to migrate the second data to its corresponding storage area for storage based on the access characteristics of the second data in response to the access characteristics of the second data in the first time period not meeting the access characteristics of the storage area where the second data is located; the second data is the data stored in the storage device.

[0007] According to a third aspect of an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above-mentioned big data storage method applied to a laboratory are implemented.

[0008] In a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned big data storage method applied to a laboratory are implemented.

[0009] The beneficial effects of the big data storage method and system for laboratory use provided by the embodiments of the present application are: The storage method of this application, based on data features and model classification, can more accurately classify and store data according to its actual attributes and usage requirements, avoiding unreasonable allocation of data in storage areas, thereby effectively reducing the storage pressure of storage devices. This application uses a dynamic data migration mechanism to enable data to adjust its storage location according to actual access conditions, migrating frequently accessed data to storage areas that are more suitable for fast access, reducing the waiting time during data access, thereby shortening the overall response time, improving the performance and efficiency of the laboratory data storage system, and better meeting the demand for fast data access in the scientific research process. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] Figure 1 A flowchart of a big data storage method applied to a laboratory provided in one embodiment of the present application; Figure 2 This is a structural block diagram of a big data storage system for use in a laboratory, provided in one embodiment of the present application; Figure 3 A schematic block diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0012] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0013] In order to make the purpose, technical solutions and advantages of this application clearer, specific embodiments will be described below with reference to the accompanying drawings.

[0014] Please refer to Figure 1 , Figure 1 A flowchart of a big data storage method applied to a laboratory is provided in one embodiment of the present application. The method is applied to a controller of a storage device and includes: S101-S103.

[0015] S101: In response to receiving first data, extracting features from the first data to obtain a plurality of target features; the first data is laboratory data to be stored.

[0016] In this embodiment, when the storage device controller detects that new laboratory data needs to be stored, the subsequent processing flow is triggered. In this embodiment of the application, "receiving" can be achieved through network transmission, a local interface (such as USB), or other data input methods, with the storage device controller ultimately acquiring the data. The first data can be various laboratory data to be stored, such as milk content test data, livestock environment data (such as sewage and treated sewage), or livestock excrement sampling test results.

[0017] Feature extraction is performed on the first data to obtain multiple target features. These target features can include basic features such as data generation time, file size, and data format; business features such as experiment type (e.g., gene sequencing, chemical reaction), equipment number, and data sensitivity (confidential / public). In this embodiment, feature extraction generates a set of structured feature vectors that describe the data.

[0018] It should be noted that the basic features can be obtained through the first data itself, and the business features are not data features inherent in the data, but features assigned to the first data by relevant personnel when storing the first data.

[0019] S102: Input the first data and multiple target features into a target classification model to obtain a classification result of the first data.

[0020] In this embodiment, the target classification model is a pre-trained machine learning model used to classify input data, and can be a decision tree model, a random forest model, or a support vector machine model. The training dataset for the target classification model is the laboratory's historical data, the features of each historical data, and the corresponding classification results. The classification results of the first data can be hot data, warm data, or cold data, etc.

[0021] In this example, hot data is frequently accessed and time-sensitive data, requiring fast read and write times. Warm data is moderately accessed and less time-sensitive data, requiring no real-time access but occasionally requiring retrieval or analysis. Cold data is rarely accessed and archived for long periods of time, used for compliance storage or historical reference.

[0022] S103: Storing the first data in a corresponding storage area in a storage device based on the classification result.

[0023] As mentioned above, the classification result can be hot data, warm data or cold data. In this embodiment, each of them can be corresponded to a storage area. For example, in response to the classification result being hot data, the first data is stored in the hot data layer; in response to the classification result being warm data, the first data is stored in the warm data layer; in response to the classification result being cold data, the first data is stored in the cold data layer.

[0024] In this embodiment, the hot data tier can be an SSD or memory, the warm data tier can be an HDD, and the cold data tier can be a tape library, cloud-based cold storage, etc. The hot data tier has the highest storage cost and the fastest read speed, while the cold data tier has the lowest storage cost and the slowest read speed. Through classification and grading, the overall storage cost can be reduced and the overall data read speed can be improved.

[0025] S104: In response to the access characteristics of the second data in the first time period not satisfying the access characteristics of the storage area where the second data is located, migrate the second data to its corresponding storage area for storage based on the access characteristics of the second data; the second data is data stored in the storage device.

[0026] In this embodiment, the first time period can be a preset time period, such as 1 day, 1 week or 1 month, which can be determined by the manager of the laboratory storage device based on the amount of data in the laboratory and actual needs. The second data is the data already stored in the storage device, which can be understood as the first data becoming the second data after the above-mentioned feature extraction, classification and storage. It should be noted that in the embodiment of the present application, the process of changing from the first data to the second data does not change any data or feature in the first data, but it is redefined with a new name because it is located in a different area.

[0027] In this embodiment, the access feature may be the frequency, time interval, or response speed of data access. In this embodiment, the access frequency is used as an example: The storage areas of the storage device include hot data tier, warm data tier and cold data tier; In response to the second data being stored in the hot data layer and the second data being accessed at a frequency less than a hot data access threshold within the first time period, migrating the second data to a corresponding storage area for storage based on the second data being accessed; In response to the second data being stored in the warm data tier and the access frequency of the second data in the first time period being greater than the first warm data access threshold or the access frequency of the second data in the first time period being less than the second warm data access threshold, migrating the second data to a corresponding storage area for storage based on the access frequency of the second data; In response to the second data being stored in the cold data layer and the second data being accessed at a frequency greater than a cold data access threshold within the first period, the second data is migrated to a corresponding storage area for storage based on the second data access frequency.

[0028] The hot data access threshold is greater than the first warm data access threshold, the first warm data access threshold is greater than the second warm data access threshold, and the second warm data access threshold is greater than the cold data access threshold. The first time period can be determined according to actual application scenarios.

[0029] It should be noted that, in this embodiment, the access frequency intervals corresponding to the hot data layer, the warm data layer, and the cold data layer are continuous, but the aforementioned hot data access threshold, the first warm data access threshold, the second warm data access threshold, and the cold data access threshold are not the boundary values ​​of the access frequency intervals corresponding to the hot data layer, the warm data layer, and the cold data layer. For example, the access frequency interval corresponding to the hot data layer is The access threshold of hot data is b, and a is greater than b. That is, when determining data migration, it should be more relaxed than when determining storage, to avoid excessive migration leading to waste of resources and sudden access leading to long waiting times.

[0030] From the above, it can be concluded that the storage method based on data characteristics and model classification in this application can more accurately classify and store data according to its actual attributes and usage requirements, avoiding unreasonable allocation of data in storage areas, thereby effectively reducing the storage pressure of storage devices. This application uses a dynamic data migration mechanism to enable data to adjust its storage location according to actual access conditions, migrate frequently accessed data to storage areas that are more suitable for fast access, reduce waiting time during data access, and thus shorten the overall response time, improve the performance and efficiency of the laboratory data storage system, and better meet the demand for fast data access in the scientific research process.

[0031] In one embodiment of the present application, the storage area of ​​the storage device includes: a hot data layer, a warm data layer, and a cold data layer; the second data includes hot data and cold data; the hot data is the data stored in the hot data layer, and the cold data is stored in the cold data layer; A big data storage method applied to a laboratory also includes: In response to the remaining capacity of the hot data layer being less than a preset capacity, determining a migration weight for each hot data based on access characteristics and data characteristics of each hot data; migrating the hot data having a migration weight less than a first migration weight to the warm data layer; In response to a wait time for access requests to the first cold data being longer than the first time period during a second time period, the second cold data is migrated to the warm data tier; wherein the first cold data and the second cold data are both cold data, and a correlation between a characteristic of the second cold data and a characteristic of the first cold data is greater than the first correlation. The first migration weight and the first correlation may be determined based on experience.

[0032] In this embodiment, when the available storage space of the hot data layer falls below the preset capacity, it indicates that the data in the hot data layer is too large. At this time, the data migration mechanism is triggered to free up space to ensure the storage needs of high-frequency data. Therefore, the embodiment of the present application can determine the migration weight based on access characteristics and data characteristics. The smaller the migration weight, the higher the migration priority. The access characteristics can be the frequency of access, and the data characteristics can be the size, type, or generation time of the data.

[0033] This embodiment also considers that if the access latency of certain data (first cold data) in the cold data tier exceeds a tolerance threshold (e.g., the first latency is 30 seconds) within a set time period, indicating that the current storage tier cannot meet access efficiency requirements, second cold data can be migrated to the warm data tier. The correlation between the features of the second cold data and the features of the first cold data is greater than the first correlation. This embodiment can calculate the correlation between the features of the second cold data and the features of the first cold data using cosine similarity. The features of the first cold data and the second cold data are both vectors in nature.

[0034] In one embodiment of the present application, the method further includes: In response to the existence of target cold data in the cold data, where the target cold data is cold data whose access characteristics meet the deletion condition, sending a deletion request to the target device; In response to the deletion request being approved, the target cold data is deleted.

[0035] In this embodiment, the target cold data is data stored in the cold data tier whose access characteristics meet preset deletion criteria. This is typically archived data that has not been accessed for a long time, has no business value, or has been replaced by newer versions. The preset deletion criteria can be that the data has never been accessed and the storage duration exceeds the deletion period. The deletion period can be determined based on the number of actual laboratories and the storage performance of the storage device, or based on experience.

[0036] In this embodiment, the target device can be a mobile phone, tablet, computer or other device of the manager of the laboratory storage device, or it can be the host corresponding to the storage device. When the deletion request is approved, the target cold data can be deleted.

[0037] From the above, it can be concluded that the present application determines the migration weight and migrates data based on multi-dimensional features, which can reasonably release the hot data layer space, ensure the storage requirements of high-frequency access data in the hot data layer, avoid data access delays due to insufficient space, and improve the stability and response speed of the data storage system. The present application predicts subsequent access needs by calculating the correlation of data features, migrates related data in advance, avoids multiple low-speed retrievals, and reduces the waiting time for scientific researchers to obtain relevant data. The present application can clean up useless data in the storage device in a timely manner by monitoring data that has not been accessed for a long time and issuing deletion requests, freeing up storage space, improving the utilization rate of the storage device, and avoiding invalid data occupying storage resources.

[0038] In one embodiment of the present application, the target classification model is a target decision tree model; The process of determining the target decision tree model includes: Take the first data set as the root node; Calculating information gain of each feature in a first data set, and determining a split feature based on the information gain of each feature; the first data set includes historical data of a laboratory; Split the root node based on the split feature to obtain multiple child nodes; For each child node, calculate the information gain of each feature in the child node, and determine the splitting feature of the child node based on the information gain of each feature; Splitting the corresponding child nodes based on the splitting characteristics of each child node until the number of node samples meets the preset conditions, thereby obtaining a first decision tree; Perform pruning operation on the first decision tree to obtain the target decision tree.

[0039] In this embodiment, a target decision tree is selected as the target classification model of the embodiment of this application. The first dataset is a collection of historical laboratory data, including multiple samples and their characteristics, as well as corresponding classification results. This classification result can be determined and labeled based on expert experience or by laboratory personnel. The first dataset serves as the basic data for building the decision tree and is used for initial splitting.

[0040] In this embodiment, information gain is a measure of a feature's contribution to data classification. The greater the information gain, the more effectively the feature distinguishes between different categories. A split feature is the feature with the highest information gain at the current node. It is used to divide a node into multiple child nodes, determining the branching direction of the decision tree and forming the basis for constructing the decision tree. The preset condition is the termination condition, which can be when the number of node samples is less than a termination threshold. The termination threshold can be set to 10.

[0041] In this embodiment, after the preset conditions are met, what is obtained is not the target decision tree used for classification in the embodiment of the present application, but the first decision tree. The present application also takes into account the uneven distribution of data in the big data storage scenario in the laboratory, and the decision tree may overfit high-frequency data. Therefore, the present application performs a pruning operation on the first decision tree obtained by training to eliminate redundant branches.

[0042] Specifically, the first decision tree is pruned to obtain a target decision tree, including: Starting from the bottom node of the first decision tree, traverse toward the top node of the first decision tree; In response to the node being a leaf node, no pruning operation is performed; In response to the node not being a leaf node, determining a first loss value of the first decision tree based on the target loss function; Pruning the node and determining a second loss value of the first decision tree based on the loss target loss function; In response to the difference between the first loss value and the second loss value being smaller than a tolerance threshold, the pruning operation on the node is canceled until all nodes of the first decision tree are traversed to obtain a target decision tree.

[0043] In this embodiment, we can start from the leaf node of the decision tree and traverse upward layer by layer to the root node. The leaf node is the last node of each branch or each subtree, and there is no pruning operation available, so pruning operation cannot be performed.

[0044] In this embodiment, the objective loss function may be: ,in, represents the target loss function, is the error term, which represents the classification error rate of the current decision tree model on the validation set. Indicates the The data proportion of the layer, Indicates the The unit storage cost of the layer, is the weight coefficient, , can be determined based on multiple experiments, for example, . Represents the hot data layer, represents the warm data layer, Representing the cold data layer, the storage cost of each data layer (that is, each storage area) should be a quantified value that can be determined based on expert experience.

[0045] In the target loss function, The lower the value, the more accurate the model prediction and the lower the storage cost. Through weighted summation, the model stores data in the low-cost layer as much as possible, reducing the total overhead. As a result, the decision tree can accurately classify data after pruning and reduce storage costs.

[0046] The first loss value is the loss of the entire tree before pruning (that is, when the current node has not been pruned). The second loss value is the loss of the entire tree after pruning (that is, after the current node is replaced with a leaf node). The tolerance threshold is the upper limit of the allowed loss increase. If the loss increase after pruning does not exceed the threshold, the pruning is retained; otherwise, it is canceled.

[0047] From the above, it can be concluded that the present application traverses from the bottom node to the top node, calculates the loss value before and after pruning for non-leaf nodes based on the target loss function, and decides whether to cancel the pruning operation based on the comparison of the loss value difference with the tolerance threshold until all nodes are traversed. This can eliminate redundant branches in the decision tree, avoid overfitting the model to the training data, and enable the target decision tree to maintain good classification performance when facing new and unseen data, thereby enhancing the generalization ability of the model and ensuring the reliability of the classification results in practical applications. The embodiment of the present application can guide the model to store data in the low-cost layer as much as possible by taking a weighted sum of the data proportion and unit storage cost of different storage layers, thereby reducing the total storage cost.

[0048] In one embodiment of the present application, determining the migration weight of each hot data based on the access characteristics and data characteristics of each hot data includes: screening the access frequencies of the hot data based on the first condition, setting the migration weight of the hot data that meets the first condition as the second migration weight, and setting the hot data that does not meet the first condition as the target hot data; the second migration weight is greater than the first migration weight; For each target hot data, a weighted calculation is performed on the access frequency of the target hot data, the inverse of the storage space occupied by the data, and the difference between the generation time and the current time to obtain the migration weight of the target hot data.

[0049] In this embodiment, it is also considered that there is a lot of hot data in the hot data layer, and calculating the migration weights one by one consumes a lot of computing resources and time. Therefore, the embodiment of the present application first filters the hot data and calculates the migration weights of the filtered data.

[0050] In this embodiment, the first condition is that the access frequency of hot data is greater than the frequency of normal hot data, and the frequency of normal hot data is a preset frequency. When the frequency of hot data exceeds the frequency of normal hot data, it means that this data is frequently accessed data and should be retained. There is no need to calculate its corresponding migration weight. Its migration weight is directly set to the second migration weight, and the second migration weight is greater than the first migration weight.

[0051] In this embodiment, a weighted calculation can be performed based on the target hot data's access frequency, the inverse of the storage space occupied by the data, and the difference between the time it was generated and the current time. Specifically, the lower the access frequency, the smaller the migration weight, and the higher the migration priority. The larger the data, the more space it frees after migration, the smaller the migration weight, and the higher the migration priority. The older the data, the smaller the migration weight, and the higher the migration priority. The weights corresponding to the access frequency, the inverse of the storage space occupied by the data, and the difference between the time it was generated and the current time can be determined based on multiple experiments.

[0052] In this embodiment, the method further includes: Determine a weight corresponding to the inverse of the storage space occupied by the data based on the first difference; the first difference is the difference between the remaining capacity of the hot data layer and the preset capacity; The first difference is positively correlated with the weight corresponding to the inverse of the storage space occupied by the data.

[0053] In this embodiment, the smaller the first difference, that is, the difference between the remaining capacity of the hot data tier and the preset capacity, the tighter the capacity is. In this case, large files should be migrated first. Therefore, the weight corresponding to the inverse of the storage space occupied by the data should be increased. Therefore, the difference between the remaining capacity of the hot data tier and the preset capacity is positively correlated with the weight corresponding to the inverse of the storage space occupied by the data.

[0054] In this embodiment, the weight corresponding to the inverse of the storage space occupied by the data may be determined based on a simple linear relationship or a mapping relationship, and the slope of the linear relationship may be determined based on multiple experiments.

[0055] From the above, it can be concluded that the embodiment of the present application filters the hot data by setting the first condition, and directly sets the hot data that meets the condition to a higher migration weight, thereby avoiding the complex weight calculation of these frequently accessed data, and can quickly distinguish between the hot data that needs to be retained first and the hot data that does not need to be deeply calculated, thereby reducing the amount of calculation, improving data processing efficiency, and saving computing resources and time costs. The present application makes storage space management more efficient by adjusting the weights, and can adjust the migration strategy in time according to actual conditions, ensuring that the hot data layer always has enough space to store high-frequency access data, and ensuring the stable operation of the data storage system.

[0056] Corresponding to the above embodiment, a method for storing big data in a laboratory is described. Figure 2 This is a block diagram of a big data storage system for a laboratory provided in one embodiment of the present application. For ease of explanation, only the parts related to the embodiment of the present application are shown. Figure 2 The big data storage system 20 applied to the laboratory includes: a feature extraction module 21, a classification module 22, a storage module 23 and a storage adjustment module 24.

[0057] The feature extraction module 21 is configured to extract features from the first data in response to receiving the first data to obtain a plurality of target features; the first data is laboratory data to be stored; A classification module 22, configured to input the first data and a plurality of target features into a target classification model to obtain a classification result of the first data; A storage module 23, configured to store the first data in a corresponding storage area in a storage device based on the classification result; The storage adjustment module 24 is used to migrate the second data to its corresponding storage area for storage based on the access characteristics of the second data in the first time period in response to the access characteristics of the second data not meeting the access characteristics of the storage area where the second data is located; the second data is the data stored in the storage device.

[0058] In one embodiment of the present application, the storage area of ​​the storage device includes: a hot data layer, a warm data layer, and a cold data layer; the second data includes hot data and cold data; the hot data is the data stored in the hot data layer, and the cold data is stored in the cold data layer; A big data storage system 20 for use in a laboratory further includes: a storage migration module configured to, in response to a remaining capacity of a hot data tier being less than a preset capacity, determine a migration weight for each hot data item based on access characteristics and data characteristics of each hot data item; and migrate hot data items with a migration weight less than a first migration weight to a warm data tier; In response to a waiting time for an access request to the first cold data being longer than the first time period within a second time period, the second cold data is migrated to the warm data layer; wherein the first cold data and the second cold data are both data in the cold data layer, and a correlation between a feature of the second cold data and a feature of the first cold data is greater than the first correlation.

[0059] In one embodiment of the present application, the target classification model is a target decision tree model; a big data storage system 20 applied to a laboratory further includes: a target classification model determination module for taking the first data set as a root node; Calculating information gain of each feature in a first data set, and determining a split feature based on the information gain of each feature; the first data set includes historical data of a laboratory; Split the root node based on the split feature to obtain multiple child nodes; For each child node, calculate the information gain of each feature in the child node, and determine the splitting feature of the child node based on the information gain of each feature; Splitting the corresponding child nodes based on the splitting characteristics of each child node until the number of node samples meets the preset conditions, thereby obtaining a first decision tree; Perform pruning operation on the first decision tree to obtain the target decision tree.

[0060] In one embodiment of the present application, the target classification model determination module is specifically configured to traverse the first decision tree from the bottom node of the first decision tree as a starting point toward the top node of the first decision tree; In response to the node being a leaf node, no pruning operation is performed; In response to the node not being a leaf node, determining a first loss value of the first decision tree based on the target loss function; Pruning the node and determining a second loss value of the first decision tree based on the loss target loss function; In response to the difference between the first loss value and the second loss value being smaller than a tolerance threshold, the pruning operation on the node is canceled until all nodes of the first decision tree are traversed to obtain a target decision tree.

[0061] In one embodiment of the present application, the storage migration module is specifically configured to screen the access frequency of each hot data based on a first condition, set the migration weight of the hot data that meets the first condition as a second migration weight, and set the hot data that does not meet the first condition as the target hot data; the second migration weight is greater than the first migration weight; For each target hot data, a weighted calculation is performed on the access frequency of the target hot data, the inverse of the storage space occupied by the data, and the difference between the generation time and the current time to obtain the migration weight of the target hot data.

[0062] In one embodiment of the present application, a big data storage system 20 for a laboratory further includes: a weight adjustment module for determining a weight corresponding to the inverse of the storage space occupied by the data based on a first difference; the first difference being the difference between the remaining capacity of the hot data layer and a preset capacity; The first difference is positively correlated with the weight corresponding to the inverse of the storage space occupied by the data.

[0063] In one embodiment of the present application, a big data storage system 20 for a laboratory further includes: a data deletion module for sending a deletion request to a target device in response to the existence of target cold data in the cold data, where the target cold data is cold data whose access characteristics meet a deletion condition; In response to the deletion request being approved, the target cold data is deleted.

[0064] See also Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided in one embodiment of the present application. Figure 3 The electronic device 300 in the embodiment shown may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memory 304 is used to store computer programs, which include program instructions. The processor 301 is used to execute the program instructions stored in the memory 304. The processor 301 is configured to call the program instructions to execute the functions of the modules in the above-mentioned system embodiments, such as Figure 2 The functions of the feature extraction module 21, the classification module 22, the storage module 23 and the storage adjustment module 24 are shown.

[0065] It should be understood that in the embodiment of the present application, the processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0066] The input device 302 may include a touchpad, a fingerprint collection sensor (for collecting user fingerprint information and fingerprint direction information), a microphone, etc. The output device 303 may include a display (LCD, etc.), a speaker, etc.

[0067] The memory 304 may include a read-only memory and a random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store parameters of the target classification model, a preset capacity, and a first duration.

[0068] In a specific implementation, the processor 301, input device 302, and output device 303 described in the embodiment of the present application can execute the implementation method described in an embodiment of a big data storage method applied to a laboratory provided in the embodiment of the present application, and can also execute the implementation method of the electronic device described in the embodiment of the present application, which will not be repeated here.

[0069] In another embodiment of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, all or part of the process of the method in the above embodiment is implemented. The computer program can also be used to instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above method embodiments are implemented. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium.

[0070] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the aforementioned embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, the computer-readable storage medium can include both an internal storage unit of the electronic device and an external storage device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or is about to be output.

[0071] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0072] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the electronic devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0073] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces or units, or can be an electrical, mechanical or other form of connection.

[0074] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present application.

[0075] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0076] The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for storing big data in a laboratory, characterized in that: Controllers used in storage devices include: In response to receiving first data, performing feature extraction on the first data to obtain a plurality of target features; the first data is laboratory data to be stored; Inputting the first data and the plurality of target features into a target classification model to obtain a classification result of the first data; storing the first data in a corresponding storage area in a storage device based on the classification result; In response to the access characteristics of the second data in the first time period not meeting the access characteristics of the storage area where the second data is located, the second data is migrated to its corresponding storage area for storage based on the access characteristics of the second data; the second data is the data stored in the storage device.

2. A method for storing large amounts of data in a laboratory according to claim 1, characterized in that: The storage area of ​​the storage device includes: a hot data layer, a warm data layer and a cold data layer; the second data includes hot data and cold data; the hot data is the data stored in the hot data layer, and the cold data is stored in the cold data layer; The method further comprises: In response to the remaining capacity of the hot data layer being less than a preset capacity, determining a migration weight for each hot data based on access characteristics and data characteristics of each hot data; migrating the hot data having the migration weight less than the first migration weight to the warm data layer; In response to a waiting time for an access request to the first cold data being longer than the first time period within a second time period, the second cold data is migrated to the warm data layer; wherein the first cold data and the second cold data are both data in the cold data, and a correlation between a feature of the second cold data and a feature of the first cold data is greater than the first correlation.

3. The method for storing large amounts of data in a laboratory according to claim 1, wherein: The target classification model is a target decision tree model; The process of determining the target decision tree model includes: Take the first data set as the root node; Calculating the information gain of each feature in the first data set, and determining a split feature based on the information gain of each feature; the first data set includes historical data of the laboratory; Splitting the root node based on the splitting feature to obtain multiple child nodes; For each child node, calculate the information gain of each feature in the child node, and determine the splitting feature of the child node based on the information gain of each feature; Splitting the corresponding child nodes based on the splitting characteristics of each child node until the number of node samples meets the preset conditions, thereby obtaining a first decision tree; Perform a pruning operation on the first decision tree to obtain a target decision tree.

4. A method for storing large amounts of data in a laboratory according to claim 3, characterized in that: The pruning operation is performed on the first decision tree to obtain a target decision tree, comprising: Starting from the bottom node of the first decision tree, traverse toward the top node of the first decision tree; In response to the node being a leaf node, no pruning operation is performed; In response to the node not being a leaf node, determining a first loss value of the first decision tree based on a target loss function; Pruning the node and determining a second loss value of the first decision tree based on a loss target loss function; In response to the difference between the first loss value and the second loss value being smaller than a tolerance threshold, the operation of pruning the node is canceled until all nodes of the first decision tree are traversed to obtain the target decision tree.

5. The method for storing big data in a laboratory according to claim 2, wherein: The step of determining the migration weight of each hot data based on the access characteristics and data characteristics of each hot data includes: screening the access frequencies of the hot data based on the first condition, setting the migration weight of the hot data that meets the first condition as the second migration weight, and setting the hot data that does not meet the first condition as the target hot data; the second migration weight is greater than the first migration weight; For each target hot data, a weighted calculation is performed on the access frequency of the target hot data, the inverse of the storage space occupied by the data, and the difference between the generation time and the current time to obtain the migration weight of the target hot data.

6. A method for storing big data in a laboratory according to claim 5, characterized in that: Also includes: Determining a weight corresponding to the inverse of the storage space occupied by the data based on the first difference; The first difference is the difference between the remaining capacity of the hot data layer and the preset capacity; The first difference is positively correlated with a weight corresponding to the inverse of the storage space occupied by the data.

7. A method for storing big data in a laboratory according to claim 2, characterized in that: Also includes: In response to the existence of target cold data among the cold data, the target cold data being cold data whose access characteristics meet the deletion condition, sending a deletion request to the target device; In response to the deletion request being approved, the target cold data is deleted.

8. A big data storage system for laboratory use, characterized in that: include: a feature extraction module, configured to extract features from the first data in response to receiving the first data, to obtain a plurality of target features; the first data being laboratory data to be stored; a classification module, configured to input the first data and the plurality of target features into a target classification model to obtain a classification result of the first data; A storage module, configured to store the first data in a corresponding storage area in a storage device based on the classification result; A storage adjustment module is used to migrate the second data to its corresponding storage area for storage based on the access characteristics of the second data in a first time period in response to the access characteristics of the second data not meeting the access characteristics of the storage area where the second data is located; the second data is the data stored in the storage device.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Multi-source data processing method, system, equipment and product of humanoid robot

    CN121957481A