A laboratory data collection management method, system, equipment and medium

Through the data collection method combining RFID and multi-type sensors, combined with the improved fuzzy C-means clustering algorithm and stochastic gradient descent method, the problems of frequent manual operations and low efficiency in laboratory sample management are solved, the automatic collection and efficient management of sample data are realized, and the real-time and security of the data are improved.

CN120234324BActive Publication Date: 2025-09-16STATE GRID FUJIAN ELECTRIC POWER RES INST
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510716638.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-16
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

Traditional laboratory sample management methods involve a lot of manual operations, inflexible scheduling, and inefficient sample storage and retrieval. Especially in high-throughput experiments and drug development, it is difficult to schedule and track samples efficiently and quickly, and the existing system lacks intelligent scheduling capabilities.

Method used

RFID scanners and multiple types of sensors are used to automatically collect data, and anomaly classification is performed through the improved fuzzy C-means clustering algorithm and stochastic gradient descent method. Combined with multi-objective optimization and dynamic adjustment strategies, the full life cycle management and anomaly detection of data are achieved.

Benefits of technology

It realizes the automated collection of sample data, improves collection efficiency and storage security, reduces transmission delay and retransmission probability, improves the real-time and traceability of data, and ensures data integrity and query response time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234324B_ABST
    Figure CN120234324B_ABST
Patent Text Reader

Abstract

The present invention discloses a laboratory data acquisition management method, system, device and medium, which relates to the field of data acquisition management technology, including associating data with sample unique identifiers to generate standardized data packets with attached decision variables; multi-objective optimization balances delay and retransmission rate, dynamically adjusts batch write volume and database and table partitioning strategy; marks abnormal data, uses clustering algorithm to classify abnormalities, and adaptively adjusts detection thresholds. The method of the present invention realizes adaptive throughput adjustment to maintain data integrity through dynamic partitioning of message queues; maintains stable write latency in high-concurrency scenarios through dynamic scaling of batch write volume; shortens query response time through SSD / HDD database partitioning strategy and time partitioning combined indexing; improves abnormality classification accuracy and reduces false alarm rate through clustering algorithm; ensures data rollback accuracy in the event of database failure through transaction retry and backoff to avoid dirty data generation; and optimizes the network in real time through dynamic calculation of network status.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data acquisition management, and in particular to a laboratory data acquisition management method, system, equipment and medium. Background Art

[0002] Traditional laboratory sample management methods have problems such as many manual operations, inflexible scheduling, and low efficiency in sample storage and retrieval. These problems are particularly prominent in the fields of high-throughput experiments, drug development, clinical testing, etc., and as the number of samples increases, the difficulty of management also increases significantly. In the existing technology, although there are some sample tracking systems based on barcodes and RFID technologies, most of them lack intelligent scheduling functions and cannot effectively optimize scheduling based on laboratory needs, sample storage conditions, environmental conditions and other factors, resulting in a waste of sample storage space and retrieval efficiency. Especially when the laboratory needs to schedule a large number of samples, how to efficiently and quickly schedule, manage and track these samples while ensuring the safety of the samples remains a technical problem that needs to be solved urgently. Summary of the Invention

[0003] In view of the above-mentioned problems, the present invention is proposed.

[0004] Therefore, the technical problem solved by the present invention is: how to solve various problems in laboratory data collection and management, and improve the collection efficiency, processing capability, storage security, real-time and traceability of experimental data.

[0005] To solve the above technical problems, the present invention provides the following technical solutions: a laboratory data acquisition and management method, comprising collecting experimental data, associating the data with a unique sample identifier to generate a standardized data packet with attached decision variables; establishing a transmission delay model and a retransmission probability model, balancing the delay and retransmission rate through multi-objective optimization, and dynamically adjusting the batch write volume and the database and table partitioning strategy when writing data to the database; marking abnormal data based on preset thresholds and change rate rules, classifying anomalies using a modified fuzzy C-means clustering algorithm, and adaptively adjusting the detection threshold through the stochastic gradient descent method.

[0006] As a preferred solution of the laboratory data collection and management method described in the present invention, the collection of experimental data includes automatically collecting experimental data using RFID scanners, sensors, and IoT devices; the RFID scanners are installed at key nodes where samples enter and exit the laboratory, and the scanning range covers the entire process of sample movement; the sensors include temperature sensors, pressure sensors, pH sensors, flow meters, and spectrometers; the sensors are calibrated according to the layout of the experimental area during installation, and the temperature sensor is calibrated at three points in a constant temperature box at a standard temperature of 25°C to generate a calibration coefficient and store it in the LIMS; the pH sensor is calibrated using a standard buffer solution, and if the deviation between three consecutive measured values ​​and the standard value exceeds ±0.2, an automatic alarm is triggered and data collection is suspended; the IoT gateway is deployed at the center of the experimental area, and sensor data is periodically collected through the Modbus protocol. If the network delay exceeds 200ms, it switches to event trigger mode and uploads the data when the data change is ≥5%.

[0007] As a preferred solution of the laboratory data collection and management method described in the present invention, the data identification includes: when the data packet is generated, the system assigns a globally unique identifier to each sample and binds it to the RFID tag; the data packet metadata includes the experiment type, operator ID, and equipment status; the decision variables are attached to the data packet header; the decision variables include the ACK waiting time threshold , Maximum number of retransmissions , Number of message queue partitions ;ACK waiting time threshold Including dynamic calculation based on network status; maximum number of retransmissions Including setting the initial value and upper limit value. If two consecutive transmissions fail, Change the maximum number of retransmissions; the number of message queue partitions Including dynamic adjustment based on data inflow rate and processing rate.

[0008] As a preferred solution of the laboratory data acquisition management method described in the present invention, the establishment of a transmission delay model and a retransmission probability model includes defining a transmission delay model and designing a transmission delay objective function; constructing a retransmission probability model and designing a retransmission probability objective function; the multi-objective optimization includes establishing a multi-objective optimization problem based on data integrity verification and system load constraints, and obtaining the optimal solution for delay and retransmission rate through a non-dominated sorting genetic algorithm.

[0009] As a preferred solution of the laboratory data collection and management method described in the present invention, the dynamic adjustment of batch writing volume and database and table division strategy includes setting the batch writing volume Initial value, if write delay Greater than the delay high threshold , then according to Gradually reduce; if write delay Less than the delay low threshold , then according to Gradually increase; the database partitioning strategy is divided according to the experiment type, storing high-frequency data in the SSD database and low-frequency data in the HDD database; the table partitioning strategy is divided according to the time range, generating a new table every month, and establishing a joint index for the sample ID and timestamp; setting a transaction submission timeout If the transaction fails, then Retry; after 3 retries, the transaction is rolled back and logged.

[0010] As a preferred solution of the laboratory data collection and management method described in the present invention, the marking of abnormal data includes: LIMS performing real-time abnormal data detection on the collected data and determining temperature anomalies on the temperature sensor data; setting a normal range for the pH sensor; calculating the rate of change per unit time for any sensor data and defining a preliminary abnormality determination for the data per unit time; if any data triggers an abnormality, marking the data point as abnormal.

[0011] As a preferred solution of the laboratory data collection and management method described in the present invention, the abnormality classification includes classifying the abnormal data according to the characteristics of the abnormal data through a modified fuzzy C-means clustering algorithm, and constructing a feature vector for each abnormal data point; the improved fuzzy C-means clustering algorithm includes introducing a time continuity penalty term based on the standard fuzzy C-means clustering algorithm to construct an improved objective function; the adaptive adjustment of the detection threshold includes updating the threshold parameter using the stochastic gradient descent method; if the loss does not decrease continuously, the expert review process is triggered.

[0012] As a preferred solution of the laboratory data acquisition and management system described in the present invention, it includes: a data acquisition module, a data transmission module, and a data anomaly detection module; the data acquisition module is used to collect experimental data and associate the data with a unique sample identifier to generate a standardized data packet with an accompanying decision variable; the data transmission module is used to establish a transmission delay model and a retransmission probability model, balance the delay and retransmission rate through multi-objective optimization, and dynamically adjust the batch write volume and the library and table partitioning strategy when writing data to the database; the data anomaly detection module is used to mark abnormal data based on preset thresholds and change rate rules, use a modified fuzzy C-means clustering algorithm to classify anomalies, and adaptively adjust the detection threshold through the stochastic gradient descent method.

[0013] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement a step of a laboratory data collection and management method.

[0014] A computer-readable storage medium stores a computer program, which implements the steps of a laboratory data acquisition management method when executed by a processor.

[0015] The beneficial effects of the present invention are as follows: the laboratory data acquisition management method provided by the present invention realizes automatic data acquisition of the entire life cycle of samples from feeding to experiment through RFID full-node coverage and linkage with multiple types of sensors, eliminating manual entry errors; ensures data acquisition accuracy through sensor calibration; reduces data transmission delay in laboratory measurements and controls retransmission probability through dual-objective optimization of non-dominated sorting genetic algorithm; realizes adaptive throughput adjustment through dynamic partitioning of message queues, and can maintain data integrity in burst traffic scenarios; maintains stable write delay in high concurrency scenarios through dynamic scaling of batch write volume; shortens query response time through SSD / HDD library partitioning strategy and time partition joint indexing; introduces time penalty terms through improved fuzzy C-means clustering algorithm, improves the accuracy of anomaly classification and reduces false alarm rate; ensures data rollback accuracy in the event of database failure through transaction retry backoff and avoids the generation of dirty data; and optimizes the network in real time through dynamic calculation of network status. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 The present invention provides an overall flow chart of a laboratory sample management method according to an embodiment of the present invention.

[0018] Figure 2 A system solution module diagram of a laboratory sample management system provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0019] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.

[0020] Example 1, with reference to Figure 1 , as an embodiment of the present invention, provides a laboratory data collection and management method, comprising:

[0021] S1: Collect experimental data and associate the data with the sample unique identifier to generate a standardized data package with decision variables.

[0022] Furthermore, collecting experimental data includes automatically collecting experimental data using RFID scanners, sensors, and IoT devices.

[0023] RFID scanners are installed at key nodes where samples enter and leave the laboratory, and the scanning range covers the entire process of sample movement.

[0024] Sensors include temperature sensors, pressure sensors, pH sensors, flow meters, and spectrum analyzers.

[0025] When the sensor is installed, it is calibrated according to the layout of the experimental area. The temperature sensor is calibrated at three points in a constant temperature box at a standard temperature of 25°C to generate a calibration coefficient and store it in the LIMS laboratory information management system; the calibration coefficient , expressed as:

[0026]

[0027] in, represents the slope, represents the intercept, Indicates raw data; the pH sensor is calibrated using standard buffer. If the deviation between three consecutive measured values ​​and the standard value exceeds ±0.2, an automatic alarm is triggered and data collection is suspended.

[0028] The IoT gateway is deployed in the center of the experimental area and collects sensor data periodically through the Modbus protocol. If the network delay exceeds 200ms, it switches to event trigger mode and uploads the data when the data change is ≥5%.

[0029] It should be noted that data identification includes that when the data packet is generated, the system assigns a globally unique identifier to each sample and binds it to the RFID tag.

[0030] Data package metadata includes experiment type, operator ID, and device status.

[0031] The decision variables are appended to the packet header.

[0032] Decision variables include ACK wait time threshold , Maximum number of retransmissions , Number of message queue partitions .

[0033] ACK wait time threshold Including dynamic calculation based on network status, expressed as:

[0034]

[0035] in, is the average delay in the past 10 minutes, is the maximum delay.

[0036] Maximum number of retransmissions Including setting the initial value and upper limit value. If two consecutive transmissions fail, Change the maximum number of retransmissions.

[0037] Number of message queue partitions Including dynamic adjustment based on data inflow rate and processing rate, expressed as:

[0038]

[0039] in, Indicates the data inflow rate, Indicates the data processing rate.

[0040] S2: Establish a transmission delay model and a retransmission probability model, balance the delay and retransmission rate through multi-objective optimization, and dynamically adjust the batch write volume and database and table sharding strategy when writing data to the database.

[0041] Furthermore, establishing a transmission delay model and a retransmission probability model includes defining a transmission delay model and designing a transmission delay objective function.

[0042] The transmission delay objective function is expressed as:

[0043]

[0044] in, represents the objective function value of transmission delay, represents the delay estimation function, Represents the weight coefficient.

[0045] Construct a retransmission probability model and design the retransmission probability objective function.

[0046] Assume the retransmission probability is , then the retransmission probability objective function is expressed as:

[0047]

[0048] in, represents the objective function value of the retransmission probability, Indicates the corresponding weight.

[0049] The multi-objective optimization includes establishing a multi-objective optimization problem based on data integrity verification and system load limitation, and obtaining the optimal solution of delay and retransmission rate through a non-dominated sorting genetic algorithm.

[0050] Data integrity check, the check pass rate is satisfied, expressed as:

[0051]

[0052] The system load limit is not exceeded, which is expressed as:

[0053]

[0054] Parameter range, expressed as:

[0055] ,

[0056] Taking comprehensive consideration, a multi-objective optimization problem is established and expressed as:

[0057]

[0058] in, .

[0059] The non-dominated sorting genetic algorithm includes setting parameters as population size, crossover rate, and mutation rate, outputting the optimal solution set based on iteration, and selecting the solution with the highest comprehensive score, i.e. Minimum as the final parameter.

[0060] It should be noted that data is validated, and if pH values ​​are extremely abnormal (e.g., outside the set standard deviation of ±2), an automatic calibration is performed. Abnormal data is isolated and marked as "requiring manual review." If multiple devices or sensors report similar abnormalities, the system automatically recommends pausing the experiment and performing an equipment inspection.

[0061] It should also be noted that the dynamic adjustment of batch write volume and database and table sharding strategy includes setting the batch write volume Initial value, if write delay Greater than the delay high threshold , then according to Gradual reduction;

[0062] If write delay Less than the delay low threshold , then according to Increase gradually.

[0063] The database partitioning strategy is based on the experiment type, storing high-frequency data in the SSD database and low-frequency data in the HDD database.

[0064] The table partitioning strategy is based on the time range, generating a new table every month and establishing a joint index for the sample ID and timestamp.

[0065] Set transaction commit timeout If the transaction fails, then Try again.

[0066] After 3 retries, the transaction is rolled back and logged.

[0067] It should also be noted that data written to the database management system is synchronously and optimally transmitted via wireless or wired networks to the LIMS data server and stored in the database. Data storage utilizes standardized data formats. Data storage and database management modeling is implemented to achieve efficient database writes and data consistency, reduce write latency, and ensure transaction integrity.

[0068] Data is transmitted to the LIMS data server via wireless or wired networks and stored in a database. Data is stored in a standardized format to facilitate subsequent query, analysis, and report generation. This module ensures data security, integrity, and reliability.

[0069] In addition to single data thresholds, the rate of change of data can also serve as an important indicator. This is especially true for dynamic data collected by flow meters, thermometers, and other devices. Sudden changes in the rate of change may indicate equipment failure or environmental anomalies. For flow meter data, assuming the normal rate of change for a flow meter is 1-5 L / s, if the rate of change exceeds this range (e.g., significant fluctuations or sudden changes), the system will process the data using the following rules:

[0070] Activate pre-set rules to check for mechanical or transmission failures. Automatically adjust sampling frequency to ensure smoother data acquisition. Trigger data rollbacks. If data fluctuations persist and do not recover, trigger manual review and equipment inspection.

[0071] Objective function design, minimum write latency goal, establish write latency function:

[0072]

[0073] in, Indicates the write delay function value, represents the average write latency, Represents weight.

[0074] Data consistency goal, using data consistency verification indicators To express it, the goal is:

[0075]

[0076] in, represents the data consistency objective function value, Indicates the consistency pass rate, is the weight.

[0077] Constraints: The batch write volume B does not exceed the system processing capacity. , the transaction commit time T_trans must be lower than the system response requirement - the cache refresh interval P_cache matches the data update frequency.

[0078] The multi-objective problem constructed is:

[0079]

[0080] Where y= .

[0081] It should also be noted that LIMS systems can dynamically adjust thresholds based on real-time data feedback and historical data trends. For example, if the system detects increased volatility in a sensor's data, it can automatically adjust the sensor's threshold to avoid excessive alarms.

[0082] S3: Abnormal data is marked based on preset thresholds and change rate rules, anomalies are classified using the improved fuzzy C-means clustering algorithm, and the detection threshold is adaptively adjusted through the stochastic gradient descent method.

[0083] Furthermore, marking abnormal data includes LIMS performing real-time abnormal data detection on collected data and determining temperature anomalies on temperature sensor data.

[0084] Set the normal range to , temperature anomaly judgment, expressed as:

[0085]

[0086] in, Indicates the lower temperature threshold, Indicates the upper temperature threshold. Indicates the temperature sensor at time point The measured value of

[0087] For pH sensors, set the normal range, expressed as:

[0088]

[0089] in, Indicates the pH sensor at the time point The measured value of

[0090] For any sensor data, calculate the rate of change per unit time, expressed as:

[0091]

[0092] in, Indicates the current time point The sensor measurement value, Indicates the previous time point The measured value of

[0093] If any data triggers an anomaly, the data point is marked as an anomaly.

[0094] Set the normal speed range to (For example, for a flow meter, it can be set to [1,5L / s]). The anomaly detection function is:

[0095]

[0096] The definition of the preliminary abnormality judgment of the unit time data is expressed as:

[0097] .

[0098] It should be noted that anomaly classification includes classifying the abnormal data according to the characteristics of the abnormal data through a modified fuzzy C-means clustering algorithm and constructing a feature vector for each abnormal data point.

[0099] The feature vector of each abnormal data point is expressed as:

[0100]

[0101] Where: μ is the historical mean of the corresponding data; |Δr(t)| is the absolute value of the data change rate; d(t) represents the duration of the anomaly (for example, the number of consecutive abnormal data points).

[0102] The improved fuzzy C-means clustering algorithm includes introducing a time continuity penalty term based on the standard fuzzy C-means clustering algorithm and constructing an improved objective function.

[0103] The improved objective function is expressed as:

[0104]

[0105] in: Indicates the degree of membership of the data point at time t to the abnormal category k; represents the center of category k; m represents the fuzzy index, usually m>1; λ represents the weight of the time penalty term; represents the time continuity penalty function, which is designed as:

[0106]

[0107] This is used to ensure that the ownership relationship of data at adjacent moments does not fluctuate dramatically.

[0108] Adaptive adjustment of the detection threshold includes updating the threshold parameters using a stochastic gradient descent method;

[0109] If the losses do not continue to decrease, the expert review process will be triggered.

[0110] It should also be noted that the abnormal data are classified into three categories using the improved fuzzy C-means clustering algorithm:

[0111] Abnormal mutation: large fluctuations in a short period of time, usually accompanied by random extremely high |Δr(t)|.

[0112] Progressive drift: Data deviates continuously from normal values, showing the appearance of long-term cumulative deviation.

[0113] Single point anomaly: An isolated error, usually a momentary deviation from normal.

[0114] It should also be noted that in actual operation, the system will count false positives (incorrectly marking normal data as anomalies) and false negatives (failing to detect real anomalies), and define a loss function to measure the impact of these errors. This loss function is based on the number of false positives and false negatives, and sets weights to balance the impact of the two.

[0115] Gradient descent updates parameters, automatically adjusting preset threshold parameters like temperature, pH, and rate based on the gradient of the loss function. In other words, if the system finds that the current parameter settings are causing too many false positives or false negatives, it calculates the impact of each parameter on the loss function and adjusts these thresholds appropriately until the system achieves optimal detection performance. The process and results of each parameter adjustment are recorded, forming a "regular feedback loop."

[0116] Through the above data analysis and processing, when the system detects data anomalies, it will make the following decisions: Alarm system: Triggers an alarm message to laboratory staff. Automatic adjustment system: If conditions permit, automatically adjusts experimental environment parameters or equipment status, such as adjusting temperature and pressure. Manual intervention: When the system determines that the abnormal situation is complex, it notifies manual intervention.

[0117] It should be noted that the system features user management and permission control to ensure strict management of laboratory personnel's access, editing, and deletion rights to experimental data. Only authorized personnel can perform specific operations, ensuring data security and compliance. The system integrates with the laboratory's automated equipment and environmental monitoring systems to monitor equipment operating status and experimental environment parameters (such as temperature, humidity, and air pressure) in real time, and promptly report any anomalies. This module helps laboratory managers promptly detect equipment failures or environmental anomalies, reducing experimental errors.

[0118] Example 2 is an embodiment of the present invention, which provides a laboratory data acquisition management method. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0119] First, the laboratory implemented a comprehensive automated data acquisition system to collect, identify, transmit, store, analyze, and detect anomalies in real time for all key experimental data. This system also enabled traceability and feedback adjustments for anomalous data. During the experimental preparation phase, RFID scanners were installed at the sample entry and exit areas to ensure that each sample was uniquely identified upon entering the laboratory. Furthermore, temperature sensors, pressure sensors, pH sensors, flow meters, and optical spectrum analyzers were installed in the experimental area, and all sensors were rigorously calibrated to ensure their output data was within standard ranges. An IoT gateway was deployed within the environmental monitoring system to centrally transmit sensor data to the data acquisition terminal. The acquisition module's software system utilizes an embedded control program to acquire real-time data using fixed-cycle and event-triggered modes. It then matches RFID scan data with sensor data, generating standardized data packets with timestamps, device IDs, sample IDs, and other metadata. Each data packet also includes decision variable parameters such as the ACK wait time threshold (T_ack), the maximum number of retransmissions (R_max), the number of partitions in the message queue (N_p), the data inflow rate (λ), and the message processing rate (μ). Next, the data is transmitted in real time to the LIMS system's data server through standardized APIs and middleware (such as Kafka or RabbitMQ), and data transmission is optimized using transmission delay models and retransmission probability models. During the transmission process, the system uses preset weights to minimize the delay from the data collection point to the target server, while reducing the retransmission probability through multi-objective optimization and ensuring that the data verification pass rate meets the set threshold. The uploaded data will enter the database management module via a wireless or wired network. The system uses batch writing, transaction management, and cache refresh strategies for data storage to ensure low write latency and data consistency that meets predetermined requirements. During the writing process, the database management component dynamically adjusts parameters such as batch writing volume, transaction submission timeout, cache refresh interval, and the number of shards / tables to achieve a balance between low latency and high consistency. After data storage is completed, the LIMS system immediately analyzes and processes the data in real time. The system first performs preliminary anomaly detection on various data points, such as temperature, pH, and flow rate, based on preset normal ranges. For example, if the temperature is set to 20°C to 30°C, the pH is set to 6 to 8, and the flow rate change rate is set to 1 to 5 L / s, any data outside these ranges will be automatically marked as an anomaly. Subsequently, for data marked as anomaly, the system constructs a feature vector, taking into account the data's deviation from the historical mean, the rate of change per unit time, and the duration of the anomaly. It then uses a modified fuzzy C-means clustering algorithm (which introduces a time continuity penalty based on traditional clustering) to classify the anomalies into three categories: sudden anomalies, gradual shifts, and single-point anomalies. The system then ranks the severity of each anomaly.Finally, the system adaptively adjusts the threshold parameters through the gradient descent method according to the false alarms and missed alarms, forming a regular feedback loop, thereby continuously optimizing the entire data collection and anomaly detection process.

[0120] Table 1 Test data record table

[0121]

[0122] As can be seen in Table 1, various decision variable parameters and database write parameters were finely tuned under different experimental conditions to demonstrate the significant advantages of the present invention's solution in optimizing data transmission and writes. First, looking at the T_ack (ACK wait time threshold) data, the settings for each experimental condition ranged from 180 milliseconds to 210 milliseconds, significantly lower than the typical 250 milliseconds of conventional systems. This demonstrates that the present invention optimizes data confirmation response speed and can more quickly identify transmission anomalies.

[0123] Secondly, regarding R_max (maximum number of retransmissions), this invention typically sets it to 2 to 4 times. Compared to conventional systems with higher retransmission times, this invention reduces the retransmission rate while effectively ensuring stable and real-time data transmission. Looking at the number of partitions N_p in the message queue and the data inflow rate λ, the experimental data is distributed between 4 and 5 partitions and 50 to 60 data points per second. This demonstrates that the system is able to properly distribute and process data, ensuring that even in high data traffic conditions, latency is not increased due to overloading of a single processing node.

[0124] Example 3 is the third embodiment of the present invention, which differs from the first two embodiments in that:

[0125] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0126] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0127] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, and then editing, interpreting, or processing in another suitable manner as necessary, and then storing it in a computer memory.

[0128] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the aforementioned embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or combination of the following technologies known in the art may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.

[0129] Example 4, with reference to Figure 2 , which is the fourth embodiment of the present invention, provides a laboratory sample management system, including a data acquisition module, a data transmission module, and a data anomaly detection module.

[0130] Among them, the data acquisition module is used to collect experimental data, and associate the data with the sample unique identifier to generate a standardized data packet with decision variables; the data transmission module is used to establish a transmission delay model and a retransmission probability model, balance the delay and retransmission rate through multi-objective optimization, and dynamically adjust the batch write volume and database and table partitioning strategy when writing data to the database; the data anomaly detection module is used to mark abnormal data based on preset thresholds and change rate rules, use the improved fuzzy C-means clustering algorithm to classify anomalies, and adaptively adjust the detection threshold through the stochastic gradient descent method.

[0131] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A laboratory data collection and management method, characterized in that: include: Collect experimental data and associate the data with the sample unique identifier to generate a standardized data package with decision variables; Establish a transmission delay model and a retransmission probability model, balance delay and retransmission rate through multi-objective optimization, and dynamically adjust the batch write volume and database and table sharding strategy when writing data to the database; Abnormal data is marked based on preset thresholds and change rate rules, and anomalies are classified using the improved fuzzy C-means clustering algorithm. The detection threshold is adaptively adjusted through the stochastic gradient descent method. The associating the data with the sample unique identifier includes, when the data packet is generated, the system assigns a globally unique identifier to each sample and binds it to the RFID tag; Data packet metadata includes experiment type, operator ID, and device status; The decision variable is attached to the data packet header; Decision variables include ACK wait time threshold , Maximum number of retransmissions , Number of message queue partitions ; ACK wait time threshold Including dynamic calculation based on network status, expressed as: ,in, is the average delay in the past 10 minutes, is the maximum delay; Maximum number of retransmissions Including setting the initial value and upper limit value. If two consecutive transmissions fail, Change the maximum number of retransmissions; Number of message queue partitions Including dynamic adjustment based on data inflow rate and processing rate, expressed as: ,in, Indicates the data inflow rate, Indicates the data processing rate; The establishment of the transmission delay model and the retransmission probability model includes defining the transmission delay model and designing the transmission delay objective function; The transmission delay objective function is expressed as: ,in, represents the delay estimation function, represents the weight coefficient; Construct a retransmission probability model and design a retransmission probability objective function; Assume the retransmission probability is , then the retransmission probability objective function is expressed as: ,in, Indicates the corresponding weight; The multi-objective optimization includes establishing a multi-objective optimization problem based on data integrity verification and system load limitation, and obtaining the optimal solution of delay and retransmission rate through a non-dominated sorting genetic algorithm.

2. The laboratory data collection and management method according to claim 1, wherein: The collecting of experimental data includes automatically collecting experimental data using RFID scanners, sensors, and IoT devices; RFID scanners are installed at key points where samples enter and leave the laboratory, and the scanning range covers the entire sample movement process; The sensors include temperature sensor, pressure sensor, pH sensor, flow meter, and spectrum analyzer; When the sensor is installed, it is calibrated according to the layout of the experimental area. The temperature sensor is calibrated at three points in a constant temperature box at a standard temperature of 25°C to generate the calibration coefficient and store it in the LIMS. The pH sensor is calibrated using a standard buffer solution. If the deviation between the measured value and the standard value exceeds ±0.2 for three consecutive times, an automatic alarm is triggered and data collection is suspended; The IoT gateway is deployed in the center of the experimental area and collects sensor data periodically through the Modbus protocol. If the network delay exceeds 200ms, it switches to event trigger mode and uploads the data when the data change is ≥5%.

3. The laboratory data collection and management method according to claim 2, wherein: The dynamic adjustment of batch writing volume and database and table sharding strategy includes setting the batch writing volume. Initial value, if write delay Greater than the delay high threshold , then according to Gradual reduction; If write delay Less than the delay low threshold , then according to gradually increase; The database partitioning strategy is based on the experiment type, storing high-frequency data in the SSD database and low-frequency data in the HDD database; The table partitioning strategy is based on time ranges, with new tables generated monthly and joint indexes created for sample IDs and timestamps. Set transaction commit timeout If the transaction fails, then Retry; After 3 retries, the transaction is rolled back and logged.

4. The laboratory data collection and management method according to claim 3, wherein: The marking of abnormal data includes LIMS performing real-time abnormal data detection on the collected data and determining temperature anomaly on the temperature sensor data; For pH sensors, set the normal range; For any sensor data, calculate the rate of change per unit time and define the preliminary abnormality judgment of the data per unit time; If any data triggers an anomaly, the data point is marked as an anomaly.

5. The laboratory data collection and management method according to claim 4, wherein: The abnormality classification includes classifying the abnormal data according to the characteristics of the abnormal data by using a modified fuzzy C-means clustering algorithm and constructing a feature vector for each abnormal data point; The improved fuzzy C-means clustering algorithm includes introducing a time continuity penalty term based on the standard fuzzy C-means clustering algorithm to construct an improved objective function; The adaptive adjustment of the detection threshold comprises updating the threshold parameters using a stochastic gradient descent method; If the losses do not continue to decrease, the expert review process will be triggered.

6. A system using the laboratory data collection and management method according to any one of claims 1 to 5, characterized in that: Including data acquisition module, data transmission module, and data anomaly detection module; The data acquisition module is used to collect experimental data and associate the data with the sample unique identifier to generate a standardized data package with decision variables; The data transmission module is used to establish a transmission delay model and a retransmission probability model, balance the delay and retransmission rate through multi-objective optimization, and dynamically adjust the batch writing volume and database and table sharding strategy when writing data to the database; The data anomaly detection module is used to mark abnormal data based on preset thresholds and change rate rules, classify anomalies using a modified fuzzy C-means clustering algorithm, and adaptively adjust the detection threshold through a stochastic gradient descent method.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the laboratory data collection and management method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the laboratory data collection and management method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Fault automatic detection and repair method for self-healing intelligent power line

    CN118739184A

  • Power distribution network fault positioning method based on artificial intelligence and storage medium

    CN118884129A

  • LSM engine data cache optimization method

    CN119829626A