Intelligent sample sampling management system and method based on Java language
Through an intelligent sample sampling management system based on Java language, the sampling frequency is dynamically adjusted using the random forest algorithm and communication API, the efficiency and accuracy of sample sampling management in the existing technology are solved, and fast response and efficient sample management are achieved.
Patent Information
- Application Number
- CN202510308594.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art relies on statistical methods and manual operations in sample sampling management, resulting in low efficiency in rapid response and data processing, difficult to adapt to changing production conditions, affecting the accuracy and timeliness of decision-making, and manual recording and reporting are prone to errors, reducing the reliability and stability of sampling management.
The intelligent sample sampling management system based on the Java language is adopted, including anomaly detection module, sampling optimization module, notification management module and sampling analysis module. The random forest algorithm is used to identify abnormal data points, dynamically adjust the sampling frequency, real-time notification is realized through Java concurrency package and communication API, and combined with time series prediction and pattern recognition, the sampling process is optimized.
It improves the speed of abnormal state recognition, reduces the risk of error sampling, ensures that the sampling frequency is synchronized with the production environment, improves resource utilization efficiency and decision-making transparency, and enhances the accuracy and efficiency of the sampling management system.
Smart Images

Figure CN120257002A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sample sampling, and particularly to an intelligent sample sampling management system and method based on the Java language. Background Art
[0002] Sample sampling technology is a methodology widely used in multiple fields such as data collection, quality control, and production process management. It is particularly crucial in industries such as statistics, biology, pharmacy, materials science, and food safety. This technology mainly involves the process of extracting representative samples from a large number of samples or data sets for more in-depth analysis or testing. The correct sample sampling method can significantly improve the efficiency and accuracy of analysis and ensure the representativeness and reliability of the results.
[0003] Among them, an intelligent sample sampling management system based on the Java language refers to using the Java programming language to provide an intelligent interface and tools for managing and automating the sample sampling process. Its main uses include the generation of automated sampling plans, real-time data collection and analysis, monitoring and detection of sample status, and automated generation of reports. This system manages human behavior according to specifications, can greatly improve the accuracy and efficiency of sample sampling management, reduce human errors, ensure that all operations comply with industry standards, and at the same time provide stable and reliable data to support decision-making.
[0004] Traditional technologies mainly rely on statistical methods and manual operations, and perform poorly in terms of rapid response and data processing efficiency. Manual sampling is prone to errors when dealing with large-scale data sets, and the efficiency is relatively low. It is not sensitive to changing production conditions, resulting in sampling decisions often being based on outdated or incomplete data, directly affecting the accuracy and timeliness of decisions. Errors in the process of manual recording and report generation also increase the uncertainty in data processing, limiting the reliability and stability of sampling management, and thus affecting the efficiency and safety of the overall sampling operation. Summary of the Invention
[0005] The purpose of the present invention is to solve the disadvantages existing in the prior art, and to propose an intelligent sample sampling management system and method based on the Java language.
[0006] To achieve the above purpose, the present invention adopts the following technical solution: An intelligent sample sampling management system based on the Java language, the system includes:
[0007] The anomaly detection module continuously receives sample data based on the data stream interface, analyzes the status of data points using the random forest algorithm, identifies data points that deviate from the normal pattern, and judges the abnormal status of the sample to obtain an anomaly data identification mark;
[0008] Based on the abnormal data identification mark, the sampling optimization module calls the Java concurrent package, analyzes the abnormal degree and occurrence frequency of data points, calculates the number of sampling points that need to be adjusted, and dynamically optimizes the sampling frequency of the target batch to obtain a sampling adjustment configuration;
[0009] Based on the sampling adjustment configuration, the notification management module uses the Java network and communication API to send abnormal and sampling adjustment notifications in real time, monitors the sending status of the notifications, and records the associated sampling operation logs simultaneously to obtain detailed operation log details;
[0010] Based on the detailed operation log details, the sampling analysis module combines historical sampling data, uses JTimeSeries or Deeplearning4j to predict the time series of sample data, identifies sampling patterns, and analyzes the correlation between sample characteristics and sampling effects to determine the optimal sampling process and timing, obtaining a sampling optimization record.
[0011] The improvement of the present invention is that the step of analyzing the status of data points is specifically as follows:
[0012] Based on the data stream interface, continuously receive sample data, set up a data buffer, and store the instant data stream to obtain a data stream set;
[0013] Apply the random forest algorithm to analyze the data stream set, calculate the status index of each data point, using the formula:
[0014]
[0015] Obtain a set of status scores, where S i represents the status score of data point i, w j is the weight of feature j, a ij is the current monitored value of data point i on feature j, and m is the total number of features;
[0016] According to the set of status scores, filter out the data points with scores lower than the threshold, and identify the data points that deviate from the normal pattern to obtain a set of data points with a deviation pattern.
[0017] The improvement of the present invention is that the step of obtaining the abnormal data identification mark is specifically as follows:
[0018] Mark the data points that deviate from the normal pattern, assign a unique non-conventional behavior mark to them, and obtain the screened abnormal data;
[0019] Based on the processed abnormal data, judge the abnormal status of the sample, and associate the mark and abnormal status of the data point to obtain an abnormal data identification mark.
[0020] The improvements of the present invention are as follows. The analysis steps for the degree and frequency of anomalies are specifically as follows:
[0021] Based on the anomaly data identification markers, collect the number of anomalies and timestamps for each data point, and use the Java concurrent package to implement parallel processing. Integrate the anomaly events and corresponding times of the data points to obtain a preliminary anomaly data set.
[0022] Based on the preliminary anomaly data set, use the formula:
[0023]
[0024] Calculate the degree of anomaly for each data point to obtain the analysis results of the degree and frequency of anomalies. Among them, ES i represents the degree of anomaly of data point i, nS i represents the number of anomalies of data point i within the monitoring period, TS is the total time of the monitoring period, and fS i is the occurrence frequency of data point i during the monitoring period.
[0025] The improvements of the present invention are as follows. The steps for obtaining the sampling adjustment configuration are specifically as follows:
[0026] Based on the degree and frequency of anomalies, determine the number of data points that need to increase the sampling frequency, and set priorities according to the level of anomaly degree to obtain a sampling adjustment target list.
[0027] Based on the sampling adjustment target list, use the formula:
[0028] FS′ i = FS0 + α X ·ES i ;
[0029] Perform sampling frequency adjustment to obtain a new sampling frequency. Among them, FS′ i represents the adjusted sampling frequency of data point i, FS0 is the baseline sampling frequency, ES i is the degree of anomaly of data point i, and α X is the adjustment coefficient.
[0030] Apply the new sampling frequency to the target batch and adjust the associated sampling settings to optimize the efficiency and accuracy of sampling data collection to obtain the sampling adjustment configuration.
[0031] The improvements of the present invention are as follows. The steps for real-time sending of anomaly and sampling adjustment notifications are specifically as follows:
[0032] Based on the sampling adjustment configuration, configure the Java network and communication API, and formulate trigger rules for different types of notifications according to the degree of anomaly and sampling data to obtain a notification trigger rule set.
[0033] Based on the notification triggering rule set, the formula is adopted:
[0034]
[0035] Dynamically adjust the order and frequency of notification sending, obtain the notification sending queue and status record, among which NV h is the priority of the hth notification, DV h is the data trigger level, W is the weight set according to the urgency of the notification, and SV is the current system load;
[0036] The notification sending queue and status record are analyzed to verify the sending efficiency and success rate of the notification, and the notifications that are not successfully sent are reordered and resent to obtain the notification sending record.
[0037] The present invention is improved in that the steps of obtaining the operation log details are specifically as follows:
[0038] Capture activities associated with sampling adjustment, record the request and response of each sampling adjustment, including timestamp, operation type and status, and obtain a preliminary log data set;
[0039] The preliminary log data set is processed to remove invalid or duplicate entries, and the logs are classified and sorted to optimize the readability and usability of the log data to obtain operation log details.
[0040] The present invention is improved in that the steps of obtaining the sampling optimization record are specifically as follows:
[0041] Based on the operation log details, extract the time and parameters of the sampling adjustment, perform time series analysis through JTimeSeries, identify the periodic patterns and abnormal changes of the data, and obtain the time series analysis results;
[0042] Based on the time series analysis results, the formula is adopted:
[0043]
[0044] Predict the sampling process and timing to obtain the predicted sampling optimization value PU i , among which, XU ij is the value of the jth independent variable in the i-th data, β j is the jth independent variable XU ij The regression coefficient, Z i is the adjustment factor, e is the base of the natural logarithm, β0 is the intercept or baseline coefficient, and n U is the total number of independent variables;
[0045] Based on the predicted sampling optimization value, analyze the correlation between sample characteristics and sampling effect, determine the optimal sampling process and timing, and obtain the sampling optimization record.
[0046] An intelligent sample sampling management method based on the Java language, the intelligent sample sampling management method based on the Java language is executed based on the above-mentioned intelligent sample sampling management system based on the Java language, and includes the following steps:
[0047] S1: Continuously receive sample data based on the data stream interface, use the random forest algorithm to analyze the status of data points, identify data points deviating from the normal mode, and judge the abnormal status of the sample to obtain an abnormal data identification mark;
[0048] S2: Based on the abnormal data identification mark, call the Java concurrent package to analyze the degree and frequency of data point abnormalities, calculate the number of sampling points that need to be adjusted, and dynamically optimize the sampling frequency for the target batch to obtain a sampling adjustment configuration;
[0049] S3: Based on the sampling adjustment configuration, use the Java network and communication API to send abnormal and sampling adjustment notifications in real time, monitor the sending status of the notifications, and record the associated sampling operation logs at the same time to obtain the operation log details;
[0050] S4: Based on the operation log details, combined with historical sampling data, use JTimeSeries or Deeplearning4j to predict the time series of sample data, analyze the correlation between sample characteristics and sampling effect, and obtain the sampling mode analysis result;
[0051] S5: Based on the sampling mode analysis result, evaluate the effects of different sampling modes, and determine the optimal sampling time point and process according to the sample characteristics and data accuracy requirements to obtain the sampling optimization record.
[0052] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0053] In the present invention, by using the random forest algorithm to immediately identify data deviations, the recognition speed of abnormal states is improved, the risk of incorrect sampling is reduced, and dynamic adjustment is performed through real-time analysis results to ensure that the sampling frequency and efficiency are always synchronized with the current production environment. This dynamic adjustment mechanism effectively improves the resource utilization efficiency. By integrating the communication API to automate the notification process, the transparency and response speed of decision support are enhanced, and time series prediction and pattern recognition make the sampling process more accurate, thereby improving the decision-making quality and efficiency of the sampling management system. Description of the Drawings
[0054] Figure 1This is the module diagram of the intelligent sample sampling management system based on the Java language proposed by the present invention;
[0055] Figure 2 This is the flowchart for analyzing the status of data points in the present invention;
[0056] Figure 3 This is the flowchart for obtaining the identification mark of abnormal data in the present invention;
[0057] Figure 4 This is the flowchart for analyzing the degree of abnormality and the frequency of occurrence in the present invention;
[0058] Figure 5 This is the flowchart for obtaining the sampling adjustment configuration in the present invention;
[0059] Figure 6 This is the flowchart for real-time sending of abnormal and sampling adjustment notifications in the present invention;
[0060] Figure 7 This is the flowchart for obtaining the details of the operation log in the present invention;
[0061] Figure 8 This is the flowchart for obtaining the sampling optimization record in the present invention. Detailed implementation manners
[0062] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0063] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "a plurality of" is two or more unless otherwise specifically defined.
[0064] Embodiment
[0065] Please refer to Figure 1 , the present invention provides a technical solution: An intelligent sample sampling management system based on the Java language includes:
[0066] The anomaly detection module continuously receives sample data based on the data stream interface, analyzes the status of data points using the random forest algorithm, identifies data points that deviate from the normal pattern, and determines the anomaly status of the samples to obtain an anomaly data identification mark;
[0067] Based on the anomaly data identification mark, the sampling optimization module calls the Java concurrent package java.util.concurrent, analyzes the anomaly degree and occurrence frequency of data points, calculates the number of sampling points that need to be adjusted, and dynamically optimizes the sampling frequency for the target batch to obtain a sampling adjustment configuration;
[0068] Based on the sampling adjustment configuration, the notification management module uses Java network and communication APIs, including JavaMail API and WebSocket, to send anomaly and sampling adjustment notifications in real time, monitors the sending status of the notifications, and records the associated sampling operation logs to obtain operation log details;
[0069] The sampling inspection module assigns inspection tasks, generates a material requisition form according to the inspection items in the inspection task. The inspection personnel receive consumables in the consumables warehouse according to the material requisition form. After the consumables are used, the QR code on the consumables is scanned to cancel the consumption of the consumables by quantity. The cancellation data is automatically uploaded to the system to facilitate the statistics of inventory information;
[0070] The experimental personnel assigned to the inspection task enter the laboratory through face recognition. The face recognition information is saved in the system, and the experimental personnel are authorized;
[0071] The inspection personnel scan the inspection samples, record each link of the inspection and associate the current inspection equipment to form an instrument usage record. The system is connected to all equipment in the laboratory. The system automatically collects information such as the weighing information, constant volume information, and liquid concentration information collected by the equipment, and finally determines the content or value of the substance to be measured in the inspection sample;
[0072] Based on the operation log details, the sampling analysis module combines historical sampling data, uses JTimeSeries or Deeplearning4j to predict the time series of sample data, identifies sampling patterns, and analyzes the correlation between sample characteristics and sampling effects to determine the optimal sampling process and timing to obtain a sampling optimization record;
[0073] The report approval module compares the current inspection data with the technical standard requirement data in the product standard. If it exceeds the technical standard requirement in the product standard, the sample is a suspected unqualified sample. The system automatically assigns a retest task. Samples that are still unqualified after retesting are registered as unqualified samples. At the same time, the system automatically uploads the report number and backs up the report.
[0074] The abnormal data identification markers include deviation level, abnormal point index, and deviation type. The sampling adjustment configuration specifically includes frequency adjustment value, efficiency gain indicator, and target batch identifier. The operation log details include notification effect evaluation results, log storage location, and log access time. The sampling optimization record specifically refers to the sampling adjustment basis, implementation timing, and effect feedback analysis results.
[0075] Please refer to Figure 2 , and the steps for analyzing the status of data points are specifically as follows:
[0076] Based on the data stream interface, continuously receive sample data, set up a data buffer, and store the instantaneous data stream to obtain a data stream set;
[0077] The data stream interface obtains continuous data from sensors, monitoring devices, or third-party data sources. The data buffer is designed as a first-in-first-out (FIFO) structure, which can gradually cache and temporarily store the incoming data, avoiding data loss or delay caused by data stream overflow. By monitoring the update frequency of the cache and the effective timestamp of the data stream, the integrity and consistency of data transmission are monitored. When storing the instantaneous data stream, preliminary marking and filtering processing are performed on duplicate, incorrect, or damaged data, while retaining the time series information for subsequent data analysis. After the above operations, a data stream set with accurate timestamps and no data loss is obtained.
[0078] Apply the random forest algorithm to analyze the data stream set, calculate the status indicators of each data point, using the formula:
[0079]
[0080] Obtain the status score set, where S i represents the status score of data point i, which is used to subsequently determine whether the data point deviates from the normal mode. w j is the weight of feature j, and a ij is the current monitored value of data point i on feature j, which is the original value directly obtained from the data. m is the total number of features;
[0081] There are three data points, each data point has m = 3 features, the feature weights are w1 = 0.5, w2 = 0.3, w3 = 0.2, and the observed values of each data point on these three features are respectively:
[0082] a 11 = 10, a 12 = 20, a 13 = 30;
[0083] a 21 = 15, a 22 = 25, a 23 = 35;
[0084] a 31 = 20, a 32 = 30, a 33 = 40;
[0085] Calculate the status score S for each data point i :
[0086] For data point 1:
[0087] S1 = (0.5 × 10) + (0.3 × 20) + (0.2 × 30) = 5 + 6 + 6 = 17;
[0088] For data point 2:
[0089] S2 = (0.5 × 15) + (0.3 × 25) + (0.2 × 35) = 7.5 + 7.5 + 7 = 22;
[0090] For data point 3:
[0091] S3 = (0.5 × 20) + (0.3 × 30) + (0.2 × 40) = 10 + 9 + 8 = 27;
[0092] The results show that the comprehensive status scores of data points 1, 2, and 3 are 17, 22, and 27 respectively. The scores will be used to compare with a set threshold to determine whether each data point deviates from the normal pattern. If S i is lower than the threshold, it indicates that there is an abnormality in the data point.
[0093] According to the set of status scores, filter out the data points with scores lower than the threshold and identify the data points that deviate from the normal pattern to obtain a set of data points in the deviation pattern;
[0094] First, set a threshold, which can be determined according to the statistical distribution or quantiles of historical data. Select the 90% or 95% quantile in the score set as the threshold, and then compare the status scores point by point. If the status score of a data point is lower than the threshold, it is determined as a data point deviating from the normal pattern. For the data points that meet the conditions, retain their position information and original values in the data stream, mark them as abnormal data points, and at the same time establish a set of data points in the deviation pattern. The set contains the status scores, timestamps, observed values of each feature, and associated historical status scores of the data points.
[0095] Please refer to Figure 3 , the specific steps for obtaining the abnormal data identification mark are as follows:
[0096] Mark the data points that deviate from the normal pattern, assign a unique non - conventional behavior mark to them, and obtain the filtered abnormal data;
[0097] Associate the unique identifier of the data point (such as the data point ID) with its status score to ensure that each data point can be individually identified. Then, assign an anomaly flag to each data point whose status score is lower than the threshold. This process relies on a preset flag rule library, which includes anomaly level classification, flag encoding format, etc. The anomaly flag uses a standardized format numbering, for example, using a two - part encoding (anomaly category code + data point number) for identification. For example, it is marked as "E01 - 001", where E01 represents the anomaly category and 001 is the unique number of the data point. Record the flagged data point and its corresponding anomaly status score, and store it in the anomaly data record table to ensure that all anomaly data can be quickly located and traced during subsequent processing, and finally generate a processed anomaly data set.
[0098] Based on the processed anomaly data, judge the anomaly status of the sample, and associate the flag of the data point with the anomaly status to obtain an anomaly data identification flag;
[0099] By classifying and counting the flagged anomaly data set, extract the anomaly category flag of the data point and the corresponding status score. Compare the status score with the threshold range, and judge the anomaly status through classification rules. The classification criteria include parameters such as the difference in status score, anomaly frequency, and anomaly distribution range. For example, a flag with a large difference between the status score and the threshold is marked as "high risk", and a flag with a small difference is marked as "low risk". Associate the anomaly category of each data point with the flag, and finally form a complete anomaly data identification flag table.
[0100] Please refer to Figure 4 , and the analysis steps for the anomaly degree and occurrence frequency are specifically as follows:
[0101] Based on the anomaly data identification flag, collect the number of anomalies and timestamps of each data point, and use the Java concurrent package java.util.concurrent to achieve parallel processing. Integrate the anomaly events and corresponding times of the data points to obtain a preliminary anomaly data set;
[0102] Associate the data point identifier with its anomaly events, obtain the anomaly occurrence time and number of times by querying the record database, map each anomaly event of the data point to a timeline, and mark the timestamp of the data point with millisecond precision. Then, call the multithreaded framework in the Java concurrent package java.util.concurrent, and perform parallel calculations on the data point anomaly events in a segmented manner. Each thread is responsible for aggregating the anomaly data within a certain period of time, and after the thread is completed, the results are merged to ensure that the data processing process does not overlap. Finally, integrate all the anomaly events of the data points and the corresponding time records to form a preliminary anomaly data set.
[0103] Based on the preliminary abnormal data set, the formula is used:
[0104]
[0105] Calculate the abnormality degree of each data point to obtain the analysis results of abnormality degree and frequency. Among them, ES i represents the abnormality degree of data point i, which combines the number of abnormal occurrences and the occurrence frequency, and is used to evaluate the abnormal nature of the data point relative to the entire observation period. nS i represents the number of abnormal occurrences of data point i during the monitoring period. The value is the number of events identified by analysis that deviate from the normal behavior. TS is the total time of the monitoring period, indicating the entire time period for data collection. fS i is the occurrence frequency of data point i during the monitoring period, which is the ratio of the number of times the data point appears to the total number of records, reflecting the activity degree of the data point in the entire data set;
[0106] During a certain monitoring period, the number of abnormal occurrences nS of data point i i = 5, the total time TS of the monitoring period = 24 hours (one day), and the occurrence frequency fS of data point i i = 0.1 (the number of times the data point appears accounts for 10% of the total number of times). Substitute it into the formula for calculation:
[0107]
[0108] This result indicates that the abnormality degree of data point i during the monitoring period is 0.0208, and the value is relatively low, indicating that although data point i has abnormal records, its abnormality degree and frequency in the overall data are not very high, which helps to determine which data points require further attention and analysis.
[0109] Please refer to Figure 5 , and the specific steps for obtaining the sampling adjustment configuration are as follows:
[0110] Based on the abnormality degree and occurrence frequency, determine the number of data points that need to increase the sampling frequency, and set priorities according to the high and low abnormality degrees to obtain the sampling adjustment target list;
[0111] By statistically calculating the abnormality degree scores and occurrence frequencies of each data point during the monitoring period, calculate the number of data points whose abnormality degree exceeds the threshold, and sort the data points according to the high and low abnormality degrees. This includes comparing the numerical values of the abnormality degrees, setting the priority of the sorting rules, combining the normalization results of the number of abnormal occurrences and the occurrence frequencies of the data points during the monitoring period, screening and marking the data points one by one, and generating a sampling adjustment target list containing the data point ID, abnormality degree value, and occurrence frequency in descending order of the abnormality degree. This list will be used as the input data for sampling optimization for subsequent use.
[0112] Adjust the target list based on sampling, using the formula:
[0113] FS′ i = FS0 + α X ·ES i ;
[0114] Perform sampling frequency adjustment to obtain a new sampling frequency. Among them, FS′ i represents the adjusted sampling frequency of data point i, which is used to determine the frequency at which the data point should be sampled in the updated sampling plan. FS0 is the baseline sampling frequency, that is, the sampling frequency defaulted by the system before considering the adjustment of abnormal data. ES i is the abnormality degree of data point i, and the value reflects the deviation degree of the data point from the normal behavior, which is used to determine the weight of the data point in the new sampling configuration. α X is the adjustment coefficient, which is used to adjust the influence of the abnormality degree ES on the sampling frequency;
[0115] The abnormality degree ES of a data point i i = 0.2, the baseline sampling frequency FS0 = 1.0 times per hour, and the adjustment coefficient α X = 5. Substitute into the formula for calculation:
[0116] FS′ i = 1.0 + 5·0.2 = 2;
[0117] This result indicates that the new sampling frequency of data point i is 2 times per hour, which is an adjustment based on the abnormality degree, aiming to increase the monitoring frequency of data points with more frequent abnormal behaviors to ensure that key data will not be missed.
[0118] Apply the new sampling frequency to the target batch and adjust the associated sampling settings to optimize the sampling data collection efficiency and accuracy, obtaining the sampling adjustment configuration;
[0119] First, call the sampling adjustment target list. According to the data point ID in the list and the corresponding new sampling frequency assignment value, match the sampling frequency with the operation parameters of the current batch, and update the sampling configuration of each data point in the monitoring system one by one, including reallocating the sampling trigger time, sampling interval, and batch sampling point number. Then, synchronize the new configuration parameters to the sampling control through the interface, adjust the relevant data storage and recording settings to ensure that the sampling trigger and data transmission are consistent with the new configuration, and finally complete the sampling frequency optimization configuration of the target batch.
[0120] Please refer to Figure 6 , and the steps for real-time sending of anomaly and sampling adjustment notifications are specifically as follows:
[0121] Configure the Java network and communication APIs, including the JavaMail API and WebSocket, based on sampling adjustment, and formulate trigger rules for different types of notifications according to the degree of exception and sampling data to obtain a notification trigger rule set.
[0122] Configure the Java network and communication APIs, call the JavaMail API and WebSocket, set the sending interfaces for emails and real-time notifications respectively, extract the degree of exception and sampling data as notification trigger parameters, set a quantitative grading standard for the degree of exception, divide the degree of exception into multiple levels, identify the data points that meet the trigger conditions by parsing the sampling data, match them with the quantitative grading standard, screen out the eligible notification types and notification content, and generate corresponding trigger conditions for different exception types according to the set rules, thereby generating a notification trigger rule set that includes notification types, trigger conditions, and notification content.
[0123] Based on the notification trigger rule set, use the formula:
[0124]
[0125] Dynamically adjust the order and frequency of notification sending to obtain a notification sending queue and status record, where NV h is the sending priority of the h-th notification, used to determine its order in the sending queue, and DV h is the data trigger level, reflecting the importance or degree of exception of the data points that trigger this notification, W is the weight set according to the urgency of the notification, and SV is the current system load, represented by the number of notifications being processed by the system, used to balance the system resources and the need for notification sending;
[0126] There are three notifications with data trigger levels DV h being 5, 3, and 2 respectively, the weight W is 1.5 for all notifications, and the current system load SV is 4. According to the formula calculation, for the first notification:
[0127]
[0128] For the second notification:
[0129]
[0130] For the third notification:
[0131]
[0132] The results show that the first notification has the highest priority for sending because it has the highest data trigger level, while the second and third notifications have lower priorities, reflecting the relatively low data trigger levels, helping to determine the position of each notification in the sending queue and ensuring that they are prioritized according to their importance and urgency.
[0133] Analyze the notification sending queue and status records to verify the notification sending efficiency and success rate, and reorder and resend the notifications that were not successfully sent to obtain the notification sending records;
[0134] Perform status retrieval on each notification in the sending queue, determine whether the notification sending result is successful by recording the sending time, status code and number of failures, extract the notifications that failed to be sent, reorder the failed notification queue, set the number of sending attempts according to the reordered notification priority, and record the timestamp and status code of each sending operation, resend the notifications that were not successfully sent in turn, and compare the resending results with the initial status, update the notification sending record, and finally obtain the notification sending record including the sending status, number of resending and status code.
[0135] See also Figure 7 , the specific steps for obtaining operation log details are as follows:
[0136] Capture activities associated with sampling adjustment, record the request and response of each sampling adjustment, including timestamp, operation type and status, and obtain a preliminary log data set;
[0137] The communication interaction with the system during each sampling adjustment process is recorded, including capturing the request sending time, response time, operation type, and status code returned by the system. The data is arranged in sequence by timestamp to ensure the time continuity of the log. At the same time, the system event tracker is called to mark the operation record associated with the log according to the ID of the sampling request. All data is recorded in the initial log table. By comparing the corresponding IDs of the request and response, the complete operation records are screened out, and the unmatched or missing data is marked for subsequent processing, forming a preliminary log data set containing timestamp, operation type, system status and response information.
[0138] Process the preliminary log data set, remove invalid or duplicate entries, and classify and sort the logs to optimize the readability and usability of the log data and obtain the operation log details;
[0139] Read log entries one by one, and eliminate invalid data entries, including response failure records, duplicate timestamps or operation type entries. Ensure the integrity of the log time sequence by sorting and deduplicating timestamps. At the same time, call the operation classification standard to classify the logs according to the operation type, and classify the log entries into "request sent", "response received" and "operation retry". Then optimize the data sorting according to the status code and timestamp, eliminate abnormal records, integrate valid data and store it as a structured log, and obtain the cleaned, classified and sorted operation log details for further data review and analysis.
[0140] See also Figure 8 , the specific steps for obtaining sampling optimization records are:
[0141] Based on the operation log details, extract the time and parameters of sampling adjustment, perform time series analysis through JTimeSeries, identify the periodic patterns and abnormal changes of the data, and obtain the time series analysis results;
[0142] By parsing each record in the operation log, we obtain the timestamp, operation parameters and related events of the sampling adjustment, normalize the time field through data extraction, and filter out the valid time and parameters. Then, we use JTimeSeries to identify the periodicity and abnormal fluctuation of the data. The analysis process compares the sampling frequency, time interval and trend curve of parameter changes at each time point, calculates the deviation value of each data point, identifies potential periodic patterns and abnormal change points, and generates time series analysis results.
[0143] Based on the results of time series analysis, the formula is used:
[0144]
[0145] Predict the sampling process and timing to obtain the predicted sampling optimization value PU i , among which, XU ij is the value of the jth independent variable in the i-th data, such as the sampling frequency or characteristic data at a specific time point, which is used to calculate the sampling optimization prediction value of each sample, β j is the jth independent variable XU ij The regression coefficient of Z indicates the influence of this variable on the sampling optimization value. i is the adjustment factor used to adjust the forecast process nonlinearly, e is the base of the natural logarithm, β0 is the intercept or baseline coefficient, which provides the basic forecast of the sampling optimization value without considering the influence of other variables, and n U is the total number of independent variables;
[0146] There are three independent variables, where β0 = 0.5 (baseline effect), β1 = 0.3 (effect of the first independent variable), β2 = -0.2 (effect of the second independent variable), β3 = 0.1 (effect of the third independent variable), and Z i = 0.05 (adjustment factor), and the independent variable values XU i1 = 5, XU i2 = 3, and XU i3 = 2. Substitute these values into the formula for calculation:
[0147]
[0148] The result shows that, based on the current model and the input independent variable values, the predicted sampling optimization value is 1.261, indicating that the sampling timing and process need to be optimized according to this predicted value to improve efficiency.
[0149] Based on the predicted sampling optimization value, analyze the correlation between sample characteristics and sampling effect, and determine the optimal sampling process and timing to obtain a sampling optimization record;
[0150] By matching the sample characteristic data with the predicted sampling optimization value, extract the specific characteristic parameters of each sample, including the physical attributes of the sample, the historical sampling success rate, and the range of variation of the characteristic data. Subsequently, compare the data with the sampling effect of the sample one by one, calculate the correlation index, select the characteristic parameters with statistical significance, clarify the main factors affecting the sampling effect through correlation calculation, and adjust the time and process according to the sampling optimization value. Record the optimization results to generate a sampling optimization record, recording the influence degree of each characteristic parameter on the sampling process and the correlation analysis results.
[0151] An intelligent sample sampling management method based on the Java language, including the following steps:
[0152] S1: Based on the data stream interface, continuously receive sample data, use the random forest algorithm to analyze the status of data points, identify data points that deviate from the normal mode, and judge the abnormal status of the sample to obtain an abnormal data identification mark;
[0153] S2: Based on the abnormal data identification mark, call the Java concurrent package to analyze the degree and frequency of data point abnormalities, calculate the number of sampling points that need to be adjusted, and dynamically optimize the sampling frequency for the target batch to obtain a sampling adjustment configuration;
[0154] S3: Based on the sampling adjustment configuration, use the Java network and communication API to send abnormal and sampling adjustment notifications in real time, monitor the sending status of the notifications, and record the associated sampling operation logs simultaneously to obtain the operation log details;
[0155] S4: Based on the operation log details, combined with historical sampling data, use JTimeSeries or Deeplearning4j to predict the time series of the sample data, analyze the correlation between the sample characteristics and the sampling effect, and obtain the sampling mode analysis result;
[0156] S5: Based on the sampling mode analysis result, evaluate the effect of the differential sampling mode, and determine the optimal sampling time point and process according to the sample characteristics and the data accuracy requirements, and obtain the sampling optimization record.
[0157] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still belong to the protection scope of the technical solution of the present invention.
Claims
1. An intelligent sample sampling management system based on the Java language, characterized in that, The system includes: The anomaly detection module continuously receives sample data based on the data stream interface, analyzes the status of data points using the random forest algorithm, identifies data points that deviate from the normal pattern, and determines the anomaly status of the sample to obtain an anomaly data identification mark. Based on the anomaly data identification mark, the sampling optimization module calls the Java concurrent package, analyzes the anomaly degree and occurrence frequency of data points, calculates the number of sampling points that need to be adjusted, and dynamically optimizes the sampling frequency for the target batch to obtain a sampling adjustment configuration. Based on the sampling adjustment configuration, the notification management module uses the Java network and communication API to send anomaly and sampling adjustment notifications in real time, monitors the sending status of the notifications, and records the associated sampling operation logs to obtain operation log details. Based on the operation log details, the sampling analysis module combines historical sampling data, uses JTimeSeries or Deeplearning4j to predict the time series of sample data, identifies the sampling pattern, and analyzes the correlation between sample characteristics and sampling effects to determine the optimal sampling process and timing, obtaining a sampling optimization record.
2. The intelligent sample sampling management system based on the Java language according to claim 1, wherein The specific steps for analyzing the status of data points are as follows: Based on the data stream interface, continuously receive sample data, set up a data buffer, and store the instant data stream to obtain a data stream set. Apply the random forest algorithm to analyze the data stream set, calculate the status index of each data point, using the formula: Obtain a set of status scores, where S i represents the status score of data point i, and w j is the weight of feature j, and a ij is the current monitored value of data point i on feature j, and m is the total number of features; According to the status score set, filter out the data points with scores lower than the threshold, and identify the data points that deviate from the normal pattern to obtain a set of data points with a deviation pattern.
3. The intelligent sample sampling management system based on the Java language according to claim 1, wherein The specific steps for obtaining the anomaly data identification mark are as follows: Mark the data points that deviate from the normal pattern, assign a unique non-conventional behavior mark to them, and obtain the screened anomaly data. Based on the processed anomaly data, determine the anomaly status of the sample, and associate the mark of the data point with the anomaly status to obtain an anomaly data identification mark.
4. The intelligent sample sampling management system based on the Java language according to claim 1, wherein The specific steps for analyzing the anomaly degree and occurrence frequency are as follows: Based on the anomaly data identification mark, collect the number of anomalies and timestamps of each data point, use the Java concurrent package to implement parallel processing, and integrate the anomaly events of the data points and the corresponding times to obtain a preliminary anomaly data set. Based on the preliminary anomaly data set, use the formula: Calculate the abnormality degree of each data point to obtain the abnormality degree and frequency analysis results, where ES i represents the abnormality degree of data point i, and nS i represents the number of abnormalities of data point i during the monitoring period. TS is the total time of the monitoring period, and fS i is the occurrence frequency of data point i during the monitoring period.
5. The intelligent sample sampling management system based on the Java language according to claim 1, characterized in that, The specific steps for obtaining the sampling adjustment configuration are as follows: Based on the anomaly degree and occurrence frequency, determine the number of data points that need to increase the sampling frequency, and set priorities according to the high and low anomaly degrees to obtain a sampling adjustment target list. Based on the sampling adjustment target list, use the formula: FS′ i = FS0 + α X ·ES i ; Adjust the sampling frequency to obtain a new sampling frequency, where FS′ i represents the adjusted sampling frequency of data point i, FS0 is the baseline sampling frequency, and ES i is the degree of abnormality of data point i, and α X is the adjustment coefficient; Apply the new sampling frequency to the target batch, and adjust the associated sampling settings to optimize the sampling data collection efficiency and accuracy to obtain a sampling adjustment configuration.
6. The intelligent sample sampling management system based on the Java language according to claim 1, characterized in that The specific steps for sending anomaly and sampling adjustment notifications in real time are as follows: Based on the sampling adjustment configuration, configure the Java network and communication API, and formulate trigger rules for different types of notifications according to the anomaly degree and sampling data to obtain a notification trigger rule set. Based on the notification trigger rule set, use the formula: Dynamically adjust the order and frequency of notification sending to obtain a notification sending queue and status records, where NV h is the sending priority of the h-th notification, and DV h is the data trigger level, W is the weight set according to the urgency of the notification, and SV is the current system load; The notification sending queue and status record are analyzed to verify the sending efficiency and success rate of the notification, and the notifications that are not successfully sent are reordered and resent to obtain the notification sending record.
7. The intelligent sample sampling management system based on the Java language according to claim 1, characterized in that, The steps for obtaining the operation log details are as follows: Capture activities associated with sampling adjustment, record the request and response of each sampling adjustment, including timestamp, operation type and status, and obtain a preliminary log data set; The preliminary log data set is processed to remove invalid or duplicate entries, and the logs are classified and sorted to optimize the readability and usability of the log data to obtain operation log details.
8. The intelligent sample sampling management system based on the Java language according to claim 1, characterized in that The steps for obtaining the sampling optimization record are specifically as follows: Based on the operation log details, extract the time and parameters of the sampling adjustment, perform time series analysis through JTimeSeries, identify the periodic patterns and abnormal changes of the data, and obtain the time series analysis results; Based on the time series analysis results, the formula is adopted: Predict the sampling process and timing to obtain the predicted sampling optimization value PU i , where XU ij is the value of the j-th independent variable in the i-th data, β j is the regression coefficient of the j-th independent variable XU ij , Z i is the adjustment factor, e is the base of the natural logarithm, β0 is the intercept or baseline coefficient, and n U is the total number of independent variables; Based on the predicted sampling optimization value, the correlation between sample characteristics and sampling effect is analyzed, and the optimal sampling process and timing are determined to obtain a sampling optimization record.
9. An intelligent sample sampling management method based on the Java language, characterized in that, The intelligent sample sampling management system based on Java language according to any one of claims 1 to 8 is implemented, comprising the following steps: Based on the data stream interface, the sample data is continuously received, and the random forest algorithm is used to analyze the data point status, identify the data points that deviate from the normal mode, and judge the abnormal state of the sample to obtain the abnormal data identification mark; Based on the abnormal data identification mark, call Java and send a package to analyze the abnormal degree and frequency of the data point, calculate the number of sampling points that need to be adjusted, and dynamically optimize the sampling frequency for the target batch to obtain the sampling adjustment configuration; Based on the sampling adjustment configuration, using Java network and communication API, sending exception and sampling adjustment notifications in real time, monitoring the sending status of the notifications, and recording the associated sampling operation logs to obtain operation log details; Based on the operation log details and in combination with historical sampling data, use JTimeSeries or Deeplearning4j to predict the time series of sample data, analyze the correlation between sample characteristics and sampling effects, and obtain sampling pattern analysis results; Based on the sampling pattern analysis results, the effect of the differential sampling pattern is evaluated, and according to the sample characteristics and data accuracy requirements, the optimal sampling time point and process are determined to obtain the sampling optimization record.