Resource dynamic allocation system and method for differentiated data
By dividing the laboratory into operational phases, setting data priorities, and constructing a knowledge graph, the problem of unreasonable resource allocation was solved, the efficiency and stability of laboratory operations were improved, and precise resource matching and scientific management were achieved.
Patent Information
- Application Number
- CN202511362806.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-09-23
AI Technical Summary
The laboratory lacks a way to accurately divide the operation into stages based on the data growth pattern, making it impossible to carry out targeted management based on the data characteristics of different stages. Data processing lacks a scientific priority determination mechanism, and resource allocation fails to fully integrate with the business scenarios of each department, resulting in low resource utilization, a single dimension of effect evaluation, and an imperfect adjustment mechanism, making it difficult to guarantee operational efficiency and stability.
By dividing the operation into stages based on data growth, setting data priorities, constructing a knowledge graph for resource allocation, and evaluating and adjusting parameters through quantitative indicators, a closed-loop management system is formed to ensure that resource allocation is aligned with business needs.
This has led to an overall improvement in the efficiency of laboratory operations, reduced data processing delays and resource waste, enhanced the scientific nature and adaptability of resource allocation, and ensured the stability of long-term operational results.
Smart Images

Figure CN120853862B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of big data analysis, in particular to a resource dynamic allocation system and method for differentiated data. BACKGROUND
[0002] In the process of laboratory operation, various types of data will continue to be generated as the business advances. The data growth characteristics and business needs in different periods are significantly different. Reasonable operation stage definition, data processing strategy and resource allocation scheme are crucial to ensuring operational efficiency. Currently, laboratory operation management often faces multiple challenges: first, there is a lack of effective way to accurately divide the operation stage according to the data growth law, making it difficult to carry out targeted management according to the data characteristics of different stages; second, there is a lack of scientific priority determination mechanism for data processing, which cannot distinguish the importance of data combined with data attributes, and is prone to delay in processing key data; third, resource allocation fails to fully match the business scenarios of each department, lacks a systematic allocation tool based on historical resource data, and the resource utilization rate and department demand satisfaction are low; fourth, the effect evaluation dimension after resource allocation is single and the adjustment mechanism is imperfect, making it difficult to identify problems and optimize them in a timely manner, which restricts the efficiency and stability of laboratory operation as a whole. SUMMARY
[0003] The purpose of the present application is to provide a resource dynamic allocation system and method for differentiated data to solve the problems in the prior art.
[0004] To achieve the above purpose, the present application provides the following technical scheme: a resource dynamic allocation method for differentiated data, an intelligent management method comprising the following steps:
[0005] Step S1, dividing the laboratory operation stage according to the data growth, setting the data priority of each operation stage for each operation stage data;
[0006] Step S1-1, collecting the data volume data of each laboratory operation stage in the operation of the historical multiple same-scale laboratories, calculating the data volume growth between each stage by subtracting the total data volume of the adjacent two stages;
[0007] Step S1-2, calculating the mean value of the data volume growth between each stage, setting the mean value as threshold a, adding a buffer value b to the threshold a, setting the threshold interval [a-b, a+b] combined with the threshold value and the buffer value, a is the threshold value between each operation stage, b is the buffer value set according to the business scenario of each operation stage;
[0008] Step S1-3, set a time interval t for the laboratory operation cycle, divide the time window according to the time interval t, calculate the data growth of each adjacent time window, when the growth of two adjacent time windows is in the threshold interval, determine the operation stage of the adjacent time window according to different threshold intervals;
[0009] Step S1-4, collect historical data of each operation stage, and clean the data, specifically:
[0010] For missing values, use mean filling, for outliers, use 3σ principle to identify, identify data exceeding as outliers, as mean, as standard deviation, then remove outliers, and remove duplicate values;
[0011] Step S1-5, set data priority according to data attributes, specifically:
[0012] Step S1-5-1, divide the data attributes into three dimensions of data processing time, latest delivery time and current operation stage core target;
[0013] Step S1-5-2, use quantitative index method to quantify the parameters of data processing time, latest delivery time and current operation stage core target to the interval [1, x];
[0014] Step S1-5-3, use weighted average method to calculate the priority score of each type of data, specifically:
[0015] ;
[0016] In the formula, P is the priority score of the data, A is the quantitative value of the data processing time, B is the quantitative value of the latest delivery time, C is the quantitative value of the core target matching degree, is the weight of the data processing time dimension, is the weight of the latest delivery time dimension, is the weight of the core target dimension;
[0017] Step S1-5-4, according to the priority score of the data, sort the priority score in ascending order, and then use quartile method to divide the threshold interval [x1, x2], [x2, x3], [x3, x4], according to the threshold interval of the priority value of each type of data Set the data priority.
[0018] The laboratory operation stage is divided by the data growth amount, the threshold interval is calculated combined with the historical data, and the time window is matched, so that the operation stage can be accurately defined, the management confusion caused by ambiguous stage division is avoided, the data is cleaned to ensure the data quality, and a reliable data basis is provided for subsequent analysis, and the data priority is set to provide a basis for subsequent resource allocation.
[0019] Step S2, processing the data by calling resources according to the data priority;
[0020] Step S2-1, marking the data with priority according to the data priority of each operation stage, and outputting a structured data set of 'data + priority';
[0021] Step S2-2, sorting the structured data set of 'data + priority' in ascending order, and classifying the data according to the priority to construct structured data sets of different priorities;
[0022] Step S2-3, sorting the structured data in the structured data set in descending order according to the P value, and the system processes the data from left to right in the structured data set of high level first.
[0023] The phenomenon of wasting resources by low-level data is avoided, the key data can be processed first, the delay of key tasks is reduced, and the pertinence and efficiency of data processing are improved.
[0024] Step S3, dividing and matching the resource data according to the business scenarios of each department of the laboratory;
[0025] Step S3-1, dividing the departments according to the different business processes of the laboratory, constructing a department table, extracting the business scenarios based on the business processes of each department, and clearly defining the business goals and data association of each business scenario to construct a 'department-business scenario' set;
[0026] Step S3-2, setting the words with the highest frequency and containing professional terms in the department business process as keywords, extracting multiple keywords of the business scenario, and constructing a 'department-business scenario keyword' set;
[0027] Step S3-3, text extraction of resources in the business scenario;
[0028] Step S3-4, setting the words related to the business process as keywords, and extracting the keywords from the text data converted by the NLG technology using a semantic recognition model;
[0029] Step S3-5, calculating the similarity between the business scenario keywords and the text data keywords, specifically:
[0030] Step S3-5-1, use the word segmentation tool Jieba to split the phrase sentence keywords into word units;
[0031] Step S3-5-2, remove mood auxiliaries, transition words, special symbols and low-frequency words, and keep the core semantic words;
[0032] Step S3-5-3, set the keyword sequence length, unify the keyword sequence length, fill the sequence shorter than the keyword sequence length with blank characters, and truncate the sequence longer than the keyword sequence length;
[0033] Step S3-5-4, collect all business scenario keywords and text data keywords, generate an original word set, remove duplicates from the original word set, filter low-frequency words, and build a global word table. Assign a unique ID and index to each word;
[0034] Step S3-5-5, set the data size of the keyword as the complexity of the keyword, set the word vector dimension according to the complexity of the keyword, establish the matching relationship of "keyword complexity-word vector dimension", set the word vector dimension to 50-100 for keywords associated with only 1-2 business scenarios and data size between x1-x2, set the dimension to 100-200 for keywords associated with 3-5 business scenarios and data size between x2-x3, and set the dimension to 200-300 for keywords associated with more than 5 business scenarios and data size exceeding x3. x1, x2, x2 are interval thresholds for dividing the complexity of keywords according to data size. Each word in the global word table is randomly generated a fixed-dimension word vector to build a word vector matrix;
[0035] Step S3-5-6, extract any two business scenario keywords and text data keywords to form a keyword pair, set the business scenario keyword and the matching text data keyword pair as the positive sample, and the business scenario keyword and the non-matching text data keyword pair as the negative sample, and build the training sample;
[0036] Step S3-5-7, use a triple loss function to build a training model, specifically:
[0037] ;
[0038] In the formula, L is the loss value, is the maximum value operation, is a similarity calculation function, q is an encoded semantic vector, is a negative sample vector, is a positive sample vector, is the minimum gap between the similarity of the positive sample and the similarity of the negative sample, and is a threshold value set by professionals according to the business scenario.
[0039] Step S3-5-8, encode the keywords according to the training model, and then calculate the keyword pair similarity, specifically:
[0040]
[0041] In the formula, is the final calculated similarity, is the business scenario keyword, is the text data keyword, is the length of the business scenario keyword, is the length of the text data keyword;
[0042] Step S3-6, take the average of the keyword pair as the similarity threshold, and when the threshold is exceeded, determine that the resource data meets the business scenario of the department.
[0043] The accurate matching of resource data and department business scenarios is achieved, avoiding the mismatch of resources leading to the inability to meet business needs. Through threshold determination, it ensures that the resource allocation is in line with the actual business of the department, and improves the fit of resource use and business needs.
[0044] Step S4, for different operation stages, construct a knowledge graph to allocate resources, and use the knowledge graph to allocate resources for each operation stage;
[0045] Step S4-1, divide the resources into one-level categories, specifically hardware resources and software resources, divide the hardware resources into two-level categories, specifically computing resources and storage resources, the specific entities of computing resources are CPU computing power, GPU computing power, and memory capacity, the specific entities of storage resources are storage capacity and storage device read-write speed, and the specific entities of software resources are the number of data processing tools authorized and the frequency of data processing tools called;
[0046] Step S4-2, for each operation stage of the laboratory division, collect historical resource quantities of each operation stage in the historical operation process of the same scale, and use the data cleaning method of step S1-4 to clean the historical resources;
[0047] Step S4-3, use the three-layer ontology structure defined in step S4-1 as the mapping skeleton, traverse each resource after cleaning, and map CPU computing power, GPU computing power, memory capacity, storage capacity, storage device read-write speed, the number of data processing tools authorized, and the frequency of data processing tools to the skeleton according to the upper level, to construct a knowledge graph;
[0048] Step S4-4, extract the historical resource quantities of each operation stage in the knowledge graph as the basis for allocating resources.
[0049] By constructing a knowledge graph to allocate resources, systematically sorting out resource categories and associated relationships, avoiding resource management fragmentation, and allocating resources based on historical resource data, the resource supply can adapt to the actual needs of different operation stages, improving the scientificity and rationality of resource allocation.
[0050] Step S5, analyze and evaluate the resource allocation result after adjusting the parameters, and when the evaluation is not up to standard, analyze the reasons and propose solutions;
[0051] Step S5-1, divide the core evaluation indicators of the resource allocation result into data processing timeliness, resource utilization rate, and department demand satisfaction degree;
[0052] Step S5-2, quantize the data processing timeliness, resource utilization rate, and department demand satisfaction degree, specifically:
[0053] Data processing timeliness: ;
[0054] In the formula, when T act ≤T exp , S tim ≥1, is the unified value range, and finally min(S tim ,1), is the zero correction term, is the quantized value of data processing timeliness, is the actual total time consumption of a certain type of data processing task, is the expected total time consumption of the task;
[0055] Resource utilization rate:
[0056] ;
[0057] In the formula, is the quantized value of resource utilization rate, is the weight of equipment resources, is the weight of personnel resources, is the actual running time of resource equipment, is the total available time of equipment per month, is the actual working time of personnel, is the total rated time of personnel per month;
[0058] Department demand satisfaction degree: ;
[0059] In the formula, is the quantized value of department demand satisfaction degree, n is the amount of resource demand proposed by the department, is the weight of the jth demand, is the actual satisfaction amount of the jth demand, the total quantity of department applications for the jth demand, to prevent zero correction terms;
[0060] Step S5-3, standardizing the quantified values of data processing timeliness, resource utilization rate, and department demand satisfaction degree;
[0061] Step S5-4, integrating the quantified values of data processing timeliness, resource utilization rate, and department demand satisfaction degree, specifically:
[0062]
[0063] wherein, is the core evaluation index of the resource allocation result, is the weight of data processing timeliness, is the weight of resource utilization rate, is the weight of department demand degree, is the quantified value of data processing timeliness, is the quantified value of resource utilization rate, is the quantified value of department demand satisfaction degree;
[0064] Step S5-5, for the core evaluation index, setting a threshold value according to the business scenario, and when the core evaluation index exceeds the threshold value, it is determined that the resource allocation result is qualified, and the resource allocation result is solidified as the standard result of the operation stage, which is used for subsequent data processing in the same stage, records the running data corresponding to the resource allocation result, and stores it in the historical sample library;
[0065] Step S5-6, when the core evaluation index does not meet the standard, extracting the three sub-index threshold values preset by the business scenario, which are set by professionals according to the business scenario, comparing each sub-index with the corresponding sub-index threshold value, and dividing the unqualified reasons according to the sub-index that does not meet the corresponding sub-index threshold value into:
[0066] When S tim-comp < S tim-threshold is determined to be unqualified in data processing timeliness, S tim-comp is the quantified value of data processing timeliness after standardization, and S tim-threshold is the threshold value of data processing timeliness: specific analysis of insufficient hardware resources CPU, GPU computing power and storage capacity, data processing tool authorization, calling frequency idling or overload, and corresponding addition of resources according to system log and software monitoring platform;
[0067] When S res-comp < S res-threshold is determined to be unqualified in resource utilization rate, S res-comp S res-threshold Resource utilization threshold: check the model input data for data deviation;
[0068] When S dmd-comp <S dmd-threshold is determined to be a department demand satisfaction that does not meet the standard, S dmd-comp S is a quantitative value of department demand satisfaction after standardization processing dmd-threshold S is a threshold value of department demand satisfaction: use step S3 to recalculate the matching.
[0069] By quantifying the core evaluation indicators by dimension, the objective evaluation and accurate analysis of resource allocation results are realized, and targeted solutions are proposed to ensure that resource allocation continues to adapt to operational needs and improves overall operational stability.
[0070] The system includes an operation stage and data priority division module, a resource calling processing module, a resource data matching module, a resource allocation module, and a resource allocation result evaluation and adjustment module.
[0071] The operation stage and data priority division module is used to divide each operation stage of the laboratory operation, and set the data priority of each operation stage according to the data attribute;
[0072] The resource calling processing module is used to generate a structured set according to the data priority, and then call the resource to process the data;
[0073] The resource data matching module is used to divide the laboratory departments, extract the business scenarios of each department, extract the keywords and match the resource data;
[0074] The resource allocation module is used to construct a resource allocation knowledge graph, and allocate resources for each operation stage according to the knowledge graph;
[0075] The resource allocation result evaluation and adjustment module is used to evaluate the resource allocation result, analyze the reasons for the substandard condition and propose solutions.
[0076] The operation stage and data priority division module includes an operation stage division unit and a data priority setting unit;
[0077] The operation stage division unit is used to calculate the threshold interval according to the historical data growth, combine the time window to divide the operation stage and clean the historical data;
[0078] The data priority setting unit is used to quantify the data attribute, calculate the priority score by weighting algorithm, and then divide the threshold interval to set the data priority.
[0079] The resource calling processing module comprises a data priority structured set generation unit and a priority data processing unit;
[0080] The data priority structured set generation unit is used for marking the priority of each stage data and outputting a structured data set of "data+priority".
[0081] The priority data processing unit is used for sorting the structured set, and then processing the set data according to the priority from high to low in descending order of P value.
[0082] The resource data matching module comprises a business scenario keyword extraction unit and a keyword similarity matching unit.
[0083] The business scenario keyword extraction unit is used for constructing a "department-business scenario keyword" set and extracting text data keywords.
[0084] The keyword similarity matching unit is used for calculating the similarity of the business scenario and the text data keywords, and determining whether the resource data meets the department business scenario according to a threshold.
[0085] The resource allocation module comprises a knowledge graph construction unit and a knowledge graph resource allocation unit.
[0086] The knowledge graph construction unit is used for dividing resource categories, cleaning historical resource data, and constructing a knowledge graph by mapping resources to a skeleton.
[0087] The knowledge graph resource allocation unit is used for extracting historical resource quantities of each operation stage of the knowledge graph, and allocating resources by taking the historical resource quantities as a reference.
[0088] The resource allocation result evaluation unit is used for quantifying core evaluation indexes, standardizing and integrating the indexes, and setting a threshold to determine whether the resource allocation parameters are qualified.
[0089] The evaluation non-compliance adjustment unit is used for comparing sub-indexes with a compliance threshold, dividing non-compliance reasons, and proposing data processing, resource adjustment or re-matching schemes.
[0090] Compared with the prior art, the present application has the following advantages:
[0091] 1. The present application divides operation stages, sets data priority, matches resources, allocates resources according to a knowledge graph, and evaluates the allocation results, forming a closed-loop management, effectively improving the overall efficiency of laboratory operation, and reducing data processing delay and resource waste.
[0092] 2. The present application quantifies indexes, replaces subjective judgment, reduces errors, and improves the scientificity and objectivity of data processing and resource allocation.
[0093] 3. This invention adapts to the characteristics of different operating periods by dividing the work into stages, and optimizes the allocation parameters through evaluation and adjustment, so as to flexibly respond to business changes and demand adjustments in laboratory operations and ensure the stability and adaptability of long-term operating results. Attached Figure Description
[0094] Fig. 1 This is a flowchart illustrating the resource dynamic allocation method for differentiated data according to the present invention.
[0095] Fig. 2 This is a schematic diagram of the structure of the resource dynamic allocation system for differentiated data according to the present invention. Detailed Implementation
[0096] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0097] Example 1: As Figs. 1-2 As shown, the present invention provides a technical solution for a dynamic resource allocation method for differentiated data, the method comprising the following steps:
[0098] Step S1: Divide the laboratory operation into stages based on the amount of data growth, and set the data priority for each operation stage.
[0099] Step S1-1: Collect data on the amount of data for each stage of operation of multiple laboratories of the same scale in history, and calculate the increase in data amount between each stage by subtracting the total amount of data from the total amount of data for two adjacent stages.
[0100] Step S1-2: Calculate the average data volume growth between each stage, set the average value as the threshold a, add a buffer value b for the threshold a, and combine the threshold and the buffer value to set the threshold range [ab, a+b], where a is the threshold between each operation stage and b is the buffer value set according to the business scenario for each operation transition stage.
[0101] Step S1-3: Set a time interval t for the laboratory operation cycle, divide the time window according to the time interval t, calculate the data growth of each adjacent time window, and when the growth of two adjacent time windows is within the threshold range, determine the operation stage of the adjacent time window according to the different threshold ranges.
[0102] Steps S1-4: Collect data from various historical operational phases and clean the data, specifically as follows:
[0103] For missing values, use mean filling, for outliers, use 3σ principle to identify, identify data exceeding as outliers, as mean, as standard deviation, then remove outliers, and remove duplicate values;
[0104] Step S1-5, according to the data attribute, set the data priority, specifically:
[0105] Step S1-5-1, divide the data attribute into three dimensions of data processing time, latest delivery time and current operation stage core target;
[0106] Step S1-5-2, use quantitative index method to quantify the parameters to the interval [1,x] for data processing time, latest delivery time and current operation stage core target three dimensions;
[0107] Step S1-5-3, use weighted average method to calculate the priority score of each type of data, specifically:
[0108]
[0109] In the formula, P is the priority score of the data, A is the quantitative value of the data processing time, B is the quantitative value of the latest delivery time, C is the quantitative value of the core target matching degree, is the weight of data processing time dimension, is the weight of the latest delivery time dimension, is the weight of the core target dimension;
[0110] Step S1-5-4, according to the priority score of the data, sort the priority score in ascending order, and then use the quartile method to divide the threshold interval [x1, x2], [x2, x3], [x3, x4], according to the threshold interval of the priority value of each type of data Set the data priority.
[0111] By dividing the laboratory operation stage by the data growth amount, combining with the historical data to calculate the threshold interval and matching the time window, the operation stage can be accurately defined, avoiding the management confusion caused by fuzzy stage division. Data cleaning is performed on the data to ensure data quality and provide a reliable data basis for subsequent analysis. By setting the data priority, it provides a basis for subsequent resource allocation.
[0112] Step S2, according to the data priority, call resources to process the data;
[0113] Step S2-1, according to the data priority of each operation stage, mark the data priority, and output the structured data set of "data + priority";
[0114] Step S2-2, for the structured data set of "data + priority", the data is classified according to the priority after ascending order sorting, and the structured data sets of different priorities are constructed;
[0115] Step S2-3, for the structured data sets of different priorities, the structured data in the structured data set is sorted in descending order according to the P value, and the system processes the data from left to right in the structured data set of high level in priority.
[0116] Avoiding the phenomenon of wasting resources by low-level data, ensuring that key data can be processed in priority, reducing the delay of key tasks, and improving the pertinence and efficiency of data processing.
[0117] Step S3, divide and match the resource data according to the business scenarios of each department of the laboratory;
[0118] Step S3-1, divide the departments according to the different business processes of the laboratory, construct a department table, extract the business scenarios based on the business processes of each department, and clarify the business goals and data association of each business scenario, and construct a "department-business scenario" set;
[0119] Step S3-2, set the words with the highest frequency and containing professional terms in the department business process as keywords, extract multiple keywords of the business scenario, and construct a "department-business scenario keyword" set;
[0120] Step S3-3, text extraction of resources in business scenarios;
[0121] Step S3-4, set the words related to the business process as keywords, and use the semantic recognition model to extract the keywords for the text data converted by the NLG technology;
[0122] Step S3-5, calculate the similarity between the business scenario keywords and the text data keywords, specifically:
[0123] Step S3-5-1, use the word segmentation tool Jieba to split the phrase keywords into word units;
[0124] Step S3-5-2, remove mood auxiliaries, transition words, special symbols and low-frequency words, and keep the core semantic words;
[0125] Step S3-5-3, set the keyword sequence length, unify the keyword sequence length, fill the sequence shorter than the keyword sequence length with blank characters, and truncate the sequence longer than the keyword sequence length;
[0126] Step S3-5-4, collect all the word units of the business scenario keywords and the text data keywords, generate an original word set, remove duplicates from the original word set, filter low-frequency words, construct a global word table, and assign a unique ID and index to each word;
[0127] Step S3-5-5, set the data size of the keyword as the complexity of the keyword, set the word vector dimension according to the complexity of the keyword, establish a matching relationship between the keyword complexity and the word vector dimension, set the word vector dimension to 50-100 for keywords associated with only 1-2 business scenarios and data size between x1-x2, set the dimension to 100-200 for keywords associated with 3-5 business scenarios and data size between x2-x3, and set the dimension to 200-300 for keywords associated with more than 5 business scenarios and data size exceeding x3, x1, x2, and x2 are interval thresholds for dividing the complexity of the keywords according to the data size, each word in the global word table is traversed, a fixed-dimension word vector is randomly generated for each word in the global word table, and a word vector matrix is constructed;
[0128] Step S3-5-6, extract any two business scenario keywords and text data keywords to form a keyword pair, set the business scenario keyword and the matching text data keyword pair as a positive sample, and set the business scenario keyword and the non-matching text data keyword pair as a negative sample, and construct a training sample;
[0129] Step S3-5-7, use a triple loss function to construct a training model, specifically:
[0130] ;
[0131] In the formula, L is the loss value, is a max operation, is a similarity calculation function, q is an encoded semantic vector, is a negative sample vector, is a positive sample vector, is a control for the minimum gap between the similarity of the positive sample and the similarity of the negative sample, and is a threshold value set by professionals according to the business scenario;
[0132] Step S3-5-8, encode the keywords according to the training model, and then calculate the similarity of the keyword pair, specifically:
[0133] ;
[0134] In the formula, is the final calculated similarity, is the business scenario keyword, is the text data keyword, The module length of the business scenario keyword, The module length of the text data keyword;
[0135] Step S3-6, the average of the keyword pair is taken as a similarity threshold value, and when the threshold value is exceeded, it is determined that the resource data meets the business scenario of the department.
[0136] The accurate matching of resource data and department business scenarios is achieved, avoiding the mismatch of resources leading to the inability to meet business needs. Through threshold determination, it is ensured that the resource allocation is in line with the actual business of the department, and the matching degree of resource use and business needs is improved.
[0137] Step S4, for different operation stages, construct a knowledge graph to allocate resources, and use the knowledge graph to allocate resources for each operation stage;
[0138] Step S4-1, the resources are classified into one level, specifically hardware resources and software resources, the hardware resources are classified into two levels, specifically computing resources and storage resources, the specific entities of computing resources are CPU computing power, GPU computing power and memory capacity, the specific entities of storage resources are storage capacity and storage device read-write speed, and the specific entities of software resources are the number of data processing tools authorized and the frequency of data processing tools called;
[0139] Step S4-2, for each operation stage of the laboratory, collect historical resource quantities of each operation stage in the historical operation process of the same scale, and use the data cleaning method of step S1-4 to clean the historical resources;
[0140] Step S4-3, according to the three-layer ontology structure defined in step S4-1, which is the mapping skeleton, traverse each resource after cleaning, and map CPU computing power, GPU computing power, memory capacity, storage capacity, storage device read-write speed, the number of data processing tools authorized, and the frequency of data processing tools to the skeleton according to the upper level, to construct a knowledge graph;
[0141] Step S4-4, extract the historical resource quantities of each operation stage in the knowledge graph as the basis for allocating resources.
[0142] By constructing a knowledge graph to allocate resources, the resource categories and associated relationships are systematically sorted out, avoiding resource management fragmentation, and based on historical resource data allocation, the resource supply is adapted to the actual needs of different operation stages, improving the scientificity and rationality of resource allocation.
[0143] Step S5, analyze and evaluate the resource allocation results after adjusting the parameters, and when the evaluation is not up to standard, analyze the reasons and propose solutions;
[0144] Step S5-1, the core evaluation indexes of the resource allocation result are divided into data processing timeliness, resource utilization, and department demand satisfaction;
[0145] Step S5-2, the data processing timeliness, resource utilization, and department demand satisfaction are quantified, specifically:
[0146] Data processing timeliness:
[0147] In the formula, when T act ≤T exp , S tim ≥1, is the unified value range, and finally min(S tim , 1) is taken, is a zero correction term, is the quantized value of data processing timeliness, is the actual total time consumption of a certain type of data processing task, is the expected total time consumption of the task;
[0148] Resource utilization:
[0149]
[0150] In the formula, is the quantized value of resource utilization, is the weight of equipment resources, is the weight of personnel resources, is the actual running time of resource equipment, is the total available time of equipment per month, is the actual working time of personnel, is the total rated time of personnel per month;
[0151] Department demand satisfaction:
[0152] In the formula, is the quantized value of department demand satisfaction, n is the resource demand quantity proposed by the department, is the weight of the jth demand, is the actual satisfaction quantity of the jth demand, is the total quantity of the jth demand applied by the department, is a zero correction term;
[0153] Step S5-3, the quantized values of data processing timeliness, resource utilization, and department demand satisfaction are standardized;
[0154] Step S5-4, the quantized values of data processing timeliness, resource utilization, and department demand satisfaction are integrated, specifically:
[0155] ;
[0156] In the formula, is the core evaluation index of the resource allocation result, is the weight of data processing timeliness, is the weight of resource utilization, is the weight of department demand degree, is the quantitative value of data processing timeliness, is the quantitative value of resource utilization, is the quantitative value of department demand satisfaction;
[0157] Step S5-5, for the core evaluation index, the core evaluation index is set according to the business scene. When the core evaluation index exceeds the threshold, it is determined that the resource allocation result is qualified, and the resource allocation result is solidified as the standard result of the operation stage, which is used for subsequent data processing in the same stage, the running data corresponding to the resource allocation result is recorded, and stored in the historical sample library;
[0158] Step S5-6, when the core evaluation index does not meet the standard, the three sub-index threshold values preset by the business scene are extracted. The three sub-index threshold values are set by professionals according to the business scene. The comparison is made between each sub-index and the corresponding sub-index threshold value. When a certain sub-index does not reach the corresponding sub-index threshold value, the unqualified reason is divided into:
[0159] When S tim-comp < S tim-threshold , it is determined that the data processing timeliness does not meet the standard. S tim-comp is the quantitative value of data processing timeliness after standardization processing, and S tim-threshold is the threshold value of data processing timeliness: specific analysis is that the hardware resources CPU, GPU computing power and storage capacity are insufficient, the data processing tool is authorized, the calling frequency is idle or overloaded, and the resources are added according to the analysis of system log and software monitoring platform;
[0160] When S res-comp < S res-threshold , it is determined that the resource utilization does not meet the standard. S res-comp is the quantitative value of resource utilization after standardization processing, and S res-threshold is the threshold value of resource utilization: whether the data deviation occurs is checked by analyzing the model input data;
[0161] When S dmd-comp < S dmd-threshold , it is determined that the department demand satisfaction does not meet the standard. S dmd-comp is the quantitative value of department demand satisfaction after standardization processing, and S dmd-thresholdTo meet the threshold value of department requirement satisfaction: use step S3 to recalculate the matching.
[0162] By quantifying the core evaluation indicators by dimensions, objective evaluation and precise analysis of resource allocation results are achieved, targeted solutions are proposed, resource allocation is continuously adapted to operational needs, and overall operational stability is improved.
[0163] The system includes an operation stage and data priority division module, a resource calling processing module, a resource data matching module, a resource allocation module, and a resource allocation result evaluation and adjustment module.
[0164] The operation stage and data priority division module is used to divide each operation stage of laboratory operation, and set the data priority of each operation stage according to the data attributes.
[0165] The resource calling processing module is used to generate a structured set according to the data priority, and then call resources to process data.
[0166] The resource data matching module is used to divide laboratory departments, extract business scenarios of each department, extract keywords, and match resource data.
[0167] The resource allocation module is used to construct a resource allocation knowledge graph, and allocate resources for each operation stage according to the knowledge graph.
[0168] The resource allocation result evaluation and adjustment module is used to evaluate the resource allocation results, analyze the reasons for the substandard situation, and propose solutions.
[0169] The operation stage and data priority division module includes an operation stage division unit and a data priority setting unit.
[0170] The operation stage division unit is used to calculate the threshold interval according to the historical data growth, combine the time window to divide the operation stage, and clean the historical data.
[0171] The data priority setting unit is used to quantify the data attributes, calculate the priority score by weighting algorithm, and then divide the threshold interval to set the data priority.
[0172] The resource calling processing module includes a data priority structured set generation unit and a priority data processing unit.
[0173] The data priority structured set generation unit is used to mark the priority of each stage data, and output the structured data set of "data + priority".
[0174] The priority data processing unit is used to sort the structured set, and then process the set data according to the priority from high to low in descending order of P value.
[0175] The resource data matching module comprises a business scenario keyword extraction unit and a keyword similarity matching unit.
[0176] The business scenario keyword extraction unit is configured to construct a "department-business scenario keyword" set and extract text data keywords.
[0177] The keyword similarity matching unit is configured to calculate the similarity between the business scenario and the text data keywords, and determine whether the resource data meets the department business scenario according to a threshold.
[0178] The resource allocation module comprises a knowledge graph construction unit and a knowledge graph resource allocation unit.
[0179] The knowledge graph construction unit is configured to divide resource categories, clean historical resource data, and construct a knowledge graph by mapping resources to a skeleton.
[0180] The knowledge graph resource allocation unit is configured to extract historical resource quantities of each operation stage of the knowledge graph, and allocate resources based on the extracted historical resource quantities.
[0181] The resource allocation result evaluation unit is configured to quantify core evaluation indicators, standardize and integrate the indicators, and set a threshold to determine whether the resource allocation parameters are qualified.
[0182] The evaluation non-compliance adjustment unit is configured to compare sub-indicators with a threshold, divide non-compliance reasons, and propose data processing, resource adjustment or re-matching solutions.
[0183] Embodiment two: Selecting historical data of a hospital in the past three years, three potential operation stages are preliminarily divided according to "business intensity", and the monthly average data quantity of each stage is counted:
[0184] Daily diagnosis and treatment data period: 48000 in 2021, 50000 in 2022, and 52000 in 2023;
[0185] Holiday before and after diagnosis and treatment peak period: 76000 in 2021, 78000 in 2022, and 80000 in 2023;
[0186] Special diagnosis and treatment activity period: 115000 in 2021, 118000 in 2022, and 122000 in 2023;
[0187] Calculate the data growth of adjacent stages:
[0188] Daily diagnosis and treatment data period-holiday before and after diagnosis and treatment peak period: 28000;
[0189] Peak patient visits before and after holidays - special medical service campaigns: 40,000 cases;
[0190] The calculated average growth rate, 'a', is 34,000.
[0191] Set a buffer value b, which is 15% of a based on the fluctuation characteristics of medical data, i.e., b = 5100 records;
[0192] Final threshold range: [ab, a+b] = [28900, 39100];
[0193] Set the time interval t=7 days to divide the time window;
[0194] Calculate the increase in adjacent window data in April 2024:
[0195] 4.1-4.7 (Routine Clinic Hours): 51,000 - 4.8-4.14 (Post-Holiday Recovery Period, Corresponding to Post-Holiday Scenario): 80,000; Increase of 29,000, within the range of [28,900, 39,100], is determined to be the peak period before and after the holiday;
[0196] Example 3: Based on the business processes of a hospital, the functional boundaries of five core departments are defined, and a department-core business table is constructed:
[0197] Medical Affairs Department: Patient medical record management, treatment plan review, and medical quality data statistics;
[0198] Laboratory Department: Specimen receipt and registration, test procedure execution, test report generation and review;
[0199] Inpatient Department: Patient admission registration, inpatient medical records, and discharge settlement data compilation;
[0200] Pharmacy Department: Drug receiving management, prescription dispensing and verification, and drug inventory data statistics;
[0201] Radiology Department: Image examination appointment, image data acquisition, image report archiving and storage;
[0202] For each department's core business processes, specific business scenarios are broken down, the correspondence between "business objectives and data relationships" is clarified, and a set of "departments-business scenarios" is constructed:
[0203] Medical Affairs Department: The business scenario is "Patient Medical Record Data Archiving", with the goal of "completing the archiving of patient medical record data within 1 working day after patient discharge". The data association includes "Patient Medical Record Cover Page", "Detailed Treatment Record Sheet", and "Medical Record Review Log".
[0204] Labs: Business scenario "Lab report data review", goal is "Complete report data verification within 2 hours after testing is completed", data associations include "Lab specimen information table", "Lab result raw data", "Lab item cross-reference table";
[0205] Inpatient Department: Business scenario "Inpatient expense data reconciliation", goal is "Complete expense data verification within 3 working days before patient discharge", data associations include "Inpatient expense detail table", "Order execution record", "Expense exception feedback form";
[0206] Pharmacy: Business scenario "Pharmaceutical inventory data inventory", goal is "Complete pharmaceutical inventory data reconciliation 5 days before the end of the month", data associations include "Pharmaceutical warehouse entry form", "Pharmaceutical warehouse exit ledger", "Inventory warning record table";
[0207] Radiology Department: Business scenario "Imaging report data archiving", goal is "Archive to hospital data center within 1 hour after imaging examination is completed", data associations include "Imaging examination application form", "Imaging data meta information", "Imaging diagnosis report".
[0208] For the core data resources in the business scenarios of each department, extract structured / unstructured text content:
[0209] Medical Department: Extract "Patient medical record home page" text, "Patient ID: P202405001, Name: Zhang San, Gender: Male, Age: 56 years old, Admission date: 2024-05-08, Discharge date: 2024-05-15, Main diagnosis: Type 2 diabetes, Medical record audit status: Passed, Archiving person in charge: Dr. Li, Archiving time: 2024-05-16";
[0210] Labs: Extract "Lab report" text, "Specimen number: S2024051001, Patient ID: P202405002, Test item: Blood routine, White blood cell count: 5.8×10 9 / L, Red blood cell count: 4.5×10¹² / L, Platelet count: 230×10 9 / L, Test time: 2024-05-10 09:30, Audit time: 2024-05-10 10:15, Test physician: Dr. Wang";
[0211] Inpatient Department: Extract the text of the "Inpatient Fee Detail Form", "Patient ID: P202405003, Name: Li Si, Hospital Number: H202405003, Days of Hospitalization: 7 days, Bed Fee: 1400 yuan, Examination Fee: 800 yuan, Medicine Fee: 2100 yuan, Total Fee: 4800 yuan, Fee Verification Status: To be verified, Verification Responsible Person: Liu Nurse, Verification Time: 2024-05-14";
[0212] Pharmacy: Extract the text of the "Drug Warehouse Entry Form", "Warehouse Entry Form Number: R202405001, Drug Name: Metformin Hydrochloride Tablets, Specification: 0.5g x 30 tablets, Manufacturer: XX Pharmaceutical Factory, Warehouse Entry Quantity: 500 boxes, Warehouse Entry Date: 2024-05-05, Acceptance Person: Zhao Pharmacist, Current Warehouse Inventory Quantity: 800 boxes";
[0213] Radiology Department: Extract the text of the "Imaging Examination Application Form", "Application Form Number: I202405001, Patient ID: P202405004, Name: Wang Wu, Examination Project: Chest CT, Examination Site: Chest, Application Department: Respiratory Medicine Department, Application Doctor: Sun Doctor, Examination Time: 2024-05-12 14:00, Report Archiving Time: 2024-05-12 15:30".
[0214] Use NLG to extract text keywords:
[0215] Determine the source of the NLG text:
[0216] Laboratory NLG text: "In May 2024, the laboratory handled 3000 test specimens, with a 'blood routine test report' audit completion rate of 100%, an average audit time of 40 minutes, and an abnormal result feedback timeliness rate of 98%";
[0217] Pharmacy NLG text: "In May 2024, the pharmacy completed the inventory data check of the drug stock, involving 200 kinds of drugs and a total of 5000 boxes of inventory, among which 'Metformin Hydrochloride Tablets' had 3 inventory warnings and had completed restocking, with an inventory reconciliation completion rate of 99.8%".
[0218] Keyword extraction:
[0219] Laboratory NLG text keywords: May 2024, test specimens, blood routine test report, audit completion rate 100%, abnormal result feedback;
[0220] Pharmacy NLG text keywords: May 2024, drug inventory check, metformin hydrochloride tablets, inventory warning, reconciliation completion rate 99.8%.
[0221] Split the keywords into word units using the Jieba segmentation tool:
[0222] Business scenario keywords: "patient medical record archiving" segmentation results: patients, medical records, archiving;
[0223] Text data keywords: "inspection report data audit" segmentation results: inspection, report, data, audit;
[0224] NLG text keywords: "drug inventory check" segmentation results: drugs, inventory, check;
[0225] Remove special symbols and low-frequency words from medical text, and keep core semantic words:
[0226] Original segmentation: inspection, report, data, audit, (2024), Dr. Wang, after screening: inspection, report, data, audit;
[0227] Original segmentation: drugs, inventory, check, metformin hydrochloride tablets, Mr. Zhao, after screening: drugs, inventory, check, metformin hydrochloride tablets;
[0228] Set the same sequence length to 6 based on the average number of words in medical keywords:
[0229] Shorter than 6 words: fill with white space;
[0230] Longer than 6 words: truncated to 6 words;
[0231] Build a global word table:
[0232] Collect word units: summarize all business scenario keywords and text data keywords after screening, a total of 32, including "patients, medical records, archiving, inspection, report, data, audit, hospitalization expenses, medical order execution, drugs, inventory, check" and others;
[0233] De-duplication and filtering: remove duplicate words and filter low-frequency words, finally keep 25 words;
[0234] Assign ID and index: assign a unique ID to each word to build a global word table;
[0235] 1-patient-0, 2-medical record-1, 3-archiving-3, 4-inspection-3, 5-report-4, 6-data-5, 7-hospitalization expenses-6, 8-medical order execution-7, 9-drugs-8, 10-inventory-9;
[0236] Based on the characteristics of medical data, set "number of associated business scenarios + data volume" dual-dimension to determine complexity, and the corresponding word vector dimensions are as follows:
[0237] Number of related business scenarios: 1-2, data volume interval: 1000-3000, word vector dimension: 50-100, keyword: patient medical record;
[0238] Number of related business scenarios: 3-5, data volume interval: 3000-8000, word vector dimension: 100-200, keyword: test report data;
[0239] Number of related business scenarios: more than 5, data volume interval: more than 8000, word vector dimension: 200-300, keyword: data;
[0240] Construct a word vector matrix:
[0241] Word ID1 (patient, 100 dimensions): [0.123, 0.456, 0.789, …, 0.321] (100 floating-point numbers in total);
[0242] Word ID2 (data, 200 dimensions): [0.234, 0.567, 0.890, …, 0.432] (200 floating-point numbers in total);
[0243] Construct a "25x200" word vector matrix;
[0244] Construct training samples:
[0245] Positive samples: business scenario keywords and matching text data keywords (50 groups in total), such as ("enterprise registration", "enterprise registration data") and ("social security payment", "social security payment base");
[0246] Negative samples: business scenario keywords and non-matching text data keywords (50 groups in total), such as ("public data sharing", "household relocation") and ("enterprise registration", "social security payment");
[0247] Form 100 groups of training samples, with a positive to negative ratio of 1:1;
[0248] Train the model using a ternary loss function;
[0249] Then calculate the final similarity:
[0250] Patient medical record filing - patient medical record is 0.98, test report data review - test report is 0.85, drug inventory check - drug inventory is 0.82, patient medical record filing - image report is 0.15, and test report review - hospitalization cost is 0.12;
[0251] Calculate the similarity threshold: the average of the similarity of 100 groups of training samples is 0.72, which is set as the "medical public data matching threshold";
[0252] When the keyword pair is greater than 0.72, it is determined that the resource data conforms to the department business scene;
[0253] It is apparent for those skilled in the art that the present application is not limited to the details of the foregoing exemplary embodiments, and the present application can be implemented in other concrete forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be considered in all aspects as illustrative and not restrictive, and the scope of the present application is defined by the appended claims rather than the above description, and it is intended to encompass all changes falling within the meaning and range of equivalents of the essential elements of the claims.
Claims
1. A method for dynamic allocation of resources for differentiated data, characterized in that: The method comprises the following steps: Step S1, dividing the laboratory operation stage according to the data growth amount, setting the data priority of each operation stage for each operation stage data; Step S1-1, collecting the data amount data of each laboratory operation stage in the operation of the historical multiple same scale laboratories, calculating the data amount growth between each stage by subtracting the total data amount of the adjacent two stages; Step S1-2, calculating the average of the data amount growth between each stage, setting the average as a threshold value a, adding a buffer value b to the threshold value a, combining the threshold value and the buffer value to set the threshold interval [a-b, a+b], a is the threshold value between each operation stage, b is the buffer value set according to the business scene of each operation stage; Step S1-3, setting a time interval t for the laboratory operation cycle, dividing the time window according to the time interval t, calculating the data growth of each adjacent time window, when the growth of two adjacent time windows is in the threshold interval, determining the operation stage of the adjacent time window according to different threshold intervals; Step S1-4, collecting historical data of each operation stage, cleaning the data, specifically: For missing values, mean filling is used, and for outliers, the 3σ principle is used to identify data that exceeds as outliers, as the mean, as the standard deviation, after which the outliers are removed, and the repeated values are removed; Step S1-5, setting the data priority according to the data attribute, specifically: Step S1-5-1, dividing the data attribute into three dimensions of data processing time consumption, latest delivery time and current operation stage core target; Step S1-5-2, quantifying the parameters to the interval [1, x] using the quantitative index method for the three dimensions of data processing time consumption, latest delivery time and current operation stage core target; Step S1-5-3, calculating the priority score of each type of data using the weighted average method, specifically: ; In the formula, P is a priority value of data, A is a quantized value of data processing time consumption, B is a quantized value of the latest delivery time, C is a quantized value of core target matching degree, is a dimension weight of data processing time consumption, is a dimension weight of the latest delivery time, is a dimension weight of core target. Step S1-5-4, according to the priority score of the data, sorting the priority score in ascending order and dividing the threshold interval [x1, x2], [x2, x3], [x3, x4] using the quartile method, setting the data priority according to the threshold interval where the priority value of each type of data is located; Step S2, processing the data according to the data priority and calling the resources; Step S3, dividing and matching the resource data according to the business scene of each department of the laboratory; Step S4, for different operation stages, constructing a knowledge graph to allocate resources, using the knowledge graph to allocate resources for each operation stage; Step S5, analyzing and evaluating the resource allocation results after adjusting the parameters, evaluating when the evaluation is not up to standard, analyzing the reasons and proposing solutions.
2. The method of claim 1, wherein: The specific steps of step S2 are as follows: Step S2-1, marking the data with priority for each operation stage data priority, outputting a structured data set of "data + priority"; Step S2-2, sorting the structured data set of "data + priority" in ascending order and classifying the data according to the priority to construct structured data sets of different priorities; Step S2-3, sorting the structured data in the structured data set in descending order according to the P value for the structured data set of different priorities, the system processes the data from left to right in the structured data set of high level first.
3. The method of claim 2, wherein: The specific steps of step S3 are as follows: Step S3-1, divide departments according to different business processes of the laboratory, construct a department table, extract business scenarios based on the business processes of each department, and clarify the business targets and data correlation of each business scenario to construct a "department-business scenario" set; Step S3-2, set the words with the highest frequency and containing professional terms in the department business process as keywords, extract multiple keywords of the business scenario, and construct a "department-business scenario keyword" set; Step S3-3, text extraction is performed on the resources in the business scenario; Step S3-4, set the text related to the business process as the keyword, and use a semantic recognition model to extract the keyword from the text data converted by the NLG technology; Step S3-5, calculate the similarity between the business scenario keywords and the text data keywords, specifically: Step S3-5-1, use the word segmentation tool Jieba to split the phrase keywords into word units; Step S3-5-2, remove mood auxiliaries, transition words, special symbols, and low-frequency words, and retain core semantic words; Step S3-5-3, set the keyword sequence length, unify the keyword sequence length, fill in the blank for sequences shorter than the keyword sequence length, and truncate sequences longer than the keyword sequence length; Step S3-5-4, collect all the word units of the business scenario keywords and the text data keywords, generate an original word set, remove duplicates from the original word set, filter low-frequency words, and construct a global word table, which assigns a unique ID and index to each word; Step S3-5-5, set the data size of the keyword as the complexity of the keyword, set the word vector dimension according to the complexity of the keyword, establish a matching relationship between "keyword complexity-word vector dimension", set the word vector dimension to 50-100 for keywords associated with only 1-2 business scenarios and data size between x1-x2, set the dimension to 100-200 for keywords associated with 3-5 business scenarios and data size between x2-x3, and set the dimension to 200-300 for keywords associated with more than 5 business scenarios and data size exceeding x3, x1, x2, and x2 are interval thresholds for dividing the complexity of keywords according to data size, and a fixed-dimension word vector is randomly generated for each word in the global word table, and a word vector matrix is constructed; Step S3-5-6, extract any two business scenario keywords and text data keywords to form a keyword pair, set the business scenario keywords and matching text data keyword pairs as positive samples, and set the business scenario keywords and non-matching text data keyword pairs as negative samples to construct training samples; Step S3-5-7, use a triple loss function to construct a training model, specifically: ; In the formula, L is a loss value, is a max operation, is a similarity calculation function, q is an encoded semantic vector, is a negative sample vector, is a positive sample vector, is a control positive sample pair similarity and negative sample pair similarity minimum gap, is a threshold value artificially set by a professional according to a business scenario. Step S3-5-8, encode the keywords according to the training model, and then calculate the similarity of the keyword pairs, specifically: ; In the formula, is the similarity obtained by final calculation, is a business scenario keyword, is a text data keyword, is a module length of the business scenario keyword, is a module length of the text data keyword; Step S3-6, take the average of the keyword pairs as the similarity threshold, and determine that the resource data meets the business scenario of the department when the threshold is exceeded.
4. The method of claim 3, wherein: The specific steps of step S4 are as follows: Step S4-1, one-level classification of resources is performed, specifically hardware resources and software resources, two-level classification of hardware resources is performed, specifically computing resources and storage resources, the specific entities of computing resources are CPU computing power, GPU computing power and memory capacity, the specific entities of storage resources are storage capacity and storage device read-write speed, and the specific entities of software resources are the number of authorized data processing tools and the frequency of calling data processing tools; Step S4-2, for each operation stage of the laboratory, the historical resource quantity of each operation stage in the historical operation process of the same scale is collected, and the historical resource is data cleaned using the data cleaning method of step S1-4; Step S4-3, according to the three-layer ontology structure defined in step S4-1, that is, the "one-level category-two-level category-specific entity" three-layer ontology structure, as a mapping skeleton, each cleaned resource is traversed, and CPU computing power, GPU computing power, memory capacity, storage capacity, storage device read-write speed, the number of authorized data processing tools and the frequency of calling data processing tools are mapped to the skeleton according to the upper level, and a knowledge graph is constructed; Step S4-4, the historical resource quantity of each operation stage in the knowledge graph is extracted as a basis for allocating resources.
5. The method of claim 4, wherein: The specific steps of step S5 are as follows: Step S5-1, the core evaluation index of the resource allocation result is divided into data processing timeliness, resource utilization rate and department demand satisfaction degree; Step S5-2, the data processing timeliness, resource utilization rate and department demand satisfaction degree are quantified, specifically: Data processing timeliness: ; In the formula, when T act ≤T exp , S tim ≥1, is the unified value range, and finally min(S tim ,1), is the anti-zero correction term, is the quantization value of data processing timeliness, is the actual total time consumption of a certain type of data processing task, is the expected total time consumption of the task. Resource utilization rate: ; In the formula, is a resource utilization rate quantitative value, is a device resource weight, is a personnel resource weight, is a resource device actual running time length, is a device monthly total available time length, is a personnel actual input work time length, is a personnel monthly rated total time length; Departmental demand satisfaction: ; In the formula, is the quantitative value of the department demand satisfaction, n is the resource demand quantity proposed by the department, is the weight of the jth demand, is the actual satisfaction amount of the jth demand, is the total amount of the jth demand applied by the department, is the zero prevention correction term; Step S5-3, the quantified values of data processing timeliness, resource utilization rate and department demand satisfaction degree are standardized; Step S5-4, the quantified values of data processing timeliness, resource utilization rate and department demand satisfaction degree are integrated, specifically: ; In the formula, is a core evaluation index of the resource allocation result, is a weight of data processing timeliness, is a weight of resource utilization rate, is a weight of department demand degree, is a quantitative value of data processing timeliness, is a quantitative value of resource utilization rate, is a quantitative value of department demand satisfaction degree; Step S5-5, for the core evaluation index, the core evaluation index is set to reach the threshold value according to the business scenario, when the core evaluation index exceeds the threshold value, it is determined that the resource allocation result is qualified, and the resource allocation result is solidified as the standard result of the operation stage, which is used for subsequent data processing of the same stage, records the running data corresponding to the resource allocation result, and stores it in the historical sample library; Step S5-6, when the core evaluation index does not reach the standard, three sub-index threshold values preset by the business scenario are extracted, the three sub-index threshold values are set by professionals according to the business scenario, and each sub-index is compared with the corresponding sub-index threshold value, and the reason for not reaching the standard is divided into: When S tim-comp tim-threshold is determined as not meeting the data processing timeliness, S tim-comp is the standardized data processing timeliness quantitative value, S tim-threshold is the threshold value of data processing timeliness: specific analysis is that the hardware resources CPU, GPU computing power and storage capacity are insufficient, the data processing tool is authorized, the calling frequency is idle or overloaded, and resources are added according to the analysis of system log and software monitoring platform; When S res-comp <S res-threshold is determined as not meeting the resource utilization rate, S res-comp is the standardized resource utilization rate quantitative value, S res-threshold The threshold value of the resource utilization rate meets the standard: check the input data of the model to see if there is data deviation; When S dmd-comp < S dmd-threshold is determined as the department demand satisfaction is not up to standard, S dmd-comp is the department demand satisfaction quantitative value after standardization, S dmd-threshold is the threshold value of department demand satisfaction: use step S3 to recalculate the matching.
6. A system for differentiated data oriented resource dynamic allocation, for performing the differentiated data oriented resource dynamic allocation method of any one of claims 1-5, characterized in that: The system comprises an operation stage and data priority division module, a resource calling processing module, a resource data matching module, a resource allocation module and a resource allocation result evaluation and adjustment module; The operation stage and data priority division module is used for dividing each operation stage of the laboratory operation, and setting the data priority of each operation stage according to the data attribute; The resource calling processing module is used for generating a structured set according to the data priority, and then calling the resources to process the data; The resource data matching module is used for dividing laboratory departments, extracting business scenarios of each department, extracting keywords and matching resource data; The resource allocation module is used for constructing a resource allocation knowledge graph and allocating resources for each operation stage according to the knowledge graph; The resource allocation result evaluation and adjustment module is used for evaluating resource allocation results, analyzing reasons for unqualified conditions and proposing solutions.
7. The differentiated data oriented resource dynamic allocation system of claim 6, wherein: The operation stage and data priority division module includes an operation stage division unit and a data priority setting unit; The operation stage division unit is used for calculating a threshold interval according to a historical data growth amount, dividing operation stages in combination with a time window and performing data cleaning on historical data; The data priority setting unit is used for quantifying data attributes, calculating a priority score through a weighting algorithm, and then dividing a threshold interval to set a data priority; The resource calling processing module includes a data priority structured set generation unit and a priority data processing unit; The data priority structured set generation unit is used for marking priorities of data of each stage and outputting a structured data set of "data + priority"; The priority data processing unit is used for classifying and sorting the structured set, and then processing the set data in order of priority from high to low according to a P value descending order.
8. The differentiated data oriented resource dynamic allocation system of claim 7, wherein: The resource data matching module includes a business scenario keyword extraction unit and a keyword similarity matching unit; The business scenario keyword extraction unit is used for constructing a "department-business scenario keyword" set and extracting text data keywords; The keyword similarity matching unit is used for calculating the similarity of business scenarios and text data keywords, and determining whether the resource data meets the department business scenario according to a threshold value; The resource allocation module includes a knowledge graph construction unit and a knowledge graph resource allocation unit; The knowledge graph construction unit is used for dividing resource categories, cleaning historical resource data, and constructing a knowledge graph by mapping resources to a skeleton; The knowledge graph resource allocation unit is used for extracting historical resource amounts of each operation stage of the knowledge graph, and allocating resources based on the reference.
9. The differentiated data oriented resource dynamic allocation system of claim 8, wherein: The resource allocation result evaluation and adjustment module includes a resource allocation result evaluation unit and an evaluation unqualified adjustment unit; The resource allocation result evaluation unit is used for quantifying core evaluation indicators, standardizing and integrating the indicators, and setting a threshold to determine whether the resource allocation parameters are qualified; The evaluation unqualified adjustment unit is used for comparing sub-indicators with a threshold value, dividing unqualified reasons and proposing data processing, resource adjustment or re-matching solutions.
Citation Information
Patent Citations
Data monitoring method, apparatus, computing device, and storage medium
CN109241133A
Intelligent resource scheduling optimization method based on artificial intelligence
CN120452714A