Server scheduling method and system based on feature analysis of sequencing data

By analyzing the characteristic analysis and processing difficulty evaluation of cfDNA sequencing data and patient parameters, the most suitable server is selected to perform cancer prediction tasks, which solves the problem of improper server selection in the prior art and improves the computing efficiency and accuracy of cancer prediction.

CN120565010AInactive Publication Date: 2025-08-29SHENZHEN RAPHA BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511055443.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-08-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art in cancer prediction has resulted in improper server selection, resulting in inefficient computing efficiency, waste of resources and prediction delays due to the lack of in-depth analysis of characteristic parameters of cfDNA sequencing data and dynamic evaluation of the difficulty of processing.

Method used

By obtaining cfDNA sequencing data and patient parameters, using feature analysis algorithms to divide data groups and calculate data specific parameters, combining patient parameters and history records to predict processing difficulty, thereby selecting the most suitable server to perform cancer prediction tasks.

Benefits of technology

Accurate server allocation based on feature analysis and processing difficulty assessment is realized, which improves the computing efficiency and prediction accuracy of cancer prediction tasks, and reduces the prediction delay and resource waste risks caused by improper server selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120565010A_ABST
    Figure CN120565010A_ABST
Patent Text Reader

Abstract

The invention discloses a server scheduling method and system based on feature analysis of sequencing data. The method comprises the following steps: acquiring cfDNA sequencing data to be subjected to cancer prediction and corresponding patient parameters; based on a feature analysis algorithm, analyzing data feature parameters corresponding to the cfDNA sequencing data; predicting the prediction processing difficulty corresponding to the cfDNA sequencing data according to the patient parameters and the data characteristic parameters; determining a target server for executing a prediction task from a plurality of candidate prediction servers according to the prediction processing difficulty; and the target server is used for receiving the cfDNA sequencing data and executing a cancer prediction task based on a built-in algorithm model. Therefore, accurate server allocation based on feature analysis and processing difficulty evaluation can be realized, the calculation efficiency and prediction accuracy of a cancer prediction task are improved, and the risk of prediction delay or resource waste caused by improper server selection is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a server scheduling method and system based on feature analysis of sequencing data. Background Art

[0002] Existing technologies typically collect cfDNA sequencing data and patient parameters, select prediction servers using fixed server allocation rules or simple load balancing methods, and execute cancer prediction tasks based on standard algorithms to support diagnostic analysis. Due to a lack of in-depth analysis of data characteristics and dynamic assessment of processing difficulty, existing solutions struggle to accurately select appropriate servers for complex prediction tasks, resulting in inefficient computing and wasted resources. Inappropriate server selection can also lead to prediction delays or inaccuracies, limiting the performance and reliability of cancer prediction services. Clearly, existing technologies have shortcomings that urgently need to be addressed. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a server scheduling method and system based on feature analysis of sequencing data, which can achieve accurate server allocation based on feature analysis and processing difficulty assessment, improve the computational efficiency and prediction accuracy of cancer prediction tasks, and reduce the risk of prediction delays or resource waste caused by improper server selection.

[0004] In order to solve the above technical problems, the first aspect of the present invention discloses a server scheduling method based on feature analysis of sequencing data, the method comprising: Obtain cfDNA sequencing data and corresponding patient parameters for cancer prediction; Analyzing data characteristic parameters corresponding to the cfDNA sequencing data based on a feature analysis algorithm; Predicting the prediction processing difficulty corresponding to the cfDNA sequencing data based on the patient parameters and the data characteristic parameters; Based on the difficulty of the prediction processing, a target server for performing the prediction task is determined from multiple candidate prediction servers; the target server is used to receive the cfDNA sequencing data and perform the cancer prediction task based on a built-in algorithm model.

[0005] As an optional embodiment, in the first aspect of the present invention, analyzing the data characteristic parameters corresponding to the cfDNA sequencing data based on the feature analysis algorithm includes: Based on a preset segmentation rule, the cfDNA sequencing data is divided into multiple data groups; Calculating data-specific parameters corresponding to each of the data sets; Screening out data groups whose data-specific parameters are greater than a preset parameter threshold to obtain multiple specific data groups; The sequence position of each specific data group in the cfDNA sequencing data and the corresponding data-specific parameters are determined as data feature parameters corresponding to the cfDNA sequencing data.

[0006] As an optional embodiment, in the first aspect of the present invention, the segmentation rule is a rule based on average sequence length, a rule based on different sequencing batches, or a rule based on a similarity clustering algorithm.

[0007] As an optional embodiment, in the first aspect of the present invention, the calculating the data-specific parameter corresponding to each of the data groups includes: For each of the data groups, determining separation fragment parameters corresponding to the data group; Determine the total amount of data corresponding to the data group; Calculating the reciprocal of the average value of the data similarity between the data group and two other data groups adjacent to each other in sequence position, to obtain the incoherence parameter corresponding to the data group; A weighted sum of the separation segment parameter, the total data volume, and the incoherence parameter is calculated to obtain a data-specific parameter corresponding to the data group.

[0008] As an optional embodiment, in the first aspect of the present invention, determining the separation fragment parameters corresponding to the data group includes: For each separation phenomenon type, determining the sequence fragments that have the separation phenomenon type among all sequence fragments in the data set to obtain separation fragments; the separation phenomenon type is AT separation phenomenon or GC separation phenomenon; Calculating the ratio of the total number of fragments of all the separated fragments to the total number of all sequence fragments in the data set; Calculate the weighted sum average of the data volume of the separation fragments corresponding to all the separation phenomenon types to obtain the separation fragment parameters corresponding to the data group; wherein the weighted calculation weight of the data volume of the separation fragment corresponding to each separation phenomenon type is proportional to the corresponding ratio.

[0009] As an optional embodiment, in the first aspect of the present invention, predicting the prediction processing difficulty corresponding to the cfDNA sequencing data based on the patient parameters and the data characteristic parameters includes: Inputting the patient parameters into a trained disease prediction model to obtain an output predicted disease type; the disease prediction model is trained by a training data set including a plurality of training patient parameters and corresponding disease annotations; Determining historical prediction records corresponding to the predicted disease type from a preset historical prediction database, and calculating an average of the prediction times of all the historical prediction records to obtain a historical time parameter; Inputting the cfDNA sequencing data and the data feature parameters into a trained processing cost prediction model to obtain an output predicted processing cost; the processing cost prediction model is trained by a training data set including a plurality of training sequencing data and corresponding data feature parameter annotations and processing cost annotations; The sum of the historical time parameter and the predicted processing cost is calculated to obtain the predicted processing difficulty corresponding to the cfDNA sequencing data.

[0010] As an optional embodiment, in the first aspect of the present invention, the patient parameters include the patient's physiological parameters, the patient's historical medical records, and the patient's medical institution.

[0011] As an optional embodiment, in the first aspect of the present invention, determining a target server for performing the prediction task from a plurality of candidate prediction servers based on the prediction processing difficulty includes: For each candidate prediction server, obtaining a historical prediction processing record of the candidate prediction server; Determining the device performance of the candidate prediction server based on the processing device performance in the historical prediction processing record; Filtering out records in the historical prediction processing records whose corresponding processing difficulty is greater than or equal to the predicted processing difficulty, to obtain a plurality of matching records; Calculate the average of the predicted processing times of all the matching records to obtain the historical predicted processing time; Calculating the product of the device performance and the historical prediction time to obtain the priority of the candidate prediction server; The candidate prediction server with the highest priority is determined as the target server for executing the prediction task.

[0012] A second aspect of an embodiment of the present invention discloses a server scheduling system based on feature analysis of sequencing data, the system comprising: An acquisition module is used to obtain cfDNA sequencing data and corresponding patient parameters for cancer prediction; An analysis module, configured to analyze data characteristic parameters corresponding to the cfDNA sequencing data based on a characteristic analysis algorithm; A prediction module, configured to predict a prediction processing difficulty corresponding to the cfDNA sequencing data based on the patient parameters and the data characteristic parameters; A scheduling module is used to determine a target server for performing the prediction task from multiple candidate prediction servers based on the prediction processing difficulty; the target server is used to receive the cfDNA sequencing data and perform the cancer prediction task based on a built-in algorithm model.

[0013] As an optional embodiment, in the second aspect of the present invention, the specific manner in which the analysis module analyzes the data characteristic parameters corresponding to the cfDNA sequencing data based on the characteristic analysis algorithm includes: Based on a preset segmentation rule, the cfDNA sequencing data is divided into multiple data groups; Calculating data-specific parameters corresponding to each of the data sets; Screening out data groups whose data-specific parameters are greater than a preset parameter threshold to obtain multiple specific data groups; The sequence position of each specific data group in the cfDNA sequencing data and the corresponding data-specific parameters are determined as data feature parameters corresponding to the cfDNA sequencing data.

[0014] As an optional embodiment, in the second aspect of the present invention, the segmentation rule is a rule based on average sequence length, a rule based on different sequencing batches, or a rule based on a similar clustering algorithm.

[0015] As an optional embodiment, in the second aspect of the present invention, the specific manner in which the analysis module calculates the data-specific parameters corresponding to each of the data groups includes: For each of the data groups, determining separation fragment parameters corresponding to the data group; Determine the total amount of data corresponding to the data group; Calculating the reciprocal of the average value of the data similarity between the data group and two other data groups adjacent to each other in sequence position, to obtain the incoherence parameter corresponding to the data group; A weighted sum of the separation segment parameter, the total data volume, and the incoherence parameter is calculated to obtain a data-specific parameter corresponding to the data group.

[0016] As an optional embodiment, in the second aspect of the present invention, the specific manner in which the analysis module determines the separation fragment parameters corresponding to the data set includes: For each separation phenomenon type, determining the sequence fragments that have the separation phenomenon type among all sequence fragments in the data set to obtain separation fragments; the separation phenomenon type is AT separation phenomenon or GC separation phenomenon; Calculating the ratio of the total number of fragments of all the separated fragments to the total number of all sequence fragments in the data set; Calculate the weighted sum average of the data volume of the separation fragments corresponding to all the separation phenomenon types to obtain the separation fragment parameters corresponding to the data group; wherein the weighted calculation weight of the data volume of the separation fragment corresponding to each separation phenomenon type is proportional to the corresponding ratio.

[0017] As an optional embodiment, in the second aspect of the present invention, the prediction module predicts the specific manner of the prediction processing difficulty corresponding to the cfDNA sequencing data based on the patient parameters and the data characteristic parameters, including: Inputting the patient parameters into a trained disease prediction model to obtain an output predicted disease type; the disease prediction model is trained by a training data set including a plurality of training patient parameters and corresponding disease annotations; Determining historical prediction records corresponding to the predicted disease type from a preset historical prediction database, and calculating an average of the prediction times of all the historical prediction records to obtain a historical time parameter; Inputting the cfDNA sequencing data and the data feature parameters into a trained processing cost prediction model to obtain an output predicted processing cost; the processing cost prediction model is trained by a training data set including a plurality of training sequencing data and corresponding data feature parameter annotations and processing cost annotations; The sum of the historical time parameter and the predicted processing cost is calculated to obtain the predicted processing difficulty corresponding to the cfDNA sequencing data.

[0018] As an optional embodiment, in the second aspect of the present invention, the patient parameters include the patient's physiological parameters, the patient's historical medical records, and the patient's medical institution.

[0019] As an optional embodiment, in the second aspect of the present invention, the scheduling module determines a target server for executing the prediction task from a plurality of candidate prediction servers according to the prediction processing difficulty, including: For each candidate prediction server, obtaining a historical prediction processing record of the candidate prediction server; Determining the device performance of the candidate prediction server based on the processing device performance in the historical prediction processing record; Filtering out records in the historical prediction processing records whose corresponding processing difficulty is greater than or equal to the predicted processing difficulty, to obtain a plurality of matching records; Calculate the average of the predicted processing times of all the matching records to obtain the historical predicted processing time; Calculating the product of the device performance and the historical prediction time to obtain the priority of the candidate prediction server; The candidate prediction server with the highest priority is determined as the target server for executing the prediction task.

[0020] A third aspect of the present invention discloses another server scheduling system based on feature analysis of sequencing data, the system comprising: a memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute part or all of the steps in the server scheduling method based on feature analysis of sequencing data disclosed in the first aspect of the present invention.

[0021] The fourth aspect of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute some or all of the steps in the server scheduling method based on feature analysis of sequencing data disclosed in the first aspect of the present invention.

[0022] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: The present invention obtains cfDNA sequencing data and patient parameters, analyzes data characteristic parameters and predicts processing difficulty based on patient parameters and data characteristic parameters, so as to select a target server from candidate prediction servers based on the predicted processing difficulty to perform cancer prediction tasks, thereby achieving accurate server allocation based on feature analysis and processing difficulty assessment, improving the computational efficiency and prediction accuracy of cancer prediction tasks, and reducing the risk of prediction delays or resource waste due to improper server selection. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1 This is a flow chart of a server scheduling method based on feature analysis of sequencing data disclosed in an embodiment of the present invention.

[0025] Figure 2 It is a structural diagram of a server scheduling system based on feature analysis of sequencing data disclosed in an embodiment of the present invention.

[0026] Figure 3 This is a structural diagram of another server scheduling system based on feature analysis of sequencing data disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0028] The terms "first," "second," and so on, in the description and claims of the present invention and the accompanying drawings are used to distinguish between different objects, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or device.

[0029] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0030] The present invention discloses a server scheduling method and system based on feature analysis of sequencing data. By acquiring cfDNA sequencing data and patient parameters, analyzing the data feature parameters, and predicting processing difficulty based on the patient parameters and data feature parameters, a target server is selected from candidate prediction servers based on the predicted processing difficulty to perform the cancer prediction task. This method achieves precise server allocation based on feature analysis and processing difficulty assessment, improves the computational efficiency and prediction accuracy of cancer prediction tasks, and reduces the risk of prediction delays or resource waste caused by improper server selection. These are described in detail below.

[0031] Example 1 See also Figure 1 , Figure 1 This is a flow chart of a server scheduling method based on feature analysis of sequencing data disclosed in an embodiment of the present invention. Figure 1 The server scheduling method based on feature analysis of sequencing data described above can be applied to a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). Figure 1As shown, the server scheduling method based on feature analysis of sequencing data may include the following operations: 101. Obtain cfDNA sequencing data and corresponding patient parameters for cancer prediction.

[0032] Optionally, the cfDNA sequencing data may include whole genome sequencing data, target region sequencing data or exome sequencing data, which is not limited in the present invention.

[0033] Optionally, the patient parameters may include age, gender, medical history, family history or medical institution information, which is not limited in the present invention.

[0034] Optionally, the acquisition process can be implemented based on database query, clinical data collection, third-party interface or real-time uploading, which is not limited by the present invention.

[0035] 102. Based on the feature analysis algorithm, analyze the data feature parameters corresponding to the cfDNA sequencing data. Optionally, the feature analysis algorithm may be a statistical analysis algorithm, a sequence feature extraction algorithm, or a machine learning algorithm, which is not limited in the present invention.

[0036] Optionally, the data feature parameters may include sequence distribution features, separation phenomenon features or variation features, which are not limited in the present invention.

[0037] Optionally, the analysis process can be optimized in combination with sequencing depth, data quality, or cancer type, which is not limited in the present invention.

[0038] 103. Predict the difficulty of predictive processing corresponding to cfDNA sequencing data based on patient parameters and data characteristic parameters. Optionally, the prediction processing difficulty may be a computational complexity, a time complexity or a resource consumption index, which is not limited in the present invention.

[0039] Optionally, the prediction process can be implemented based on a regression model, a classification model, or a hybrid model, which is not limited in the present invention.

[0040] 104. Determine a target server for executing the prediction task from multiple candidate prediction servers based on the prediction processing difficulty.

[0041] Optionally, the target server is used to receive cfDNA sequencing data and perform cancer prediction tasks based on a built-in algorithm model.

[0042] Optionally, the target server may be a local server, a cloud server, or a distributed computing node, which is not limited in the present invention.

[0043] Optionally, the built-in algorithm model may be a deep learning model, a statistical model, or a rule-based reasoning model, which is not limited in the present invention.

[0044] It can be seen that the above-mentioned embodiment of the invention obtains cfDNA sequencing data and patient parameters, analyzes data feature parameters and predicts processing difficulty based on patient parameters and data feature parameters, so as to select a target server from candidate prediction servers based on the predicted processing difficulty to perform cancer prediction tasks, thereby achieving accurate server allocation based on feature analysis and processing difficulty assessment, improving the computational efficiency and prediction accuracy of cancer prediction tasks, and reducing the risk of prediction delays or resource waste due to improper server selection.

[0045] As an optional embodiment, in the above step, analyzing the data feature parameters corresponding to the cfDNA sequencing data based on the feature analysis algorithm includes: Based on the preset segmentation rules, the cfDNA sequencing data is divided into multiple data groups; Calculate the data-specific parameters corresponding to each data set; Screening out data groups whose data-specific parameters are greater than a preset parameter threshold to obtain multiple specific data groups; The sequence position of each specific data group in the cfDNA sequencing data and the corresponding data-specific parameters are determined as the data characteristic parameters corresponding to the cfDNA sequencing data.

[0046] Optionally, the data group may be a constant-length sequence group, a variable-length sequence group, or a dynamic grouping, which is not limited in the present invention.

[0047] Optionally, the division process may be implemented based on sequence segmentation, cluster analysis or data normalization, which is not limited in the present invention.

[0048] Optionally, the parameter threshold may be a fixed threshold, a dynamic threshold, or a threshold adjusted based on data distribution, which is not limited in the present invention.

[0049] Optionally, the screening process may be implemented based on threshold filtering, a sorting algorithm, or a classification model, which is not limited in the present invention.

[0050] Optionally, the data feature parameter may be a feature vector, a feature matrix, or a multi-dimensional feature set, which is not limited in the present invention.

[0051] Optionally, the process of determining the data feature parameters may be implemented based on feature splicing, data fusion, or position mapping, which is not limited in the present invention.

[0052] It can be seen that through the above optional embodiments, by dividing the cfDNA sequencing data based on segmentation rules and calculating data-specific parameters, the specific data group is screened to determine the data feature parameters, thereby improving the accuracy and pertinence of data feature extraction through data segmentation and specificity analysis on the basis of precise server allocation, providing high-quality feature support for processing difficulty prediction, and reducing the risk of server selection error caused by feature extraction bias.

[0053] As an optional embodiment, in the above steps, the segmentation rule is a rule based on average sequence length, a rule based on different sequencing batches, or a rule based on a similarity clustering algorithm.

[0054] It can be seen that through the above optional embodiments, the types of segmentation rules are limited to effectively achieve reasonable division of sequencing data, assist in achieving accurate server allocation based on feature analysis and processing difficulty assessment, improve the computational efficiency and prediction accuracy of cancer prediction tasks, and reduce the risk of prediction delays or resource waste caused by improper server selection.

[0055] As an optional embodiment, in the above step, calculating the data-specific parameters corresponding to each data group includes: For each data set, determining the separation fragment parameters corresponding to the data set; Determine the total amount of data corresponding to the data group; Calculate the reciprocal of the average value of the data similarity between the data group and the two other data groups adjacent to each other in the sequence position, and obtain the incoherence parameter corresponding to the data group; The weighted sum of the separation segment parameter, the total data volume and the incoherence parameter is calculated to obtain the data-specific parameter corresponding to the data group.

[0056] Optionally, the total data volume may be the number of sequence segments, the number of data bytes, or the standardized data volume, which is not limited in the present invention.

[0057] Optionally, the data similarity may be cosine similarity, Euclidean distance, or Jaccard coefficient, which is not limited in the present invention.

[0058] Optionally, the weighted sum value may be calculated using fixed weights, dynamic weights, or adaptive weights, which is not limited in the present invention.

[0059] It can be seen that through the above optional embodiments, by calculating the weighted sum of the separation fragment parameters, total data volume and incoherence parameters of the data group as data-specific parameters, the accuracy of data specificity assessment is improved through comprehensive analysis of multi-dimensional parameters on the basis of accurate data feature extraction, providing a reliable basis for the generation of data feature parameters and reducing the risk of feature deviation caused by incomplete data group analysis.

[0060] As an optional embodiment, in the above step, determining the separation fragment parameters corresponding to the data group includes: For each separation phenomenon type, determining the sequence fragments of the separation phenomenon type among all sequence fragments in the data set to obtain separation fragments; optionally, the separation phenomenon type is AT separation phenomenon or GC separation phenomenon; Calculate the ratio of the total number of fragments of all separated fragments to the total number of all sequence fragments in the data set; Calculate the weighted sum average of the data volume of the separation fragments corresponding to all separation phenomenon types to obtain the separation fragment parameters corresponding to the data group; wherein the weighted calculation weight of the data volume of the separation fragment corresponding to each separation phenomenon type is proportional to the corresponding ratio.

[0061] Optionally, the separation phenomenon type may include other phenomena, such as base bias phenomenon or sequence duplication phenomenon, which is not limited in the present invention.

[0062] Optionally, the process of determining the separated fragments may be implemented based on pattern recognition, sequence analysis or feature matching, which is not limited in the present invention.

[0063] It can be seen that through the above optional embodiments, by identifying the fragments with AT or GC separation phenomena in the data group and calculating the proportion of the number of fragments and the weighted average of the data volume as separation fragment parameters, the pertinence and accuracy of the parameters are improved through separation phenomenon analysis and weighted calculation on the basis of accurate data-specific parameter calculation, providing precise support for data feature extraction and reducing the risk of feature extraction errors caused by misjudgment of separation features.

[0064] As an optional embodiment, in the above step, predicting the prediction processing difficulty corresponding to the cfDNA sequencing data based on the patient parameters and the data characteristic parameters includes: Inputting the patient parameters into a trained disease prediction model to obtain an output predicted disease type; optionally, the disease prediction model is trained using a training dataset comprising a plurality of training patient parameters and corresponding disease annotations; Determine the historical prediction records corresponding to the predicted disease type from a preset historical prediction database, and calculate the average prediction time of all historical prediction records to obtain a historical time parameter; Inputting the cfDNA sequencing data and data feature parameters into a trained processing cost prediction model to obtain an output predicted processing cost; optionally, the processing cost prediction model is trained using a training data set including a plurality of training sequencing data and corresponding data feature parameter annotations and processing cost annotations; The sum of the historical time parameters and the predicted processing cost is calculated to obtain the predicted processing difficulty corresponding to the cfDNA sequencing data.

[0065] Optionally, the disease prediction model can be a classification model, a regression model or a deep learning model, which is not limited in the present invention.

[0066] Optionally, the predicted disease type may be a cancer type, a cancer stage, or a cancer risk level, which is not limited in the present invention.

[0067] Optionally, the training data set may include historical patient data, simulated data, or labeled data, which is not limited in the present invention.

[0068] Optionally, the historical prediction database may be a local database, a cloud database, or a distributed database, which is not limited in the present invention.

[0069] Optionally, the historical prediction record may include prediction task data, computing resource data, or prediction result data, which is not limited in the present invention.

[0070] Optionally, the processing cost prediction model may be a regression model, a neural network model or an integrated model, which is not limited in the present invention.

[0071] Optionally, the predicted processing cost may be a computing time cost, a resource consumption cost, or a comprehensive cost indicator, which is not limited in the present invention.

[0072] It can be seen that through the above optional embodiments, by inputting patient parameters into the disease prediction model to obtain the predicted disease type and combining historical prediction records to calculate historical time parameters, and combining the processing cost prediction model to output the predicted processing cost to determine the processing difficulty, the accuracy and comprehensiveness of the processing difficulty prediction are improved through disease prediction and cost evaluation on the basis of precise server allocation, providing a reliable basis for target server selection and reducing the risk of resource allocation caused by misjudgment of difficulty.

[0073] As an optional embodiment, in the above steps, the patient parameters include the patient's physiological parameters, the patient's historical medical records and the patient's medical institution.

[0074] It can be seen that through the above optional embodiments, the content of patient parameters is limited to comprehensively characterize the relevant characteristics of patients to be predicted for cancer, assist in achieving accurate server allocation based on feature analysis and processing difficulty assessment, improve the computational efficiency and prediction accuracy of cancer prediction tasks, and reduce the risk of prediction delays or resource waste due to improper server selection.

[0075] As an optional embodiment, in the above step, determining a target server for performing the prediction task from multiple candidate prediction servers based on the difficulty of the prediction process includes: For each candidate prediction server, obtaining a historical prediction processing record of the candidate prediction server; Determining the device performance of the candidate prediction server based on the processing device performance in the historical prediction processing record; Filter out the records whose corresponding processing difficulty is greater than or equal to the predicted processing difficulty in the historical predicted processing records, and obtain multiple matching records; Calculate the average predicted processing time of all matching records to obtain the historical predicted processing time; Calculate the product of device performance and historical prediction time to obtain the priority of the candidate prediction server; The candidate prediction server with the highest priority is determined as the target server for executing the prediction task.

[0076] Optionally, the historical prediction processing record may include processing task data, performance data, or time data, which is not limited in the present invention.

[0077] Optionally, the acquisition process of the historical prediction processing record can be implemented based on server logs, database queries or real-time monitoring, which is not limited in the present invention.

[0078] Optionally, the device performance may include computing power, memory capacity, storage speed or network bandwidth, which is not limited in the present invention.

[0079] Optionally, the process of determining the device performance may be implemented based on the extraction of performance indicators of the device processing process in the historical prediction processing records, statistical analysis, or hardware evaluation, which is not limited in the present invention.

[0080] It can be seen that through the above optional embodiments, by obtaining the historical processing records of candidate prediction servers and calculating the priority based on the device performance and the predicted time of the matching records, the server with the highest priority is selected as the target server. Therefore, on the basis of precise server allocation, the pertinence and efficiency of server selection are improved through performance and historical time analysis, providing optimized computing resources for cancer prediction tasks, and reducing the risk of prediction delays caused by insufficient server performance.

[0081] Example 2 See also Figure 2 , Figure 2 Schematic diagram of a server scheduling system based on feature analysis of sequencing data disclosed in an embodiment of the present invention. Figure 2 The server scheduling system based on feature analysis of sequencing data described above can be applied to a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). Figure 2 As shown, the server scheduling system based on feature analysis of sequencing data may include: The acquisition module 201 is used to obtain cfDNA sequencing data and corresponding patient parameters for cancer prediction.

[0082] The analysis module 202 is used to analyze the data characteristic parameters corresponding to the cfDNA sequencing data based on the characteristic analysis algorithm. The prediction module 203 is used to predict the prediction processing difficulty corresponding to the cfDNA sequencing data based on the patient parameters and data characteristic parameters. The scheduling module 204 is configured to determine a target server for executing the prediction task from a plurality of candidate prediction servers according to the difficulty of the prediction process.

[0083] Optionally, the target server is used to receive cfDNA sequencing data and perform cancer prediction tasks based on a built-in algorithm model.

[0084] It can be seen that the above-mentioned embodiment of the invention obtains cfDNA sequencing data and patient parameters, analyzes data feature parameters and predicts processing difficulty based on patient parameters and data feature parameters, so as to select a target server from candidate prediction servers based on the predicted processing difficulty to perform cancer prediction tasks, thereby achieving accurate server allocation based on feature analysis and processing difficulty assessment, improving the computational efficiency and prediction accuracy of cancer prediction tasks, and reducing the risk of prediction delays or resource waste due to improper server selection.

[0085] As an optional embodiment, the analysis module analyzes the data characteristic parameters corresponding to the cfDNA sequencing data based on the characteristic analysis algorithm in a specific manner, including: Based on the preset segmentation rules, the cfDNA sequencing data is divided into multiple data groups; Calculate the data-specific parameters corresponding to each data set; Screening out data groups whose data-specific parameters are greater than a preset parameter threshold to obtain multiple specific data groups; The sequence position of each specific data group in the cfDNA sequencing data and the corresponding data-specific parameters are determined as the data characteristic parameters corresponding to the cfDNA sequencing data.

[0086] It can be seen that through the above optional embodiments, by dividing the cfDNA sequencing data based on segmentation rules and calculating data-specific parameters, the specific data group is screened to determine the data feature parameters, thereby improving the accuracy and pertinence of data feature extraction through data segmentation and specificity analysis on the basis of precise server allocation, providing high-quality feature support for processing difficulty prediction, and reducing the risk of server selection error caused by feature extraction bias.

[0087] As an optional embodiment, the segmentation rule is a rule based on average sequence length, a rule based on different sequencing batches, or a rule based on a similarity clustering algorithm.

[0088] It can be seen that through the above optional embodiments, the types of segmentation rules are limited to effectively achieve reasonable division of sequencing data, assist in achieving accurate server allocation based on feature analysis and processing difficulty assessment, improve the computational efficiency and prediction accuracy of cancer prediction tasks, and reduce the risk of prediction delays or resource waste caused by improper server selection.

[0089] As an optional embodiment, the specific method in which the analysis module calculates the data-specific parameters corresponding to each data group includes: For each data set, determining the separation fragment parameters corresponding to the data set; Determine the total amount of data corresponding to the data group; Calculate the reciprocal of the average value of the data similarity between the data group and the two other data groups adjacent to each other in the sequence position, and obtain the incoherence parameter corresponding to the data group; The weighted sum of the separation segment parameter, the total data volume and the incoherence parameter is calculated to obtain the data-specific parameter corresponding to the data group.

[0090] It can be seen that through the above optional embodiments, by calculating the weighted sum of the separation fragment parameters, total data volume and incoherence parameters of the data group as data-specific parameters, the accuracy of data specificity assessment is improved through comprehensive analysis of multi-dimensional parameters on the basis of accurate data feature extraction, providing a reliable basis for the generation of data feature parameters and reducing the risk of feature deviation caused by incomplete data group analysis.

[0091] As an optional embodiment, the specific manner in which the analysis module determines the separation fragment parameters corresponding to the data group includes: For each separation phenomenon type, determining the sequence fragments of the separation phenomenon type among all sequence fragments in the data set to obtain separation fragments; optionally, the separation phenomenon type is AT separation phenomenon or GC separation phenomenon; Calculate the ratio of the total number of fragments of all separated fragments to the total number of all sequence fragments in the data set; Calculate the weighted sum average of the data volume of the separation fragments corresponding to all separation phenomenon types to obtain the separation fragment parameters corresponding to the data group; wherein the weighted calculation weight of the data volume of the separation fragment corresponding to each separation phenomenon type is proportional to the corresponding ratio.

[0092] It can be seen that through the above optional embodiments, by identifying the fragments with AT or GC separation phenomena in the data group and calculating the proportion of the number of fragments and the weighted average of the data volume as separation fragment parameters, the pertinence and accuracy of the parameters are improved through separation phenomenon analysis and weighted calculation on the basis of accurate data-specific parameter calculation, providing precise support for data feature extraction and reducing the risk of feature extraction errors caused by misjudgment of separation features.

[0093] As an optional embodiment, the prediction module predicts the specific method of the prediction processing difficulty corresponding to the cfDNA sequencing data based on the patient parameters and data characteristic parameters, including: Inputting the patient parameters into a trained disease prediction model to obtain an output predicted disease type; optionally, the disease prediction model is trained using a training dataset comprising a plurality of training patient parameters and corresponding disease annotations; Determine the historical prediction records corresponding to the predicted disease type from a preset historical prediction database, and calculate the average prediction time of all historical prediction records to obtain a historical time parameter; Inputting the cfDNA sequencing data and data feature parameters into a trained processing cost prediction model to obtain an output predicted processing cost; optionally, the processing cost prediction model is trained using a training data set including a plurality of training sequencing data and corresponding data feature parameter annotations and processing cost annotations; The sum of the historical time parameters and the predicted processing cost is calculated to obtain the predicted processing difficulty corresponding to the cfDNA sequencing data.

[0094] It can be seen that through the above optional embodiments, by inputting patient parameters into the disease prediction model to obtain the predicted disease type and combining historical prediction records to calculate historical time parameters, and combining the processing cost prediction model to output the predicted processing cost to determine the processing difficulty, the accuracy and comprehensiveness of the processing difficulty prediction are improved through disease prediction and cost evaluation on the basis of precise server allocation, providing a reliable basis for target server selection and reducing the risk of resource allocation caused by misjudgment of difficulty.

[0095] As an optional embodiment, the patient parameters include the patient's physiological parameters, the patient's historical medical records, and the patient's medical institution.

[0096] It can be seen that through the above optional embodiments, the content of patient parameters is limited to comprehensively characterize the relevant characteristics of patients to be predicted for cancer, assist in achieving accurate server allocation based on feature analysis and processing difficulty assessment, improve the computational efficiency and prediction accuracy of cancer prediction tasks, and reduce the risk of prediction delays or resource waste due to improper server selection.

[0097] As an optional embodiment, the scheduling module determines a target server for executing the prediction task from multiple candidate prediction servers according to the difficulty of the prediction process, including: For each candidate prediction server, obtaining a historical prediction processing record of the candidate prediction server; Determining the device performance of the candidate prediction server based on the processing device performance in the historical prediction processing record; Filter out the records whose corresponding processing difficulty is greater than or equal to the predicted processing difficulty in the historical predicted processing records, and obtain multiple matching records; Calculate the average predicted processing time of all matching records to obtain the historical predicted processing time; Calculate the product of device performance and historical prediction time to obtain the priority of the candidate prediction server; The candidate prediction server with the highest priority is determined as the target server for executing the prediction task.

[0098] It can be seen that through the above optional embodiments, by obtaining the historical processing records of candidate prediction servers and calculating the priority based on the device performance and the predicted time of the matching records, the server with the highest priority is selected as the target server. Therefore, on the basis of precise server allocation, the pertinence and efficiency of server selection are improved through performance and historical time analysis, providing optimized computing resources for cancer prediction tasks, and reducing the risk of prediction delays caused by insufficient server performance.

[0099] Example 3 See also Figure 3 , Figure 3 This is another server scheduling system based on feature analysis of sequencing data disclosed in an embodiment of the present invention. Figure 3 The server scheduling system based on feature analysis of sequencing data is applied to a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). Figure 3 As shown, the server scheduling system based on feature analysis of sequencing data may include: A memory 301 storing executable program code; a processor 302 coupled to the memory 301; The processor 302 calls the executable program code stored in the memory 301 to execute the steps of the server scheduling method based on feature analysis of sequencing data described in the first embodiment.

[0100] Example 4 An embodiment of the present invention discloses a computer storage medium storing a computer program for electronic data exchange, wherein the computer program enables a computer to execute the steps of the server scheduling method based on feature analysis of sequencing data described in Example 1.

[0101] Example 5 An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps of the server scheduling method based on feature analysis of sequencing data described in Example 1.

[0102] The foregoing description of specific embodiments of the present disclosure is intended to illustrate a method for performing a multi-tasking process. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0103] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0104] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0105] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0106] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0107] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0109] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0110] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0111] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0112] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0113] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.

[0114] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

[0115] Finally, it should be noted that the server scheduling method and system based on feature analysis of sequencing data disclosed in the embodiments of the present invention are only preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features therein may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A server scheduling method based on feature analysis of sequencing data, characterized in that: The method comprises: Obtain cfDNA sequencing data and corresponding patient parameters for cancer prediction; Analyzing data characteristic parameters corresponding to the cfDNA sequencing data based on a feature analysis algorithm; Predicting the prediction processing difficulty corresponding to the cfDNA sequencing data based on the patient parameters and the data characteristic parameters; Based on the difficulty of the prediction processing, a target server for performing the prediction task is determined from multiple candidate prediction servers; the target server is used to receive the cfDNA sequencing data and perform the cancer prediction task based on a built-in algorithm model.

2. The server scheduling method based on feature analysis of sequencing data according to claim 1, characterized in that: Analyzing the data characteristic parameters corresponding to the cfDNA sequencing data based on the characteristic analysis algorithm includes: Based on a preset segmentation rule, the cfDNA sequencing data is divided into multiple data groups; Calculating data-specific parameters corresponding to each of the data sets; Screening out data groups whose data-specific parameters are greater than a preset parameter threshold to obtain multiple specific data groups; The sequence position of each specific data group in the cfDNA sequencing data and the corresponding data-specific parameters are determined as data feature parameters corresponding to the cfDNA sequencing data.

3. The server scheduling method based on feature analysis of sequencing data according to claim 2, characterized in that: The segmentation rule is a rule based on average sequence length, a rule based on different sequencing batches, or a rule based on a similar clustering algorithm.

4. The server scheduling method based on feature analysis of sequencing data according to claim 2, characterized in that: The calculating of the data-specific parameters corresponding to each of the data groups includes: For each of the data groups, determining separation fragment parameters corresponding to the data group; Determine the total amount of data corresponding to the data group; Calculating the reciprocal of the average value of the data similarity between the data group and two other data groups adjacent to each other in sequence position, to obtain the incoherence parameter corresponding to the data group; A weighted sum of the separation segment parameter, the total data volume, and the incoherence parameter is calculated to obtain a data-specific parameter corresponding to the data group.

5. The server scheduling method based on feature analysis of sequencing data according to claim 4, characterized in that: Determining the separation fragment parameters corresponding to the data group includes: For each separation phenomenon type, determining the sequence fragments that have the separation phenomenon type among all sequence fragments in the data set to obtain separation fragments; the separation phenomenon type is AT separation phenomenon or GC separation phenomenon; Calculating the ratio of the total number of fragments of all the separated fragments to the total number of all sequence fragments in the data set; Calculate the weighted sum average of the data volume of the separation fragments corresponding to all the separation phenomenon types to obtain the separation fragment parameters corresponding to the data group; wherein the weighted calculation weight of the data volume of the separation fragment corresponding to each separation phenomenon type is proportional to the corresponding ratio.

6. The server scheduling method based on feature analysis of sequencing data according to claim 2, characterized in that: The predicting of the prediction processing difficulty corresponding to the cfDNA sequencing data according to the patient parameters and the data characteristic parameters includes: Inputting the patient parameters into a trained disease prediction model to obtain an output predicted disease type; the disease prediction model is trained by a training data set including a plurality of training patient parameters and corresponding disease annotations; Determining historical prediction records corresponding to the predicted disease type from a preset historical prediction database, and calculating an average of the prediction times of all the historical prediction records to obtain a historical time parameter; Inputting the cfDNA sequencing data and the data feature parameters into a trained processing cost prediction model to obtain an output predicted processing cost; the processing cost prediction model is trained by a training data set including a plurality of training sequencing data and corresponding data feature parameter annotations and processing cost annotations; The sum of the historical time parameter and the predicted processing cost is calculated to obtain the predicted processing difficulty corresponding to the cfDNA sequencing data.

7. The server scheduling method based on feature analysis of sequencing data according to claim 6, characterized in that: The patient parameters include the patient's physiological parameters, the patient's historical medical records and the patient's medical institution.

8. The server scheduling method based on feature analysis of sequencing data according to claim 1, characterized in that: Determining a target server for performing the prediction task from a plurality of candidate prediction servers according to the prediction processing difficulty includes: For each candidate prediction server, obtaining a historical prediction processing record of the candidate prediction server; Determining the device performance of the candidate prediction server based on the processing device performance in the historical prediction processing record; Filtering out records in the historical prediction processing records whose corresponding processing difficulty is greater than or equal to the predicted processing difficulty, to obtain a plurality of matching records; Calculate the average of the predicted processing times of all the matching records to obtain the historical predicted processing time; Calculating the product of the device performance and the historical prediction time to obtain the priority of the candidate prediction server; The candidate prediction server with the highest priority is determined as the target server for executing the prediction task.

9. A server scheduling system based on feature analysis of sequencing data, characterized in that: The system comprises: An acquisition module is used to obtain cfDNA sequencing data and corresponding patient parameters for cancer prediction; An analysis module, configured to analyze data characteristic parameters corresponding to the cfDNA sequencing data based on a characteristic analysis algorithm; A prediction module, configured to predict a prediction processing difficulty corresponding to the cfDNA sequencing data based on the patient parameters and the data characteristic parameters; A scheduling module is used to determine a target server for performing the prediction task from multiple candidate prediction servers based on the prediction processing difficulty; the target server is used to receive the cfDNA sequencing data and perform the cancer prediction task based on a built-in algorithm model.

10. A server scheduling system based on feature analysis of sequencing data, characterized in that: The system comprises: a memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the server scheduling method based on feature analysis of sequencing data as described in any one of claims 1 to 8.