Non-ABO blood group antigen distribution balance optimization method and system

By receiving and processing non-ABO blood type antigen detection data of red blood cell samples, calculating weights and using optimization algorithms to automatically select sample sets, the problem of uneven antigen distribution in the existing technology is solved, and the scientificity of sample selection and the accuracy of blood transfusion matching are improved.

CN120296303AInactive Publication Date: 2025-07-11杨和军
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510445995.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the selection of samples for non-ABO blood type antigen distribution depends on manual operation, and it is difficult to achieve automatic balance optimization of the Yin-Yang ratio of multiple antigens, resulting in unstable sample set quality, affecting the accuracy of antibody screening and blood transfusion matching.

Method used

By receiving multiple non-ABO blood type antigen yin and yang detection data of red blood cell samples, the weight coefficient is calculated and normalized. Combined with greedy algorithms, genetic algorithms or mixed algorithms, the target sample set is optimized to make the antigen yin and yang ratio approach the predetermined balance ratio, and the sample identification and statistical results are output.

Benefits of technology

The automatic optimization of the antigen-yin-positive ratio is achieved, which improves the scientificity and consistency of sample selection, reduces manual intervention, and improves the sensitivity of antibody screening and the accuracy of blood transfusion matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296303A_ABST
    Figure CN120296303A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical examination, and discloses a non-ABO blood group antigen distribution balance optimization method and system, and the method comprises the following steps: receiving positive and negative detection data of a plurality of non-ABO blood group antigens of a plurality of red blood cell samples; calculating a weight coefficient of each non-ABO blood group antigen according to a preset rule, and performing normalization processing on the weight coefficient to generate weight data; executing an optimization algorithm based on the weight data, a target sample number and a predetermined balance ratio, selecting a target number of samples from the plurality of red blood cell samples, and generating an optimal sample set; and outputting the sample identification information of the optimal sample set and an antigen distribution statistical result. According to the method, the optimal sample set is selected from a large number of red blood cell samples through an automatic process and an optimization algorithm, so that the negative-positive ratio of the antigen approaches a target value, manual intervention is reduced, the scientificity and consistency of sample selection are improved, and a reliable basis is provided for blood transfusion matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical testing technologies, and particularly to a method and system for optimizing the distribution balance of non-ABO blood group antigens. Background Art

[0002] In the fields of transfusion medicine and blood group research, the distribution balance of non-ABO blood group antigens (such as antigens in blood group systems like Kell, Duffy, Kidd, etc.) is of great significance for constructing a red blood cell sample bank. In the prior art, antigen positive and negative data of red blood cell samples are usually obtained through laboratory tests, and artificial screening or simple statistical methods are used to select a subset that meets the requirements from the samples for antibody screening reagent development or blood transfusion matching. For example, technicians may manually select samples based on experience, attempting to make the positive and negative ratios of multiple antigens approach a certain target value.

[0003] However, the prior art has significant deficiencies, mainly manifested in that the sample selection process depends on manual operation, and it is difficult to achieve automatic balance optimization of the distribution of multiple non-ABO blood group antigens. Manual screening is not only inefficient but also difficult to achieve equilibrium in the positive and negative distributions of multiple antigens. For example, the positive ratio of antigens may fluctuate between 0.3 and 0.7, deviating from the target value of 0.5, resulting in unstable quality of the sample set and affecting the sensitivity of antibody screening and the accuracy of blood transfusion matching. Summary of the Invention

[0004] To make up for the above deficiencies, the present invention provides a method and system for optimizing the distribution balance of non-ABO blood group antigens, aiming to improve the problem that the sample selection process in the prior art depends on manual operation and it is difficult to achieve automatic balance optimization of the distribution of multiple non-ABO blood group antigens.

[0005] In a first aspect, the present invention provides the following technical solution. A method for optimizing the distribution balance of non-ABO blood group antigens, applied to a server, includes the following steps: Receiving positive and negative detection data of multiple non-ABO blood group antigens of a plurality of red blood cell samples, where the detection data includes sample identifiers and positive and negative values of each antigen; Calculating weight coefficients of each non-ABO blood group antigen according to preset rules, and performing normalization processing on the weight coefficients to generate weight data; Based on the weight data, the target sample quantity, and a predetermined balance ratio, executing an optimization algorithm to select a target quantity of samples from the plurality of red blood cell samples to generate an optimal sample set, so that the positive and negative ratios of each non-ABO blood group antigen in the optimal sample set approach the predetermined balance ratio; Outputting sample identifier information of the optimal sample set and antigen distribution statistical results.

[0006] Preferably, the step of receiving the positive and negative test data of multiple non-ABO blood group antigens for multiple red blood cell samples includes: Obtain the positive and negative test data from a file path or a data table; Verify the existence of the sample identification column in the test data and the legality of the values in each antigen column, where the legality means that the antigen value is only 0 or 1; Extract all antigen columns except the sample identification to generate an antigen list.

[0007] Preferably, the step of calculating the weight coefficients of each non-ABO blood group antigen according to a preset rule and normalizing the weight coefficients to generate weight data includes: Receive the initial weight values of each non-ABO blood group antigen input by the user; Assign a default weight value of 1.0 to the antigens without set weights; Calculate the sum of all weight values and divide each weight value by the sum to generate normalized weight data.

[0008] Preferably, the step of performing an optimization algorithm based on the weight data, the target sample quantity, and a predetermined balance ratio includes: Automatically select an optimization algorithm according to the total sample quantity, the target sample quantity, and the number of antigen types. The optimization algorithms include a greedy algorithm, a genetic algorithm, or a hybrid algorithm; When the greedy algorithm is selected, iteratively select the sample that minimizes the deviation between the antigen distribution of the current sample set and the predetermined balance ratio based on the weight data; When the genetic algorithm is selected, search for the optimal sample set that meets the target sample quantity and the predetermined balance ratio through population initialization, crossover, mutation, and selection operations.

[0009] Preferably, the step of automatically selecting an optimization algorithm includes: When the total sample quantity is less than 100 or the ratio of the target sample quantity to the total sample quantity is greater than 0.5, select the greedy algorithm; When the total sample quantity is between 100 and 500 and the number of antigen types is greater than 10, select the hybrid algorithm; Otherwise, select the genetic algorithm.

[0010] Preferably, the step of outputting the sample identification information of the optimal sample set and the antigen distribution statistical result includes: Calculate the positive ratio, deviation, and weighted deviation of each non-ABO blood group antigen in the optimal sample set to generate antigen distribution statistical data; Save the sample identification information of the optimal sample set and the antigen distribution statistical data as a sample data table and a statistical data table respectively; Output the sample data table and the statistical data table.

[0011] Preferably, it further includes: Based on the optimal sample set, calculate the positive proportion and weighted deviation of each non-ABO blood group antigen; Generate a visualization chart, which includes a bar chart of the positive proportion of each antigen, a heat map of deviation, a pie chart of weight distribution, and a bar chart of weighted deviation; Output and save the visualization chart.

[0012] In a second aspect, the present invention provides the following technical solution. A non-ABO blood group antigen distribution balance optimization system includes: A data receiving module, configured to receive the positive and negative detection data of multiple non-ABO blood group antigens of multiple red blood cell samples, where the detection data includes sample identification and the positive and negative values of each antigen; A weight calculation module, configured to calculate the weight coefficient of each non-ABO blood group antigen according to a preset rule, and perform normalization processing on the weight coefficient to generate weight data; An optimization processing module, configured to execute an optimization algorithm based on the weight data, the target sample quantity, and a predetermined balance ratio, select a target quantity of samples from the multiple red blood cell samples, and generate an optimal sample set, so that the positive and negative proportions of each non-ABO blood group antigen in the optimal sample set approach the predetermined balance ratio; A result output module, configured to output the sample identification information of the optimal sample set and the antigen distribution statistical result.

[0013] In a third aspect, the present invention provides the following technical solution. A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned non-ABO blood group antigen distribution balance optimization method.

[0014] In a fourth aspect, the present invention provides the following technical solution. A readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned non-ABO blood group antigen distribution balance optimization method.

[0015] The present invention has the following beneficial effects: 1. In the present invention, through an automated process and an optimization algorithm, an optimal sample set is selected from a large number of red blood cell samples, so that the positive and negative proportions of antigens approach the target value, such as being optimized from 0.3 - 0.7 to 0.5, reducing manual intervention, improving the scientificity and consistency of sample selection, and providing a reliable basis for blood transfusion matching.

[0016] 2. In the present invention, the greedy, genetic or hybrid algorithm is automatically selected according to the total number of samples and the types of antigens. For example, the greedy algorithm is used for small-scale data to reduce the processing time, and the genetic algorithm is used for large-scale data to optimize the weighted deviation to 0.03, improving the applicability and computational efficiency of the system.

[0017] 3. In the present invention, the positive ratio, deviation and weighted deviation statistical data of each antigen are output, such as the positive ratio of 0.48 and the weighted deviation of 0.008, providing a basis for evaluating the distribution quality, supporting users to analyze the optimization effect, and optimizing antibody screening and blood transfusion matching.

[0018] 4. In the present invention, various data input methods are supported and the data legality is verified. For example, it is checked whether the antigen value is 0 / 1, heterogeneous data can be processed, the preprocessing burden is reduced, and the stability and generality of the system under different data sources are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flowchart of a method for optimizing the non-ABO blood group antigen distribution balance proposed by the present invention; Figure 2 is a system architecture diagram of a non-ABO blood group antigen distribution balance optimization system proposed by the present invention; Figure 3 is a data reception and verification flowchart of a method and system for optimizing the non-ABO blood group antigen distribution balance proposed by the present invention; Figure 4 is a weight calculation flowchart of a method and system for optimizing the non-ABO blood group antigen distribution balance proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] Embodiment 1 Referring to Figures 1-4 , in the first embodiment of the present invention, the present invention provides a method for optimizing the non-ABO blood group antigen distribution balance, which is applied to a server and includes the following steps: Receiving the positive and negative detection data of various non-ABO blood group antigens of multiple red blood cell samples, where the detection data includes sample identifiers and the positive and negative values of each antigen; Calculating the weight coefficients of each non-ABO blood group antigen according to preset rules, and normalizing the weight coefficients to generate weight data; Based on the weight data, the target sample quantity, and a predetermined balance ratio, execute an optimization algorithm to select a target number of samples from multiple red blood cell samples, generating an optimal sample set such that the positive-negative ratios of each non-ABO blood group antigen in the optimal sample set approach the predetermined balance ratio; Output the sample identification information of the optimal sample set and the statistical results of antigen distribution.

[0022] Specifically, the method is executed on the server side, and the balance of antigen distribution is achieved through systematic data processing and an optimization algorithm. Specifically, the step of receiving the positive-negative detection data of multiple non-ABO blood group antigens of red blood cell samples includes obtaining sample data from a laboratory information system (LIS) or other data sources. The data format can be a CSV file, a database table, or a directly input structured data table. Each row in the data represents a red blood cell sample, each column corresponds to a non-ABO blood group antigen (such as antigens in blood group systems like Kell, Duffy, Kidd, etc.), the positive-negative values are represented by 0 (negative) and 1 (positive), and the sample identification is used to uniquely identify each sample for subsequent tracking and output. The step of calculating the weight coefficients of each non-ABO blood group antigen according to preset rules takes into account the clinical importance of the antigen (such as the risk of antibody-induced hemolytic reactions), the distribution frequency (the high or low positive rate of the antigen in the population), and the research purpose (such as the blood transfusion requirements for a specific disease). An initial weight is generated through a weighting formula (such as based on expert scoring or a statistical model), and the weights are normalized to ensure that the sum of all weight values is 1, thus facilitating a fair comparison of the deviations of different antigens in subsequent optimization calculations. The step of executing the optimization algorithm based on the weight data, the target sample quantity, and the predetermined balance ratio. The target sample quantity is set by the user according to actual needs (such as 50 samples are needed to construct an antibody screening reagent). The predetermined balance ratio is usually set to 0.5 (i.e., the positive-negative ratio is 1:1), but it can also be adjusted according to specific applications (such as adjusted to 0.7 for some scenarios with a higher positive rate of certain antigens). The optimization algorithm iteratively calculates, evaluates the deviation between the positive-negative ratio of each antigen in the sample set and the target ratio, and calculates the weighted deviation in combination with the weight data, gradually screening out the sample combination with the smallest overall deviation to generate an optimal sample set. The step of outputting the sample identification information of the optimal sample set and the statistical results of antigen distribution outputs the sample identifications (such as sample numbers) of the optimal sample set in a list form, and at the same time generates statistical data, including the positive ratio, deviation, weighted deviation of each antigen, and the overall balance index (such as the average weighted deviation). These results can be used to evaluate the quality of the sample set and guide subsequent experiments or clinical applications. The implementation of this method can significantly improve the scientificity and efficiency of sample selection. For example, in antibody screening before blood transfusion, by balancing antigen distribution, false negative results can be reduced, the sensitivity of antibody detection can be improved, and at the same time, the risk of blood transfusion reactions caused by uneven antigen distribution can be reduced.

[0023] The steps of receiving the positive and negative detection data of multiple non-ABO blood group antigens for multiple red blood cell samples include: Obtain the positive and negative detection data from the file path or data table; Verify the existence of the sample identification column in the detection data and the legality of the values in each antigen column. The legality includes that the antigen value is only 0 or 1; Extract all antigen columns except the sample identification to generate an antigen list.

[0024] Specifically, the operation of obtaining the positive and negative detection data from the file path or data table supports multiple data input methods. For example, read a CSV file (such as a table file containing sample IDs and antigen detection results) by specifying the file path, or directly receive the structured data table (such as in the Pandas DataFrame format) passed in by the user through the program interface. This flexibility enables the method to adapt to the data management modes of different laboratories or research institutions. In the step of verifying the existence of the sample identification column in the detection data and the legality of the values in each antigen column, first check whether the "sample_id" column is included in the data as the sample identification. If it is missing, return an error message to avoid subsequent processing failures; then perform a legality check on the data in each antigen column to ensure that all values are 0 or 1 (representing negative and positive respectively). If non-0 / 1 values (such as missing values or abnormal values) are found, prompt the user to correct the data or automatically exclude invalid samples to ensure the standardization of the data. In addition, this step also includes a further check on the data integrity, such as verifying whether the number of samples is sufficient (at least greater than the target number of samples), whether the antigen columns are empty, etc., to avoid optimization failures caused by data problems. In the step of extracting all antigen columns except the sample identification to generate an antigen list, by traversing the column names of the data table, automatically identify all columns except the sample identification as antigen columns (such as "A1", "B1", "C1", etc.) and generate a list of antigen names for subsequent weight setting and optimization calculations. This process does not require the user to manually specify the antigen names, improving the automation and generality of the system. In practical applications, for example, a hospital laboratory provides detection data of 1000 red blood cell samples involving 10 non-ABO blood group antigens. This step can quickly verify the data format and extract the antigen list, providing reliable input for subsequent optimization. Through the above data reception and verification process, this method can effectively handle the data heterogeneity problem, ensure the robustness of the system under different data sources and formats, and at the same time provide clear error feedback to the user for data correction and system debugging.

[0025] The steps of calculating the weight coefficients of each non-ABO blood group antigen according to the preset rules and normalizing the weight coefficients to generate weight data include: Receive the initial weight values of each non-ABO blood group antigen input by the user; Assign a default weight value of 1.0 to antigens without set weights; Calculate the sum of all weight values and divide each weight value by the sum to generate normalized weight data.

[0026] Specifically, the step of receiving the initial weight values of each non-ABO blood group antigen input by the user supports the user to input weight values through a configuration file, a program interface or an interactive interface. For example, the user can set the weight of antigen Kell to 2.0 according to clinical experience (because the risk of hemolytic reaction caused by its antibody is relatively high), while the weight of antigen Duffy is 1.0 (because its clinical impact is relatively small). This flexible weight setting method allows the user to adjust the importance of antigens according to specific application scenarios (such as blood transfusion matching, antibody screening or blood group research). The step of assigning a default weight value of 1.0 to antigens without set weights automatically assigns a default value of 1.0 to antigens not provided with weights by the user to ensure that all antigens are considered in the optimization calculation and avoid calculation deviations caused by missing weights. For example, if the data contains 10 antigens and the user only sets weights for 6 of them, the remaining 4 antigens will be automatically assigned a weight of 1.0, and a corresponding prompt message will be generated to remind the user to check the integrity of the weight setting. The step of calculating the sum of all weight values and dividing each weight value by the sum to generate normalized weight data ensures that the sum of all weight values is 1 through normalization. This normalization not only facilitates the subsequent calculation of weighted deviations, but also avoids numerical instability caused by too large or too small weight values. In practical applications, for example, a research team optimized the antigens of the Kidd blood group system and set relatively high weights to prioritize the balance of its distribution. The normalized weight data generated by this step can ensure that the optimization algorithm highlights the balance requirements of key antigens while considering all antigens. The implementation of this step significantly improves the flexibility and adaptability of the system, making the optimization results more in line with actual needs, and at the same time enhancing the numerical stability of the algorithm through normalization.

[0027] The steps of executing the optimization algorithm based on the weight data, the target sample quantity and the predetermined balance ratio include: Automatically select an optimization algorithm according to the total number of samples, the target sample quantity and the number of antigen types. The optimization algorithms include a greedy algorithm, a genetic algorithm or a hybrid algorithm; When the greedy algorithm is selected, iteratively select the sample that minimizes the deviation between the antigen distribution of the current sample set and the predetermined balance ratio based on the weight data; When the genetic algorithm is selected, search for the optimal sample set that meets the target sample quantity and the predetermined balance ratio through population initialization, crossover, mutation and selection operations.

[0028] Specifically, the step of automatically selecting an optimization algorithm, based on the total number of samples, the number of target samples, and the number of antigen types, comprehensively considers the data scale and computational complexity to balance efficiency and optimization effect. The specific selection logic will be further described in claim 5. When the greedy algorithm is selected, based on the weight data, the step of iteratively selecting the sample that minimizes the deviation between the antigen distribution of the current sample set and the predetermined balance ratio is adopted, and the samples are added one by one using the greedy strategy: First, calculate the antigen distribution of the current sample set (initially empty) (if there are no samples, it is all 0), then traverse all unselected samples, calculate the antigen distribution of the new sample set after adding each sample, and calculate the weighted deviation from the target distribution according to the weight data, and select the sample with the smallest deviation to add until the target number of samples is reached. For example, if the target number is 50 and the predetermined balance ratio is 0.5, the system will preferentially select those samples that make the antigen positive ratio closer to 0.5, while considering the weighted impact of the weight on the deviation. When the genetic algorithm is selected, the step of searching for the optimal sample set that meets the target number of samples and the predetermined balance ratio through population initialization, crossover, mutation, and selection operations is adopted, using the global search ability of the genetic algorithm: First, initialize a population, each individual is a sample selection scheme (represented by a 0 / 1 vector, 1 means selected), and the population size can be set to 100; then select the parent individuals with higher fitness through the tournament selection method (the fitness is calculated based on the weighted deviation, and the smaller the deviation, the higher the fitness); then perform single-point crossover operation to generate offspring with a certain probability (such as 0.8), and perform mutation operation with a lower probability (such as 0.1) (randomly flip some bits); finally, adjust the number of selected samples in the offspring through the repair strategy to make it close to the target number, and select the individual with the highest fitness as the optimal sample set after 100 generations of iteration. In practical applications, for example, a certain laboratory needs to select 100 samples from 1000 samples for the development of an antibody screening reagent, and the samples involve 15 antigen types. The system may select the genetic algorithm to handle large-scale data and search for the global optimal solution through multi-generation evolution. The implementation of this step can flexibly select the algorithm according to the data scale, ensuring the efficient processing of small-scale data (greedy algorithm) and achieving the global optimization of large-scale data (genetic algorithm), thus significantly improving the accuracy and applicability of antigen distribution balance.

[0029] The step of automatically selecting an optimization algorithm includes: When the total number of samples is less than 100 or the ratio of the number of target samples to the total number of samples is greater than 0.5, select the greedy algorithm; When the total number of samples is between 100 and 500 and the number of antigen types is greater than 10, select the hybrid algorithm; Otherwise, select the genetic algorithm.

[0030] Specifically, when the total number of samples is less than 100 or the ratio of the number of target samples to the total number of samples is greater than 0.5, the steps of selecting the greedy algorithm take into account the computational efficiency requirements in scenarios of small-scale data or high selection ratios: for example, if the total number of samples is 80 and the number of target samples is 50 (ratio 0.625), the greedy algorithm can quickly iterate to select samples, each time selecting the sample with the smallest weighted deviation, with a low computational complexity and suitable for rapid response requirements. When the total number of samples is between 100 and 500 and the number of antigen types is greater than 10, the steps of selecting the hybrid algorithm are for medium-scale data and higher complexity scenarios. The hybrid algorithm combines the advantages of the greedy algorithm and the genetic algorithm: first, use the greedy algorithm to quickly select a part of the samples (such as 70%), and then use the genetic algorithm to optimize the remaining samples, thus reducing the calculation time while ensuring a certain optimization effect. For example, when the total number of samples is 300, the target number is 100, and the number of antigen types is 12, the hybrid algorithm can first select 70 samples using the greedy algorithm and then select 30 samples using the genetic algorithm. Otherwise, select the steps of the genetic algorithm, which is applicable to scenarios where the total number of samples is greater than 500 or the number of antigen types is small but global optimization is required. For example, when the total number of samples is 1000, the target number is 200, and the number of antigen types is 5, the genetic algorithm can effectively search for the global optimal solution through multiple generations of evolution, avoiding the local optimal problem that the greedy algorithm may fall into. In practical applications, for example, when a blood bank needs to select 300 samples from 2000 samples to construct a red blood cell bank with balanced antigens, involving 8 types of antigens, the system will select the genetic algorithm to ensure the global nature of the optimization result; while for the scenario of selecting 30 samples from 50 samples in a small laboratory, the greedy algorithm will be preferred to improve efficiency. The implementation of this step through dynamic algorithm selection takes into account the computational efficiency and optimization effect under different data scales, significantly improving the adaptability and practicality of the system, and at the same time providing users with a transparent basis for algorithm selection, facilitating the understanding and adjustment of the optimization strategy.

[0031] The steps of outputting the sample identification information of the optimal sample set and the antigen distribution statistical results include: Calculate the positive ratio, deviation, and weighted deviation of each non-ABO blood group antigen in the optimal sample set to generate antigen distribution statistical data; Save the sample identification information of the optimal sample set and the antigen distribution statistical data as a sample data table and a statistical data table respectively; Output the sample data table and the statistical data table.

[0032] Specifically, the steps of calculating the positive proportion, deviation, and weighted deviation of each non-ABO blood group antigen in the optimal sample set to generate antigen distribution statistical data are as follows: For each antigen in the optimal sample set, calculate the number of positive samples (the number of samples with a value of 1), the number of negative samples (the number of samples with a value of 0), and the positive proportion (the number of positive samples divided by the total number of samples). Then, calculate the deviation (the absolute difference between the positive proportion and the target proportion) based on a predetermined balance proportion (usually 0.5). Finally, calculate the weighted deviation (the deviation multiplied by the weight of the corresponding antigen) by combining the weight data. These statistical data can comprehensively reflect the antigen distribution quality of the sample set. The steps of saving the sample identification information and antigen distribution statistical data of the optimal sample set as a sample data table and a statistical data table respectively are as follows: Save the sample identification (such as the sample number) of the optimal sample set and its corresponding antigen detection data as a data table (for example, a sample data table in CSV format, including a "sample_id" column and columns for each antigen). At the same time, save the antigen distribution statistical data as another data table (including columns such as antigen name, number of positive samples, number of negative samples, positive proportion, deviation, weight, weighted deviation, etc.), and add a summary row to the statistical data table to record overall indicators such as the total number of samples, total weighted deviation, and average deviation. The steps of outputting the sample data table and the statistical data table support saving the data tables as files to a specified path (such as "output_samples.csv" and "output_stats.csv"), or returning them to the user through a program interface, which is convenient for the user to directly use or import into other systems for further analysis. In practical applications, for example, a research institution uses this method to select 100 samples from 500 samples, involving 10 antigens. The sample data table output by the system can be used to directly extract the selected red blood cell samples, while the statistical data table can be used to evaluate the balance of antigen distribution. For example, it is found that the positive proportion of a certain antigen is 0.48 (the target is 0.5), the deviation is 0.02, and the weighted deviation is 0.008 (the weight is 0.4), so as to judge whether the optimization effect meets the experimental requirements. The implementation of this step provides users with intuitive and comprehensive result output, which is not only convenient for sample management and experimental verification, but also supports the quantitative evaluation of the optimization effect through statistical data, significantly improving the practical value of the system.

[0033] It also includes: Based on the optimal sample set, calculate the positive proportion and weighted deviation of each non-ABO blood group antigen; Generate visualization charts, which include bar charts of the positive proportion of each antigen, heat maps of deviation, pie charts of weight distribution, and bar charts of weighted deviation; Output and save the visualization charts.

[0034] Specifically, the steps of calculating the positive proportion and weighted deviation of each non-ABO blood group antigen based on the optimal sample set are consistent with the statistical calculations in claim 6, but further prepare data for visualization: for each antigen, extract its positive proportion (the number of positive samples divided by the total number of samples), deviation (the difference from the target proportion), and weighted deviation (the deviation multiplied by the weight), and sort them by weight or weighted deviation to highlight key antigens. The steps of generating visualization charts, including a bar chart of the positive proportion of each antigen, a heat map of the deviation, a pie chart of the weight distribution, and a bar chart of the weighted deviation, are specifically implemented as follows: 1) For the bar chart of the positive proportion, use the antigen as the horizontal axis and the positive proportion as the vertical axis to draw the bar chart, and add a target balance line (such as 0.5) to intuitively show whether the distribution of each antigen is close to the target; 2) For the heat map of the deviation, use the antigen as the vertical axis and the deviation value as a single column, and use color gradient (such as from yellow to red) to represent the magnitude of the deviation to help users quickly identify antigens with larger distribution deviations; 3) For the pie chart of the weight distribution, draw the antigen weights as a pie chart proportionally, and antigens with weights less than a certain threshold (such as 0.05) are combined into an "other" category to highlight the weight distribution of the main antigens; 4) For the bar chart of the weighted deviation, use the antigen as the horizontal axis and the weighted deviation as the vertical axis, and sort them from largest to smallest weighted deviation to show the contribution of each antigen to the overall deviation. The steps of outputting and saving the visualization charts integrate the above four charts into a 2×2 subplot layout, set a unified chart style (such as a dark grid background), and add a title (such as "Optimization Results of the Distribution Balance of Non-ABO Blood Group Antigens"), and support saving the charts as high-resolution picture files (such as PNG format, DPI is 300), or directly displaying them on the user interface. In practical applications, for example, a blood bank optimized the antigen distribution of 200 samples, involving 12 antigens. The visualization chart shows that the positive proportion of a certain antigen is 0.52 (the target is 0.5), the deviation is 0.02, and the weighted deviation is the highest among all antigens. Users can adjust the weight or optimization parameters accordingly to further improve the results. The implementation of this step through multi-dimensional visualization not only intuitively presents the optimization effect, but also helps users quickly identify antigens with uneven distribution, provides data support for subsequent optimization, and improves the user experience and analysis efficiency of the system.

[0035] Example 2: Referring to Figure 2 , in the second embodiment of the present invention, the present invention provides a non-ABO blood group antigen distribution balance optimization system, including: A data receiving module, configured to receive the positive and negative detection data of multiple non-ABO blood group antigens of multiple red blood cell samples, and the detection data includes sample identification and the positive and negative values of each antigen; A weight calculation module, configured to calculate the weight coefficients of each non-ABO blood group antigen according to preset rules, and perform normalization processing on the weight coefficients to generate weight data; An optimization processing module, which is used to execute an optimization algorithm based on weight data, the target sample quantity, and a predetermined balance ratio, select a target quantity of samples from multiple red blood cell samples, generate an optimal sample set, and make the positive and negative ratios of non-ABO blood group antigens in the optimal sample set approach the predetermined balance ratio; A result output module, which is used to output the sample identification information of the optimal sample set and the statistical results of antigen distribution.

[0036] Specifically, the system runs on the server side, supports single-machine deployment or distributed deployment, can process large-scale red blood cell sample data. For example, it can process antigen detection data of tens of thousands of samples at the same time, and is applicable to hospital blood banks, blood group research institutions or transfusion medicine laboratories. The data receiving module supports multiple data input methods, including obtaining data from local files (such as CSV format), databases or laboratory information system (LIS) interfaces. The data format needs to include a sample identification column and an antigen column, and the antigen values are represented by 0 / 1 for positive and negative. This module also includes a data verification function, which can automatically detect data integrity and legality, and provide error prompts for users to correct. The weight calculation module allows users to input antigen weights through configuration files or interfaces, supports dynamic adjustment of weights to adapt to different application scenarios. For example, in the antibody screening scenario, the weights of high-risk antigens are increased. This module also includes a weight normalization function to ensure that the sum of weights is 1, which is convenient for subsequent weighted calculations. At the same time, it supports automatically assigning a default value of 1.0 to antigens without set weights to enhance the robustness of the system. The optimization processing module integrates multiple optimization algorithms (greedy algorithm, genetic algorithm, hybrid algorithm), can automatically select algorithms according to the data scale. For example, the greedy algorithm is used for small-scale data (the number of samples is less than 100), and the genetic algorithm is used for large-scale data (the number of samples is greater than 500). During the optimization process, the antigen distribution of the sample set is evaluated through weighted deviation. The goal is to make the positive and negative ratios of each antigen approach a predetermined value (such as 0.5), and supports users to customize optimization parameters (such as population size, number of iterations) to further improve the optimization effect. The result output module generates detailed optimization results, including the sample identification list of the optimal sample set and the statistical data of antigen distribution, supports saving the results as files (such as CSV format) or returning them through interfaces, and also supports visual output of the results (see claim 7 for details), which is convenient for users to intuitively analyze the optimization effect. In practical applications, for example, a certain blood bank uses this system to select 500 samples from 5000 samples, involving 20 antigens. The system can complete the optimization within a few minutes, output a balanced sample set, and generate a visual chart, significantly improving the efficiency and quality of sample selection. The implementation of this system realizes flexible combination of functions through modular design, not only supports data processing of different scales, but also improves the practicability and user experience of the system through automatic algorithm selection and result output.

[0037] The following is an introduction in combination with specific embodiments: 1: Development of Antibody Screening Reagents for Small-Scale Samples In a small laboratory, 30 samples need to be selected from 50 red blood cell samples for the development of antibody screening reagents, involving 5 non-ABO blood group antigens (A1, A2, B1, B2, C1). First, a CSV file containing sample identifiers and antigen negative / positive data (0 for negative, 1 for positive) is exported through the laboratory information system in the format of "sample_id, A1, A2, B1, B2, C1". The data receiving module validates the data legality and extracts the antigen list. The weight calculation module receives the weights input by the user {A1: 2.0, A2: 1.5, B1: 1.0, B2: 1.0, C1: 1.5}; Normalized to {A1: 0.2857, A2: 0.2143, B1: 0.1429, B2: 0.1429, C1: 0.2143}. The optimization processing module selects the greedy algorithm based on the total number of samples (50) and the selection ratio (30 / 50 = 0.6), sets the target balance ratio to 0.5, and iteratively selects the samples with the minimum weighted deviation to generate the optimal sample set. The result output module generates a sample data table (containing the identifiers and antigen data of 30 samples) and a statistical data table (showing that the positive ratio of A1 is 0.47, the deviation is 0.03, and the weighted deviation is 0.0086). The visualization module generates a bar chart showing that the positive ratios of each antigen are close to 0.5, and the deviation heat map highlights that the deviation of A1 is slightly larger. This embodiment is applicable to small-scale data, and the optimization process takes about 5 seconds, meeting the requirements of rapid screening.

[0038] 2: Optimization of Blood Transfusion Matching for Large-Scale Blood Bank Samples A large blood bank needs to select 500 samples from 5000 red blood cell samples for blood transfusion matching, involving 10 non-ABO blood group antigens (A1 to A5, B1 to B5). The data receiving module retrieves the data table from the database, containing sample identifiers and antigen negative / positive values, and extracts the antigen list after verification. The weight calculation module sets the weights of high-risk antigens (such as A1, B1) to 2.0 and the rest to 1.0, and normalizes them to {A1: 0.1429, B1: 0.1429, others: 0.0714}. Since the total number of samples is 5000, the optimization processing module selects the genetic algorithm, sets the population size to 100, iterates 100 generations, and sets the target balance ratio to 0.5 to generate the optimal sample set. The total weighted deviation is reduced from the initial 0.15 to 0.03. The result output module generates a sample data table and a statistical data table, showing that the positive ratio of A1 is 0.51, the deviation is 0.01, and the weighted deviation is 0.0014. The visualization module generates a heat map showing that the deviation of B2 is the largest at 0.03, and the pie chart of weight distribution shows that the weight proportions of A1 and B1 account for 28.6%. This embodiment is applicable to large-scale data, and the optimization takes about 3 minutes to ensure global distribution balance.

[0039] 3: Blood Type Research of Medium-Scale Samples A research institution needs to select 100 samples from 300 red blood cell samples for the study of the Kidd blood group system, which involves 12 antigens (K1 to K6, J1 to J6). The data receiving module receives data in the format of Pandas DataFrame, and extracts the antigen list after verification. The weight calculation module sets the weight of Kidd antigens (K1 to K6) to 2.0, and the rest to 1.0, and normalizes them to {K1 to K6: 0.1111, J1 to J6: 0.0556}. Since the total number of samples is 300 and the number of antigen types is 12, the optimization processing module selects a hybrid algorithm. First, it uses the greedy algorithm to select 70 samples, and then uses the genetic algorithm to select 30 samples. The target balance ratio is 0.5, and the total weighted deviation is reduced to 0.025. The result output module shows that the positive ratio of K1 is 0.49, the deviation is 0.01, and the weighted deviation is 0.0011. The visualization module generates a bar chart of weighted deviations, showing that the weighted deviation of J3 is the largest at 0.003. This embodiment is applicable to medium-scale data, and the optimization takes about 1 minute, highlighting the balance of Kidd antigens.

[0040] 4: Optimization of Antibody Detection with Customized Balance Ratio A hospital needs to select 200 samples from 1000 samples for antibody detection, which involves 8 antigens (D1 to D4, E1 to E4), and the target balance ratio is set to 0.7 (because the positive rates of some antigens are relatively high). The data receiving module obtains data from a CSV file and extracts the antigen list after verification. The weight calculation module sets equal weights and normalizes them to 0.125 each. The optimization processing module selects the genetic algorithm, iterates 50 generations, generates the optimal sample set, and the total weighted deviation is 0.04. The result output module shows that the positive ratio of D1 is 0.68, the deviation is 0.02, and the weighted deviation is 0.0025. The visualization module generates a bar chart, showing that the positive ratio is close to 0.7, and the deviation heat map highlights the deviation of E2 at 0.03. This embodiment supports customized balance ratio, and the optimization takes about 2 minutes to meet specific detection requirements.

[0041] Embodiment III In the third embodiment of the present invention, based on the same inventive concept, the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the steps of the non-ABO blood type antigen distribution balance optimization method in the above embodiment.

[0042] Embodiment IV In the fourth embodiment of the present invention, based on the same inventive concept, the present invention provides a computer device, including: a processor and a memory; the processor and the memory communicate with each other; the memory is used to store instructions; the processor is used to execute the instructions in the memory to implement the non-ABO blood type antigen distribution balance optimization method in the above embodiment.

[0043] It should be understood that each part of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0044] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for optimizing the distribution balance of non-ABO blood group antigens, applied to a server, characterized in that, It includes the following steps: Receiving the positive and negative test data of multiple non-ABO blood group antigens of multiple red blood cell samples, where the test data includes sample identifiers and positive and negative values of each antigen; Calculating the weight coefficients of each non-ABO blood group antigen according to preset rules, and normalizing the weight coefficients to generate weight data; Based on the weight data, the target sample quantity, and a predetermined balance ratio, performing an optimization algorithm to select a target number of samples from the multiple red blood cell samples to generate an optimal sample set, such that the positive and negative ratios of each non-ABO blood group antigen in the optimal sample set approach the predetermined balance ratio; Outputting the sample identifier information of the optimal sample set and the antigen distribution statistical results.

2. The non-ABO blood group antigen distribution balance optimization method according to claim 1, wherein The step of receiving the positive and negative test data of multiple non-ABO blood group antigens of multiple red blood cell samples includes: Obtaining the positive and negative test data from a file path or a data table; Verifying the existence of the sample identifier column in the test data and the legality of the values in each antigen column, where the legality includes that the antigen value is only 0 or 1; Extracting all antigen columns except the sample identifier to generate an antigen list.

3. The non-ABO blood group antigen distribution balance optimization method according to claim 1, characterized in that The step of calculating the weight coefficients of each non-ABO blood group antigen according to preset rules and normalizing the weight coefficients to generate weight data includes: Receiving the initial weight values of each non-ABO blood group antigen input by a user; Assigning a default weight value of 1.0 to antigens without set weights; Calculating the sum of all weight values, and dividing each weight value by the sum to generate normalized weight data.

4. The non-ABO blood group antigen distribution balance optimization method according to claim 3, characterized in that The step of performing an optimization algorithm based on the weight data, the target sample quantity, and a predetermined balance ratio includes: Automatically selecting an optimization algorithm according to the total number of samples, the target sample quantity, and the number of antigen types, where the optimization algorithms include a greedy algorithm, a genetic algorithm, or a hybrid algorithm; When the greedy algorithm is selected, iteratively selecting the sample that minimizes the deviation between the antigen distribution of the current sample set and the predetermined balance ratio based on the weight data; When the genetic algorithm is selected, searching for an optimal sample set that meets the target sample quantity and the predetermined balance ratio through population initialization, crossover, mutation, and selection operations.

5. The non-ABO blood group antigen distribution balance optimization method according to claim 1, wherein The step of automatically selecting an optimization algorithm includes: When the total number of samples is less than 100 or the ratio of the target sample quantity to the total number of samples is greater than 0.5, selecting the greedy algorithm; When the total number of samples is between 100 and 500 and the number of antigen types is greater than 10, selecting the hybrid algorithm; Otherwise, selecting the genetic algorithm.

6. The non-ABO blood group antigen distribution balance optimization method according to claim 1, wherein, The step of outputting the sample identifier information of the optimal sample set and the antigen distribution statistical results includes: Calculating the positive ratio, deviation, and weighted deviation of each non-ABO blood group antigen in the optimal sample set to generate antigen distribution statistical data; Saving the sample identifier information of the optimal sample set and the antigen distribution statistical data as a sample data table and a statistical data table respectively; Outputting the sample data table and the statistical data table.

7. The non-ABO blood group antigen distribution balance optimization method according to claim 1, characterized in that It also includes: Based on the optimal sample set, calculating the positive ratio and weighted deviation of each non-ABO blood group antigen; Generating visualization charts, where the visualization charts include a bar chart of the positive ratio of each antigen, a heat map of the deviation, a pie chart of the weight distribution, and a bar chart of the weighted deviation; Output and save the visualization chart.

8. A non-ABO blood group antigen distribution balance optimization system, characterized in that, The non-ABO blood group antigen distribution balance optimization method according to any one of claims 1-7, comprising: A data receiving module, configured to receive the positive and negative detection data of multiple non-ABO blood group antigens of multiple red blood cell samples, where the detection data includes sample identifiers and the positive and negative values of each antigen; A weight calculation module, configured to calculate the weight coefficients of each non-ABO blood group antigen according to preset rules, and perform normalization processing on the weight coefficients to generate weight data; An optimization processing module, configured to execute an optimization algorithm based on the weight data, the target sample quantity, and a predetermined balance ratio, select a target quantity of samples from the multiple red blood cell samples, and generate an optimal sample set, so that the positive and negative ratios of each non-ABO blood group antigen in the optimal sample set approach the predetermined balance ratio; A result output module, configured to output the sample identifier information of the optimal sample set and the antigen distribution statistical result.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the non-ABO blood group antigen distribution balance optimization method according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that, A computer program is stored on the readable storage medium, and when the computer program is executed by the processor, it implements the non-ABO blood group antigen distribution balance optimization method according to any one of claims 1 to 7.