Mediterranean anemia screening auxiliary system and method

Through multi-dimensional feature screening and deep learning models, combined with genomics database, a simplified detection parameter set is designed to solve the problem of insufficient equipment in thalassemia screening, and efficient and accurate large-scale screening is achieved.

CN120260959APending Publication Date: 2025-07-04TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510652975.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The insufficient number of testing equipment in thalassemia screening has led to slow progress in large-scale population screening, large load on testing equipment, making it difficult to cover a wide range of people.

Method used

Identify high-risk core populations through multi-dimensional feature screening conditions, use deep learning models to train samples of the same region or type of thalassin, design a simplified detection parameter set, and build a gene relationship map with genomics databases, screen and expand personnel and perform simplified detection.

Benefits of technology

It improves the efficiency and accuracy of screening, reduces non-essential testing steps, shortens detection time, optimizes the detection process, and ensures data security and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260959A_ABST
    Figure CN120260959A_ABST
Patent Text Reader

Abstract

The invention discloses an auxiliary system and method for screening thalassemia in the technical field of anemia screening, and the system comprises the following modules: a data collection and preprocessing module, a core population screening module which is used for setting screening conditions based on family relationship, geographical distribution, working properties, ethnic characteristics and age multi-dimensional characteristics, core crowds are analyzed and screened out; the gene reasoning and expansion personnel identification module is used for reasoning potential thalassemia carriers or high-risk expansion personnel according to the gene information of the core crowd, simplifying the parameter detection module, designing a simplified detection parameter set and simplifying unnecessary detection steps; a result analysis and feedback module; and a system optimization and maintenance module. The method is scientific and reasonable, high-risk core crowds and expansion personnel can be identified through multi-dimensional feature screening conditions, unnecessary detection steps of the expansion personnel are reduced, the detection time is shortened, the detection accuracy is improved, and therefore the use efficiency of instruments is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of anemia screening, and specifically relates to a thalassemia screening assistance system and method. Background Art

[0002] Thalassemia is a hereditary hemolytic anemia caused by the mutation or deletion of globin genes, resulting in the synthesis disorder of one or more globin peptide chains, and is one of the most common single-gene genetic diseases clinically. The thalassemia screening assistance system and method are of great significance for the early detection, prevention and control of the disease.

[0003] There are many thalassemia screening and detection methods. For example, blood routine test: Since thalassemia patients usually have a decrease in mean corpuscular volume and mean corpuscular hemoglobin content, and may be accompanied by an increase in hemoglobin A2 at the same time, by observing the quantity changes and morphological distributions of red blood cells and hemoglobin, it is initially judged whether thalassemia exists. Hemoglobin electrophoresis test: According to the changes in hemoglobin, it is judged whether thalassemia is present, and even the type of thalassemia can be judged. Thalassemia gene detection: By collecting blood or other tissue cells and using specific equipment to detect the genetic material in the cells, it is the gold standard for diagnosing thalassemia. Erythrocyte osmotic fragility test: It is used to detect the osmotic fragility of red blood cells and initially judge the possibility of thalassemia. Peptide chain examination: It is used to determine whether mild α-thalassemia is present by polypeptide chain detection. However, thalassemia is mainly concentrated in mountainous areas, and there are usually only a small number of professional detection devices in the corresponding areas. Therefore, when screening the population for thalassemia, it is affected by the number of detection devices, with slow progress, heavy load on detection devices, and it is difficult to conduct large-scale screening for a large number of people.

[0004] In summary, how to increase the utilization rate of instrument detection and improve the screening coverage has become an urgent problem for those skilled in the art. Therefore, it is necessary to propose a thalassemia screening assistance system and method. Summary of the Invention

[0005] In order to solve the above problems, the purpose of the present invention is to provide a thalassemia screening assistance system and method. Through the screening conditions of multi-dimensional features, the system can accurately identify the core population at high risk, train the thalassemia samples in the same region or of the same type through a deep learning model, extract the core features, and design a simplified set of detection parameters. It reduces unnecessary detection steps, shortens the detection time, and at the same time maintains a high detection accuracy, thereby improving the use efficiency of the instrument.

[0006] In order to achieve the above purpose, the technical solution of the present invention is as follows: A thalassemia screening assistance system includes the following modules: A data collection and preprocessing module, which is used to collect regional population data. The population data includes geographical grid division, family network relationships, and basic population information. The basic population information includes age, gender, ethnicity, occupation, and place of residence, and preprocesses the data.

[0007] A core population screening module, which is used to set screening conditions based on multi-dimensional characteristics such as family relationships, geographical distribution, nature of work, ethnic characteristics, and age. Generate a genetic family pedigree representing genetic relationships based on family relationships, analyze and screen out the core population. The core population is the population in which the number of gene-associated populations in the current region within the genetic family pedigree is greater than a preset value.

[0008] A gene inference and expanded personnel identification module: After performing thalassemia detection on the core population, according to the gene information of the core population, use the single-gene inheritance law to predict potential risk populations, deduce potential thalassemia carriers or high-risk expanded personnel, then construct a gene relationship map, comprehensively consider factors such as family genetic history and geographical distribution, evaluate the individual risk level, set a risk threshold, and screen out the list of expanded personnel. The list of expanded personnel is the personnel who have a deduced relationship with the patients within the core population according to the gene inheritance law.

[0009] A simplified parameter detection module, which is used to perform thalassemia detection on the expanded personnel based on simplified parameters. Use a convolutional neural network to construct a deep learning model. The learning model is trained based on data of thalassemia patients in the same region or of the same type. Based on the training results, generate several simplified detection parameter sets classified by region and thalassemia type, and judge the simplified detection parameter set adopted by the list of expanded personnel corresponding to the patients within the core population through the detection results of the patients within the core population.

[0010] A result analysis and feedback module, which is used to statistically analyze the thalassemia detection results based on simplified parameters, generate a screening feedback report, and evaluate the sensitivity and specificity of the screening according to the preset detection strategy. Randomly select 10 people from the thalassemia detection personnel based on simplified parameters for re-examination to verify the accuracy, adjust the thalassemia detection strategy, and finally summarize the screening data.

[0011] A system optimization and maintenance module: used to regularly evaluate and optimize the system performance, and update the algorithm model and database.

[0012] Furthermore, the data collection and preprocessing module also includes a preprocessing unit. The preprocessing unit is used to clean, de-duplicate, and format the basic population information, verify the integrity and consistency of the data, including checking the ID number, age range, identifying and marking missing values or outliers, and filling or removing them. The missing values include missing ID cards, missing occupations, missing regions, and missing ages. The outliers include ages greater than 75 years old and less than 2 years old.

[0013] Further, the data collection and preprocessing module further includes an encryption unit, which is used to encrypt the population data using a symmetric encryption algorithm, and the symmetric encryption algorithm includes AES and DES.

[0014] Further, the data collection and preprocessing module further includes a user interaction unit, which is used to provide an operation interface for system administrators, medical professionals, and the public to input data, view screening results, and receive system feedback, and for medical professionals to adjust thalassemia detection strategies and query historical data through the user interaction unit.

[0015] Further, the core population screening module further includes a machine learning algorithm unit, which is used to identify high-risk families and regions using clustering analysis, decision trees, and vector machines.

[0016] Further, the screening conditions in the core population screening module include whether there are known thalassemia patients in the family, whether there is a specific ethnic background, whether living in a high-incidence area, marriage and reproductive history, personal health status, diet and nutrition status, environmental exposure factors, drug use history, and medical service utilization.

[0017] Further, the gene inference and extended personnel identification module further includes a genomics database for inputting and storing historical data of thalassemia patients, and the gene inference and extended personnel identification module uses the genomics database for comparison to construct a gene relationship map.

[0018] Further, the simplified parameter detection module further includes an automated parameter adjustment unit for automatically adjusting the detection parameter set according to the data feedback during the thalassemia detection process of the extended personnel.

[0019] Further, the simplified parameter detection module further includes a transfer learning unit for applying the trained deep learning model to new regions or new types of thalassemia samples.

[0020] Further, a method for assisting thalassemia screening, based on the above-mentioned thalassemia screening assistance system, includes the following steps: Step 1, data collection and preprocessing: Collect and preprocess the population data in the region and encrypt it; Step 2, core population screening: Use machine learning algorithms to classify the population data and identify the core population in high-risk families or regions; Step 3, gene inference and extended personnel identification: Use the core personnel to construct a gene relationship map, use the single-gene inheritance law to predict potential risk populations, and screen out the extended personnel who need further testing; Step 4, simplified parameter detection: Adopt a detection step based on simplified parameters for the extended personnel; Step 5, Result Analysis and Feedback: Statistically analyze the detection results and generate a customized screening report.

[0021] After adopting the above solution, the following principles and beneficial effects are achieved: 1. By comprehensively considering multi-dimensional features such as family relationship, geographical distribution, nature of work, ethnic characteristics, and age, the system of the present invention can accurately locate and screen out the core population. Focusing on screening the core population helps to expand a larger group of additional personnel who can simplify the detection parameter set, saving subsequent detection resources and costs. This method is more comprehensive and effective than single-dimensional screening, and can significantly improve the efficiency and pertinence of screening.

[0022] 2. Using the genetic information of the core population, combining with the genomics database and the single-gene inheritance law, the present invention constructs a gene relationship map and predicts potential risk populations. By setting a risk threshold, the system can further screen out the list of additional personnel who need further testing, realizing refined screening from the population to the individual.

[0023] 3. The present invention constructs a deep learning model through a convolutional neural network, trains thalassemia samples in the same region or of the same type, extracts core features, and designs a simplified detection parameter set. This method reduces unnecessary detection steps, shortens the detection time, while maintaining a high detection accuracy, improves the utilization efficiency of the instrument, and quickly covers a large number of people.

[0024] 4. The automatic parameter adjustment unit and transfer learning unit in the present invention can automatically adjust the detection parameter set according to the data feedback during the detection process, and apply the trained deep learning model to thalassemia samples in new regions or of new types. This adaptive ability enables the system to continuously optimize the detection process and improve the screening effect.

[0025] 5. The encryption unit in the data collection and preprocessing module of the present invention encrypts the population data using a symmetric encryption algorithm. The data encryption unit ensures the security and privacy protection of the population data, preventing data leakage and abuse. The user interaction unit provides a user-friendly interface, facilitating operations such as data input, result viewing, and system feedback for system administrators, medical professionals, and the public, improving the user experience. Description of the Drawings

[0026] Figure 1 It is the system flow chart of the thalassemia screening assistance system in the embodiment of the present invention.

[0027] Figure 2 It is the flow chart of the data collection and preprocessing module of the thalassemia screening assistance system in the embodiment of the present invention.

[0028] Figure 3This is a flowchart of the auxiliary method for thalassemia screening in the embodiments of the present invention. Detailed implementation manners

[0029] The following is a further detailed description through specific implementation manners: Embodiment 1

[0030] Basically as shown in the appendix Figure 1 and Figure 2 shown: An auxiliary system for thalassemia screening includes the following modules: A data collection and preprocessing module, which is used to collect regional population data. The population data includes geographical grid division, family network relationship, and basic population information. The basic population information includes age, gender, ethnicity, occupation, and place of residence. The data collection and preprocessing module also includes a preprocessing unit, which is used to clean, de-duplicate, and format the entrance data, and verify the integrity and consistency of the data, including checking the ID number, age range, identifying and marking missing values or outliers, and filling or eliminating them. The missing values include missing ID card, missing occupation, missing area, and missing age. The outliers include age greater than 75 years old and less than 2 years old. The data collection and preprocessing module also includes an encryption unit, which is used to encrypt the population data using symmetric encryption algorithms. The symmetric encryption algorithms include AES and DES. The data collection and preprocessing module also includes a user interaction unit, which is used to provide a user-friendly interface, enabling system administrators, medical professionals, and the public to input data, view screening results, receive system feedback, and medical professionals to adjust thalassemia detection strategies and query historical data through the user interaction unit.

[0031] A core population screening module, which is used to set screening conditions based on multi-dimensional characteristics such as family relationship, geographical distribution, nature of work, ethnic characteristics, and age. The screening conditions in the core population screening module include whether there are known thalassemia patients in the family, whether there is a specific ethnic background, whether living in a high-incidence area, marriage and fertility history, personal health status, diet and nutrition status, environmental exposure factors, drug use history, and medical service utilization.

[0032] Marriage and fertility history: Ask about the applicant's marital status and the thalassemia history of the spouse's family, because the thalassemia gene carrier status of the spouse is crucial for the offspring. At the same time, understand the applicant's past fertility history, especially whether there is a history of giving birth to thalassemia children or having a miscarriage.

[0033] Personal health status: In addition to direct genetic and family factors, certain personal health conditions may also be related to the medical history of thalassemia risk, such as anemia symptoms like pale complexion, fatigue, palpitations, long-term chronic diseases such as liver diseases, kidney diseases, etc. These diseases may affect iron metabolism.

[0034] Diet and nutritional status: Thalassemia patients often need to pay special attention to iron and nutrient intake. Therefore, the applicant's daily eating habits can be inquired about, including whether there is partiality for certain foods, whether there is sufficient intake of iron and vitamins, and whether there is malnutrition.

[0035] Environmental exposure factors: Although environmental factors are not the direct cause of thalassemia, certain environmental factors may exacerbate the symptoms of thalassemia or affect the individual's health status. For example, inquire whether the applicant has long-term exposure to harmful substances such as heavy metals, certain chemical reagents, etc., and the safety of the working environment.

[0036] Drug use history: Inquire whether the applicant is currently using or has ever used drugs that may affect iron metabolism or blood system function, such as certain antibiotics, anti-tumor drugs, etc.

[0037] Medical service utilization: Understand whether the applicant regularly receives medical and health services, and whether they have undergone thalassemia-related screening or testing in the past, which helps to evaluate their awareness of the disease and the implementation of preventive measures.

[0038] The core population screening module also includes a machine learning algorithm unit. The machine learning algorithm unit is used to identify high-risk families and regions using clustering analysis, decision trees, and vector machines, generate a genetic family pedigree representing genetic relationships based on family relationships, analyze and screen out the core population. The core population is the population in which the number of gene-associated people in the current region within the genetic family pedigree is greater than a preset value.

[0039] Gene inference and extended personnel identification module: After performing thalassemia detection on the core population, according to the gene information of the core population, it uses the single-gene inheritance law to predict potential risk populations, deduce potential thalassemia carriers or high-risk extended personnel. The gene inference and extended personnel identification module also includes a genomics database for entering and storing historical data of thalassemia patients. The gene inference and extended personnel identification module uses the genomics database for comparison to construct a gene relationship map. Considering family genetic history and geographical distribution factors comprehensively, it evaluates the individual risk level, sets a risk threshold, and screens out the list of extended personnel. The list of extended personnel is the personnel who have a deduced genetic relationship with the patients within the core population.

[0040] The simplified parameter detection module is used to perform thalassemia detection on the extended personnel based on simplified parameters. It constructs a deep learning model using a convolutional neural network. The learning model is trained based on data of thalassemia patients in the same region or of the same type. Based on the training results, it generates several simplified detection parameter sets classified by region and thalassemia type, and determines the simplified detection parameter set adopted by the list of extended personnel corresponding to the patients within the core population through the detection results of the patients within the core population, reducing unnecessary detection steps.

[0041] The detection parameter set examples include, but are not limited to, key genetic marker detection: Based on the training results of thalassemia samples in the same region or of the same type using a deep learning model, some specific gene mutation sites highly correlated with thalassemia can be determined. These sites can serve as the core part of the simplified detection parameter set, and potential thalassemia patients or carriers can be quickly screened out through gene detection means. In addition to directly pathogenic gene mutations, some genetic polymorphism markers related to thalassemia risk can also be considered. These markers may indirectly affect the occurrence of thalassemia by influencing gene expression, regulatory mechanisms, etc.

[0042] Hematological index screening: Such as mean corpuscular volume (MCV), mean corpuscular hemoglobin (MCH), etc. These parameters usually show abnormal changes in thalassemia patients. By setting reasonable thresholds, objects that need further detection can be quickly screened out. Hemoglobin electrophoresis is a commonly used thalassemia screening method, which can detect the abnormal types and proportions of hemoglobin. In the simplified parameter set, only the detection of hemoglobin types highly related to thalassemia can be included.

[0043] The simplified parameter detection module also includes an automated parameter adjustment unit, which can automatically adjust the detection parameter set according to the data feedback during the detection process by additional personnel. The simplified parameter detection module also includes a transfer learning unit, which is used to apply the trained deep learning model to thalassemia samples in new regions or of new types.

[0044] The result analysis and feedback module is used to statistically analyze the thalassemia detection results based on the simplified parameters, generate a screening feedback report, and evaluate the sensitivity and specificity of the screening according to the preset detection strategy. Randomly select 10 people from the thalassemia detection personnel based on the simplified parameters for reexamination to verify the accuracy, adjust the thalassemia detection strategy, assist in further clinical diagnosis and gene sequencing, and finally summarize the screening data.

[0045] The system optimization and maintenance module: It is used to regularly evaluate and optimize the system performance, and update the algorithm model and database.

[0046] Specific implementation steps:

[0047] First, collect the population data in this region through government databases, medical institution records, and public voluntary submissions. These data include geographical grid division into streets and communities, family network relationships obtained through questionnaires or existing medical records, and basic population information including age, gender, ethnicity, occupation, place of residence, etc.

[0048] The preprocessing unit automatically cleans the collected data, removes duplicate records, formats the data format, and verifies the integrity and consistency of the data. For example, it checks whether the ID number conforms to the national standard, whether the age is within a reasonable range, identifies and marks missing values or outliers, such as records with extremely large or small ages, and performs appropriate filling or elimination processes. At the same time, the encryption unit uses the AES encryption algorithm to encrypt sensitive data to ensure data security. The system administrator inputs the screening strategy through a user-friendly interface, medical professionals can view the screening progress and results, and the public can input their basic information through specific channels and receive screening feedback.

[0049] Based on multi-dimensional characteristics such as family relationship, geographical distribution, nature of work, ethnic characteristics, and age, the system sets screening conditions. For example, people who have known thalassemia patients in their family, belong to a specific high-incidence ethnic group, and live in a historically high-incidence area are listed as key screening targets. The machine learning algorithm unit uses clustering analysis, decision trees, and support vector machines to deeply analyze the preprocessed data and identify high-risk families and regions. Through the algorithm model, the system can automatically screen out the core population that meets the screening conditions.

[0050] For the core population, the system uses the genomics database to compare gene information and construct a gene relationship map. Through the comparison and analysis, the system can predict potential thalassemia carriers or high-risk expansion personnel. Considering family genetic history and geographical distribution factors comprehensively, the system evaluates the risk level of each individual and sets a risk threshold. According to the evaluation results, a list of expansion personnel who need further testing is screened out.

[0051] A deep learning model is constructed using a convolutional neural network to train thalassemia samples in the same region or of the same type and extract core features. Based on these features, the system designs a simplified set of detection parameters to reduce unnecessary detection steps.

[0052] For example: Suppose in a certain region, the main types of thalassemia are closely related to specific gene mutations, such as certain mutation sites of β-thalassemia. Through the training of the deep learning model, it is found that these gene mutations are one of the core features of thalassemia patients in this region. Therefore, in the simplified set of detection parameters, only the detection items for these specific gene mutations can be included. For the expansion personnel, only these specific gene mutations need to be detected to initially determine whether they are thalassemia patients or carriers, thus greatly reducing the detection cost and time.

[0053] During the detection process, the automatic parameter adjustment unit automatically optimizes the set of detection parameters according to the data feedback. At the same time, the transfer learning unit applies the trained model to thalassemia samples in new regions or of new types to improve the adaptability and accuracy of the system.

[0054] The result analysis and feedback module performs statistical analysis on the test results, and randomly selects 10 people from the thalassemia test personnel based on simplified parameters for re-examination to verify the accuracy. Finally, the sensitivity and specificity of the screening are calculated to evaluate the screening effect. A detailed screening report is generated, including screening results, risk assessment, recommended measures, etc., and feedback is given to relevant personnel. For suspected patients, the system assists in further clinical diagnosis and gene sequencing. Based on the feedback results and actual needs, the system administrator can adjust the screening strategy and test parameters to optimize the screening effect.

[0055] The system optimization and maintenance module regularly evaluates the system's performance, including processing speed, accuracy and other indicators. Based on the evaluation results, the system is optimized and upgraded. Algorithm and database update: With the progress of scientific research and the accumulation of data, the system promptly updates the algorithm model and database information to maintain the advancement and accuracy of the system.

[0056] Example 2

[0057] Basically as attached Figure 3 As shown in the figure, different from the above-mentioned embodiment, a thalassemia screening auxiliary method comprises the following steps: Step 1: Data collection and preprocessing: Collect and preprocess the population data in the area and encrypt it.

[0058] Step 2: Core population screening: Use machine learning algorithms to classify population data and identify core populations in high-risk families or regions.

[0059] Step three, genetic reasoning and identification of expansion personnel: Use core personnel to build a genetic relationship map, use single-gene inheritance rules to predict potential risk groups, and screen out expansion personnel who need further testing.

[0060] Step 4: Simplified parameter detection: Adopt detection steps based on simplified parameters for the expanded personnel.

[0061] Step 5: Result analysis and feedback: Conduct statistical analysis on the test results and generate a customized screening report.

[0062] Specific implementation process: Suppose that thalassemia screening is conducted in a specific area of ​​a province. First, population data in the area is collected through multiple channels such as government departments, medical institutions, and community service centers. This data includes but is not limited to residents' ID number, age, gender, ethnicity, occupation, residential address, and family relationships.

[0063] In the data preprocessing stage, specialized data processing software is used to clean the collected data. This includes removing duplicate records, formatting data fields, checking the validity of ID numbers and the rationality of age ranges, etc. For the missing values or outliers found, necessary filling or elimination is performed according to the data context. Meanwhile, to ensure data security, the AES encryption algorithm is used to encrypt sensitive data.

[0064] The preprocessed data is input into the machine learning algorithm model. Algorithms such as decision trees and clustering analysis are utilized, and these algorithms can classify population data based on multi-dimensional features such as family relationships, geographical distributions, nature of work, ethnic characteristics, and age.

[0065] For example, the model may identify that there are multiple known thalassemia patients in a certain family, or the incidence of thalassemia in a certain region is significantly higher than that in other regions. Based on these features, the algorithm can automatically screen out the core population in high-risk families or regions.

[0066] For the screened core population, a gene relationship map is further constructed using the genomics database. By comparing the gene information of the core population, specific gene mutations related to thalassemia can be discovered.

[0067] The law of single-gene inheritance is used to predict potential high-risk populations. For example, if a core individual carries the pathogenic gene for thalassemia, then his direct relatives and close relatives have a relatively high risk of becoming carriers or patients. By comprehensively considering factors such as family genetic history and geographical distribution, the risk level of each individual is evaluated, and a list of additional individuals who need further testing is screened out.

[0068] For the screened additional individuals, a screening procedure based on simplified parameters is adopted for thalassemia screening. This usually includes some rapid, simple, and low-cost detection methods, such as hematological examinations, hemoglobin electrophoresis, etc.

[0069] To further improve the detection efficiency and accuracy, a deep learning model is constructed using deep learning techniques such as convolutional neural networks. The deep learning model is trained on thalassemia samples from the same region or of the same type, extracts core features, and designs a set of simplified detection parameter sets based on this. These parameter sets can reduce unnecessary detection steps while maintaining high detection sensitivity and specificity.

[0070] Finally, statistical analysis is performed on the test results of all augmented personnel. This includes calculating the sensitivity and specificity of the screening, evaluating the performance of different testing methods, and analyzing the geographical distribution of the screening results, etc. Based on the results of the statistical analysis, a customized screening report is generated. This report details the test results, risk levels, and subsequent recommended measures for each augmented personnel. The recommended measures include whether further clinical diagnosis or genetic sequencing is required.

[0071] The above are only embodiments of the present invention. Specific structures and common knowledge such as characteristics well known in the art are not described in detail here. Those of ordinary skill in the art know all the common general technical knowledge in the technical field to which the invention belongs before the application date or the priority date, can know all the prior arts in this field, and have the ability to apply conventional experimental means before this date. Those of ordinary skill in the art can, under the inspiration given in this application, complete and implement this solution in combination with their own abilities. Some typical well-known structures or well-known methods should not become obstacles for those of ordinary skill in the art to implement this application. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several modifications and improvements can still be made, and these should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application should be based on the content of its claims, and the specific implementation manners described in the specification can be used to interpret the content of the claims.

Claims

1. An auxiliary system for thalassemia screening, characterized in that: It includes the following modules: A data collection and preprocessing module, which is used to collect population data of the current area. The population data includes geographical grid division, family network relationships, and basic population information. The basic population information includes age, gender, ethnicity, occupation, and place of residence, and preprocesses the data; A core population screening module, which is used to set screening conditions based on multi-dimensional characteristics such as family relationships, geographical distributions, work natures, ethnic characteristics, and ages, generate a genetic family pedigree representing genetic relationships based on family relationships, analyze and screen out the core population. The core population is a population in which the number of gene-related populations in the current area within the genetic family pedigree is greater than a preset value; A gene inference and expanded personnel identification module: which is used to, after performing thalassemia detection on the core population, according to the gene information of the core population, use the single-gene inheritance law to predict potential risk populations, deduce potential thalassemia carriers or high-risk expanded personnel, then construct a gene relationship map, comprehensively consider the family genetic history and geographical distribution factors in the gene relationship map, evaluate the individual risk level, set a risk threshold, and screen out the list of expanded personnel. The list of expanded personnel is the personnel who have a deduced relationship of genetic inheritance law with the patients within the core population; A simplified parameter detection module, which is used to perform thalassemia detection on the expanded personnel based on simplified parameters, use a convolutional neural network to construct a deep learning model, and the learning model is trained based on data of thalassemia patients in the same region or of the same type. Based on the training results, generate several simplified detection parameter sets classified by region and thalassemia type, and judge the simplified detection parameter set adopted by the list of expanded personnel corresponding to the patients within the core population through the detection results of the patients within the core population; A result analysis and feedback module, which is used to statistically analyze the thalassemia detection results based on simplified parameters, generate a screening feedback report, evaluate the sensitivity and specificity of the screening according to the preset detection strategy, randomly select 10 people from the thalassemia detection personnel based on simplified parameters for re-examination, verify the accuracy, adjust the thalassemia detection strategy, and finally summarize the screening data; A system optimization and maintenance module: which is used to regularly evaluate and optimize the performance of the system, and update the algorithm model and database.

2. The thalassemia screening assistance system according to claim 1, characterized in that: The data collection and preprocessing module includes a preprocessing unit, and the preprocessing unit is used to clean, de-duplicate, and format the basic population information, verify the integrity and consistency of the data. The methods for verifying the integrity and consistency of the data include identity card number and age range verification, identify and mark missing values or outliers, and perform filling or deletion. The missing values include missing identity cards, missing occupations, missing regions, and missing ages. The outliers include ages greater than 75 years old and less than 2 years old.

3. The thalassemia screening assistance system according to claim 2, characterized in that: The data collection and preprocessing module also includes an encryption unit, and the encryption unit is used to encrypt the population data using a symmetric encryption algorithm. The symmetric encryption algorithms include AES and DES.

4. The thalassemia screening assistance system according to claim 3, wherein: The data collection and preprocessing module also includes a user interaction unit, which is used as an operation interface for system administrators, medical professionals, and the public to input data, view screening results, and receive system feedback, and for medical professionals to adjust thalassemia detection strategies and query historical data through the user interaction unit.

5. The thalassemia screening assistance system according to claim 4, wherein: The core population screening module also includes a machine learning algorithm unit, which is used to identify high-risk families and regions using clustering analysis, decision trees, and vector machines.

6. The thalassemia screening assistance system according to claim 5, wherein: The setting methods of screening conditions include whether there are known thalassemia patients in the family, whether there is a specific ethnic background, whether living in a high-incidence area, marriage and childbearing history, personal health status, diet and nutrition status, environmental exposure factors, drug use history, and medical service utilization.

7. The thalassemia screening assistance system according to claim 6, characterized in that: The gene inference and expanded personnel identification module also includes a genomics database for inputting and storing historical data of thalassemia patients. The gene inference and expanded personnel identification module uses the genomics database for comparison to construct a gene relationship map.

8. The thalassemia screening assistance system according to claim 7, wherein: The simplified parameter detection module includes an automated parameter adjustment unit for automatically adjusting the detection parameter set based on the data feedback during the thalassemia detection process of the expanded personnel.

9. The thalassemia screening assistance system according to claim 8, wherein: The simplified parameter detection module also includes a transfer learning unit, which is used to apply the trained deep learning model to thalassemia samples in new regions or new types.

10. An auxiliary method for thalassemia screening, which is based on the method of the thalassemia screening auxiliary system according to any one of claims 1 to 9, and is characterized in that: It includes the following steps: Step 1, data collection and preprocessing: Collect and preprocess the population data in the region and encrypt it; Step 2, core population screening: Classify the population data using machine learning algorithms to identify the core population in high-risk families or regions; Step 3, gene inference and expanded personnel identification: Use the core personnel to construct a gene relationship map, predict potential risk populations using the single-gene inheritance law, and screen out the expanded personnel who need further testing; Step 4, simplified parameter detection: When the detection object is on the expanded personnel list, use the detection steps based on simplified parameters for the expanded personnel; Step 5, result analysis and feedback: Statistically analyze the detection results and generate a customized screening report.