A method for determining influencing factors and related equipment
By extracting samples and labels from physical examination data, initializing individual characteristics of the sheep, and iteratively identifying the leader, the problem of mining the influencing factors of the disease from physical examination data is solved, and the efficiency of medical research and diagnosis and treatment is improved.
Patent Information
- Application Number
- CN202210162886.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-22
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2042-02-22
AI Technical Summary
How to mine out the influencing factors of a certain disease from a large amount of physical examination data to promote the rapid development of medical research, improve the decision-making efficiency of clinicians, and reduce drug treatment accidents.
By obtaining the physical examination samples and their actual classification tags under the target disease, the individual representation characteristics of individuals in the flock are initialized, the fitness value is determined, the leader is selected, and the individual representation characteristics are updated through iteratively until the preset stop condition is reached, and the influencing factors of the target disease are determined.
The best influencing factors of the target disease are gradually found from the physical examination data, and the decision-making efficiency of medical research and the accuracy of clinical diagnosis and treatment are improved.
Smart Images

Figure CN114678135B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a method for determining influencing factors and related equipment. Background Art
[0002] In recent years, with the rapid development of medical informatization, medical big data has become available for storage and analysis. Data mining technology has been applied to the healthcare sector, unlocking the knowledge embedded in medical data and applying the discovered knowledge or rules to assist in diagnosis and treatment. This technology has significant social significance and commercial value. To facilitate understanding, the following examples illustrate this concept.
[0003] For example, for the physical examination data collected from the physical examination population, the influencing factors of a certain disease (for example, fatty liver, etc.) can be mined from these physical examination data to promote the rapid development of medical research, play an assisting and auxiliary role in decision-making, improve the decision-making efficiency of clinicians, and reduce drug treatment accidents.
[0004] However, how to explore the influencing factors of a certain disease is a technical problem that needs to be solved urgently. Summary of the Invention
[0005] In view of this, an embodiment of the present application provides an influencing factor determination method and related equipment, which can achieve the purpose of mining the influencing factors of a certain disease.
[0006] To solve the above problems, the technical solutions provided in the embodiments of the present application are as follows:
[0007] An embodiment of the present application provides a method for determining influencing factors, the method comprising: obtaining at least one physical examination sample and an actual classification label of the at least one physical examination sample under a target disease; wherein the physical examination sample comprises an index value of at least one candidate index; initializing flock characterization data of a flock to be used; wherein the flock characterization data of the flock to be used comprises the number of individuals in the flock to be used and an individual representation feature of each individual in the flock to be used; the individual representation feature is used to represent the degree of correlation between each of the candidate indexes and the target disease; determining the number of individuals in the flock to be used based on the at least one physical examination sample, the actual classification label of the at least one physical examination sample under a target disease, and the individual representation feature of each individual in the flock to be used. The fitness value of each individual; according to the fitness value of each individual in the flock to be used, determining the leader of the flock to be used from all individuals in the flock to be used; according to the fitness value of each individual in the flock to be used and the leader of the flock to be used, updating the individual representation characteristics of each individual in the flock to be used, and continuing to perform the step of determining the fitness value of each individual in the flock to be used according to the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the individual representation characteristics of each individual in the flock to be used; until after it is determined that the preset stop condition is reached, determining at least one influencing factor of the target disease according to the individual representation characteristics of the leader in the flock to be used.
[0008] The present application also provides an influencing factor determination device, including:
[0009] A sample acquisition unit, configured to acquire at least one physical examination sample and an actual classification label of the at least one physical examination sample under a target disease; wherein the physical examination sample includes an indicator value of at least one candidate indicator;
[0010] A flock initialization unit, configured to initialize flock characterization data of the flock to be used; wherein the flock characterization data of the flock to be used includes the number of individuals in the flock to be used and individual representation characteristics of each individual in the flock to be used; the individual representation characteristics are used to indicate the degree of correlation between each candidate indicator and the target disease;
[0011] a fitness determination unit, configured to determine a fitness value of each individual in the to-be-used flock of sheep based on the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the individual representative characteristics of each individual in the to-be-used flock of sheep;
[0012] a leader sheep determining unit, configured to determine a leader sheep in the sheep flock to be used from all individuals in the sheep flock to be used according to the fitness value of each individual in the sheep flock to be used;
[0013] a flock updating unit, configured to update the individual representation characteristics of each individual in the to-be-used flock according to the fitness value of each individual in the to-be-used flock and the leader in the to-be-used flock, and return the step of determining the fitness value of each individual in the to-be-used flock according to the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the individual representation characteristics of each individual in the to-be-used flock;
[0014] The factor determination unit is used to determine at least one influencing factor of the target disease according to the individual representative characteristics of the leading sheep in the flock to be used after determining that the preset stop condition is reached.
[0015] An embodiment of the present application also provides an influencing factor determination device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements any implementation of the influencing factor determination method provided in the embodiment of the present application.
[0016] An embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored. When the instructions are executed on a terminal device, the terminal device executes any implementation of the influencing factor determination method provided in the embodiment of the present application.
[0017] The embodiments of the present application further provide a computer program product. When the computer program product is run on a terminal device, the terminal device executes any implementation of the method for determining influencing factors provided in the embodiments of the present application.
[0018] It can be seen that the embodiments of the present application have the following beneficial effects:
[0019] In the technical solution provided by the embodiment of the present application, first, at least one physical examination sample (for example, the index value of each candidate index) and the actual classification label of each physical examination sample under the target disease are extracted from a large amount of physical examination data, and the individual representation characteristics of each individual in the flock to be used are initialized so that the individual representation characteristics can represent the degree of correlation between each candidate index and the target disease; secondly, based on these physical examination samples and their actual classification labels, as well as the individual representation characteristics of each individual in the flock to be used, the fitness value of each individual is determined so that the fitness value can represent the universality of the influencing factors represented by the individual to the target disease; then, based on The fitness values of these individuals are used to determine the leader in the flock to be used; finally, the individual representation characteristics of these individuals are updated according to the fitness values of these individuals and the leader, and the process returns to continue executing the above-mentioned step of "determining the fitness value of each individual according to these physical examination samples and their actual classification labels, as well as the individual representation characteristics of each individual in the flock to be used" until it is determined that the preset stopping conditions are met, and the influencing factors of the target disease are determined according to the individual representation characteristics of the leader in the flock to be used. In this way, the optimal influencing factors of the target disease can be gradually found through an iterative process, thereby achieving the purpose of mining the influencing factors of a certain disease. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A flowchart of a method for determining influencing factors provided in an embodiment of the present application;
[0021] Figure 2 A schematic diagram of part of the contents of a physical examination report provided in an embodiment of the present application;
[0022] Figure 3 A schematic diagram of a binarization process provided in an embodiment of the present application;
[0023] Figure 4 A schematic diagram of factors that cause ordinary sheep to approach a leading sheep, provided in an embodiment of the present application;
[0024] Figure 5 A schematic diagram of a process for determining influencing factors provided in an embodiment of the present application;
[0025] Figure 6 A schematic diagram of the structure of an influencing factor determination device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0027] The inventors discovered through their research on large amounts of physical examination data that these data contain indicators related to certain diseases (e.g., fatty liver disease). Therefore, they were able to identify the influencing factors of these diseases from these data. This could promote the rapid development of medical research, assist in decision-making, improve clinicians' decision-making efficiency, and reduce medication errors. However, identifying the influencing factors of a disease from these large amounts of physical examination data remains a pressing technical challenge.
[0028] Based on the above findings, in order to solve the technical problems described in the background technology part, the embodiment of the present application also provides a method for determining influencing factors, which includes: first, extracting at least one physical examination sample (for example, the index value of each candidate index) and the actual classification label of each physical examination sample under the target disease from a large amount of physical examination data, and initializing the individual representation characteristics of each individual in the flock to be used, so that the individual representation characteristics can represent the degree of correlation between each candidate index and the target disease; secondly, determining the fitness value of each individual based on these physical examination samples and their actual classification labels, and the individual representation characteristics of each individual in the flock to be used, so that the fitness value can represent the influence represented by the individual. The universality of the influencing factors on the target disease; then, according to the fitness values of these individuals, the leader in the flock to be used is determined; finally, according to the fitness values of these individuals and the leader, the individual representation characteristics of these individuals are updated, and the above step of "determining the fitness value of each individual according to these physical examination samples and their actual classification labels, and the individual representation characteristics of each individual in the flock to be used" is returned to continue to execute until it is determined that the preset stopping condition is met, and the influencing factors of the target disease are determined according to the individual representation characteristics of the leader in the flock to be used. In this way, the optimal influencing factors of the target disease can be gradually found through an iterative process, thereby achieving the purpose of mining the influencing factors of a certain disease.
[0029] In addition, the embodiments of the present application do not limit the execution entity of the influencing factor determination method. For example, the influencing factor determination method provided in the embodiments of the present application can be applied to data processing devices such as terminal devices or servers. The terminal device can be a smartphone, a computer, a personal digital assistant (PDA), or a tablet computer. The server can be a standalone server, a cluster server, or a cloud server.
[0030] To facilitate understanding of the present application, the method for determining influencing factors provided in the embodiments of the present application is described below with reference to the accompanying drawings.
[0031] See also Figure 1 , which is a flow chart of a method for determining influencing factors provided by an embodiment of the present application. The method for determining influencing factors may include S1-S7:
[0032] S1: Obtain at least one physical examination sample and an actual classification label of the at least one physical examination sample under a target disease.
[0033] The above-mentioned “physical examination sample” is used to record a physical examination result of a certain person; and the “physical examination sample” includes an indicator value of at least one candidate indicator.
[0034] In addition, the embodiments of the present application are not limited to the above-mentioned "at least one candidate indicator". For example, it may include the following 73 test indicators: gender, age, white blood cell count, granulocyte count, lymphocyte count, monocyte count, eosinophil count, basophil count, granulocyte ratio, lymphocyte ratio, monocyte ratio, eosinophil ratio, basophil ratio, red blood cell count, hemoglobin concentration, hematocrit measurement, mean corpuscular volume, mean corpuscular Hb content, mean corpuscular Hb concentration, red blood cell distribution width CV, red blood cell distribution width SD, platelet count, platelet distribution width, large platelet ratio, mean platelet ratio, mean platelet volume, serum aspartate aminotransferase measurement, serum alanine aminotransferase measurement, serum alkaline phosphatase measurement, Serum gamma-glutamyl transferase, serum total protein, serum albumin, serum total bile acid, serum total bilirubin, urea, creatinine, glucose (fasting), serum triglycerides, serum total cholesterol, serum high-density lipoprotein cholesterol, serum low-density lipoprotein cholesterol, serum uric acid, diastolic blood pressure, pulse, systolic blood pressure, occult blood, specific gravity, red blood cells, red blood cells per high-power field, protein, glucose, ketone bodies, urobilinogen, bilirubin, nitrite, pH, leukocyte esterase, white blood cells, bacteria, white blood cells per high-power field, epithelial cells per high-power field, epithelial cells, physiological casts, height, weight, body mass index, waist circumference, bacteria per high-power field, physiological casts per low-power field, carcinoembryonic antigen, 20-minute DPM value of C14 breath test.
[0035] In addition, the embodiment of the present application does not limit the representation of the above-mentioned "physical examination sample". For example, it can be represented by formula (1).
[0036]
[0037] Where, d i represents the i-th physical examination sample; It represents the index value of the nth candidate index recorded in the i-th physical examination sample (that is, the test result for the n-th candidate index); n is a positive integer, n≤N, N is a positive integer, and N represents the number of candidate indicators; i is a positive integer, i≤I, I is a positive integer, and I represents the number of physical examination samples.
[0038] It should be noted that the embodiments of the present application are not limited to the above For example, it may be the index value of the nth candidate index (eg, Figure 2 5.93 shown). For another example, it can also be represented in a two-tuple form (for example, Figure 2 As shown, wait).
[0039] In addition, the embodiment of the present application does not limit the acquisition process of the above-mentioned "i-th physical examination sample", for example, it can be specifically as follows: first extract a physical examination result (such as Figure 2 Then, the i-th physical examination sample is extracted from the physical examination result so that the i-th physical examination sample can represent the physical examination result. The physical examination database is used to store a large number of physical examination results that have been completed in at least one medical institution.
[0040] It should be noted that the embodiment of the present application does not limit the extraction process of the above-mentioned "i-th physical examination sample". For example, when the physical examination results are Figure 2 When storing in the storage method shown (i.e., column method), the extraction process of the "i-th physical examination sample" can be specifically: converting the different indicators recorded in the columns of the physical examination results into different fields of a physical examination data to obtain the i-th physical examination sample.
[0041] In fact, the inventors found in their research on physical examination results that the physical examination results may have the following three situations shown in ①-③: ① Some indicators are not common test indicators (for example, serum testosterone determination, urine microalbumin determination, etc.), which makes the filling rate of these uncommon test indicators in a large number of physical examination results relatively low (for example, the filling rate of the above-mentioned "serum testosterone determination" is only 0.35%, the filling rate of the above-mentioned "urine microalbumin determination" is only 0.77%, etc.), resulting in these uncommon test indicators usually having no reference value, so in order to avoid these test indicators causing adverse effects, they should be deleted. ② The index values of some test indicators are usually text strings (for example, "++", "2+" and other strings), and text strings are usually not conducive to data processing, so in order to better use these index values, these index values can be converted into digital variables. ③ The index values of some test indicators (especially the more common test indicators) may be missing index values, so in order to avoid the adverse effects caused by these missing values, these missing values can be filled.
[0042] Based on the above findings, the present embodiment further provides a possible implementation method for obtaining the above-mentioned “i-th physical examination sample”, which may specifically include steps 11 to 14:
[0043] Step 11: Obtain the physical examination results to be used from the physical examination database.
[0044] The above-mentioned “physical examination results to be used” refers to the physical examination results required to be used when constructing the i-th physical examination sample; and the embodiment of the present application does not limit the “physical examination results to be used”, for example, it may include Figure 2 The content shown.
[0045] Step 12: Determine the i-th physical examination data to be used from the physical examination results to be used.
[0046] The above-mentioned "i-th physical examination data to be used" is used to represent the test results under various test indicators recorded in the physical examination results to be used; and the embodiment of the present application does not limit the "i-th physical examination data to be used", for example, it may include the indicator value of at least one test indicator.
[0047] In addition, the embodiment of the present application does not limit the representation method of the above-mentioned "i-th physical examination data to be used", for example, it can be represented by formula (2).
[0048]
[0049] Where D i Indicates the i-th physical examination data to be used; It represents the index value of the h-th test index recorded in the i-th physical examination data to be used (that is, the test result for the h-th test index); h is a positive integer, h≤H, H is a positive integer, H≥N, H represents the number of test indicators involved in the i-th physical examination data to be used (that is, the number of test indicators recorded in the physical examination result to be used); i is a positive integer, i≤I, I is a positive integer, and I represents the number of physical examination samples.
[0050] It should be noted that the embodiments of the present application are not limited to the above For example, it can be the index value of the hth test index (for example, Figure 2 5.93 shown). For another example, it can also be represented in a two-tuple form (for example, Figure 2 As shown, wait).
[0051] Step 13: According to the preset data cleaning rules, the i-th physical examination data to be used is cleaned to obtain the i-th physical examination sample.
[0052] The above-mentioned “preset data cleaning rules” can be pre-set; and the embodiment of the present application is not limited to the “preset data cleaning rules”. For example, they can specifically include the following three rules:
[0053] Rule 1: Eliminate the indicator value of the indicator to be eliminated; and the filling rate of the indicator to be eliminated is less than a preset filling rate threshold (for example, 30%).
[0054] It should be noted that the embodiments of the present application do not limit the process of determining the above-mentioned "filling rate of indicators to be eliminated". For example, all physical examination results in the physical examination database can be used to determine the "filling rate of indicators to be eliminated", so that the "filling rate of indicators to be eliminated" can represent the ratio between the number of indicators to be eliminated with valid indicator values in the physical examination database and the total number of indicators to be eliminated. For another example, when the above-mentioned "at least one physical examination sample" includes I physical examination samples, and the I physical examination samples are determined from I physical examination reports, the process of determining the "filling rate of indicators to be eliminated" can be: using the I physical examination reports to determine the "filling rate of indicators to be eliminated", so that the "filling rate of indicators to be eliminated" can represent the ratio between the number of indicators to be eliminated with valid indicator values in the I physical examination reports and the total number of indicators to be eliminated.
[0055] Rule 2: Convert indicator values represented by text strings into numerical values.
[0056] It should be noted that the embodiments of the present application are not limited to the conversion process of the above-mentioned "numeric conversion processing". For example, it can be implemented by any method that can convert text into a numeric value. For another example, it can also be implemented with the help of a pre-built mapping list for recording the correspondence between different texts and different numeric values.
[0057] Rule 3: When the index value of a certain test index in the above-mentioned “i-th physical examination data to be used” is in a missing state, the preset index value corresponding to the test index can be used to fill in the missing value of the test index.
[0058] It should be noted that the embodiments of the present application do not limit the above-mentioned "preset indicator value". For example, if the h-th test indicator has a normal range value, the median value of the normal range of the h-th test indicator can be determined as the preset indicator value corresponding to the h-th test indicator. For another example, if the h-th test indicator does not have a normal range value, the median corresponding to the h-th test indicator can be determined as the preset indicator value corresponding to the h-th test indicator. Wherein, h is a positive integer, h≤H, and H is a positive integer.
[0059] It should also be noted that the embodiments of the present application do not limit the determination process of the above-mentioned "median corresponding to the h-th test indicator". For example, it can be obtained by performing median statistical processing on all indicator values under the h-th test indicator recorded in all physical examination results in the physical examination database. For another example, when the above-mentioned "at least one physical examination sample" includes I physical examination samples, and the I physical examination samples are determined from I physical examination reports, the determination process of the "median corresponding to the h-th test indicator" can specifically be: performing median statistical processing on all indicator values under the h-th test indicator recorded in the I physical examination reports to obtain the median corresponding to the h-th test indicator.
[0060] Based on the relevant contents of steps 11 to 13 above, it can be known that after obtaining the physical examination results to be used from the physical examination database, the i-th physical examination data to be used can be first determined from the physical examination results to be used, so that the i-th physical examination data to be used can represent the indicator values under each test indicator recorded by the physical examination results to be used; then, according to the preset data cleaning rules, the i-th physical examination data to be used is cleaned to obtain the i-th physical examination sample, so that the i-th physical examination sample does not contain indicator values of indicators to be eliminated with a relatively low filling rate, indicator values represented in the form of text strings, and indicator values in a missing state. In this way, the i-th physical examination sample can avoid the interference caused by test indicators with a relatively low filling rate, test indicators represented in the form of text strings, and test indicators with missing indicator values, which is conducive to improving the effect of extracting influencing factors of the target disease.
[0061] The above-mentioned “target disease” refers to a disease that requires influencing factor mining and processing; and the embodiment of the present application is not limited to the “target disease”, for example, it can be fatty liver.
[0062] The "actual classification label of the i-th physical examination sample under the target disease" is used to indicate whether the physical examinee with the indicator value recorded by the i-th physical examination sample actually suffers from the target disease; and the embodiment of the present application does not limit the acquisition process of the "actual classification label of the i-th physical examination sample under the target disease". For ease of understanding, the following is an explanation with examples.
[0063] As an example, when the above-mentioned "i-th physical examination sample" is determined based on the physical examination result to be used, and the above-mentioned "target disease" is fatty liver, the acquisition process of the above-mentioned "actual classification label of the i-th physical examination sample under the target disease" can be specifically as follows: first obtain the physical examination report text corresponding to the physical examination result to be used; then search the physical examination report text for the presence of the three words "fatty liver". If the three words "fatty liver" exist, it can be determined that the examinee with the physical examination result to be used suffers from fatty liver, so the actual classification label of the i-th physical examination sample under the target disease can be The label is set to 1 so that the group of data ("the i-th physical examination sample", "the actual classification label of the i-th physical examination sample under the target disease") can be used as a positive sample later; however, if the three words "fatty liver" do not exist, it can be determined that the examinee with the physical examination result to be used does not suffer from fatty liver, so the actual classification label of the i-th physical examination sample under the target disease can be set to 0, so that the group of data ("the i-th physical examination sample", "the actual classification label of the i-th physical examination sample under the target disease") can be used as a negative sample later.
[0064] It should be noted that the above-mentioned “physical examination report text corresponding to the physical examination result to be used” refers to the physical examination report written based on the physical examination result to be used.
[0065] Based on the relevant content of S1 above, it can be known that when you want to mine the influencing factors of the target disease (for example, fatty liver, etc.) from a large amount of physical examination data, you can obtain at least one physical examination sample and the actual classification label of the at least one physical examination sample under the target disease from the large number of physical examination results stored in the physical examination database and their corresponding physical examination report texts, so that you can subsequently use these physical examination samples and their actual classification labels under the target disease to construct a large number of positive and negative samples (for example, 2000 positive samples + 2000 negative samples, etc.) for use.
[0066] S2: Initialize the flock characterization data of the flock to be used.
[0067] The above-mentioned “flock characterization data of the flock to be used” is used to characterize the characteristics of the flock to be used; and the “flock characterization data of the flock to be used” may include the number of individuals in the flock to be used and the individual representation characteristics of each individual in the flock to be used.
[0068] The above-mentioned "number of individuals in the flock to be used" is used to indicate how many individuals are in the flock to be used; and the embodiment of the application does not limit the "number of individuals in the flock to be used", for example, it can be M (for example, M=500). Wherein, M is a positive integer.
[0069] The "individual representative characteristics of the mth individual in the flock to be used" are used to describe the characteristics of the mth individual in the flock to be used, so that the "individual representative characteristics of the mth individual in the flock to be used" can express the degree of correlation between each candidate indicator and the target disease, so that the mth individual can represent the impact of all candidate indicators on the target disease. Wherein, m is a positive integer, m≤M, and M is a positive integer.
[0070] In addition, the embodiment of the present application does not limit the above-mentioned "individual representation characteristics of the mth individual in the flock to be used". For example, it may include a representation value of the degree of relevance of at least one candidate indicator to the target disease. It can be seen that when the above-mentioned "at least one candidate indicator" includes N candidate indicators, the "individual representation characteristics of the mth individual in the flock to be used" may include the representation values of the degree of relevance of the N candidate indicators to the target disease, so that the "individual representation characteristics of the mth individual in the flock to be used" can represent an impact of the N candidate indicators on the target disease.
[0071] In addition, the embodiment of the present application does not limit the representation method of the above-mentioned "individual representation characteristics of the mth individual in the flock to be used". For example, it can be represented in the vector form shown in formulas (3)-(4) so that the vector dimension of the "individual representation characteristics of the mth individual in the flock to be used" is equal to the number of candidate indicators.
[0072] SEP m =(e m,1 ,e m,2 ,…,e m,N ) (3)
[0073] e m,n =random(0,1) (4)
[0074] Where SEP m represents the individual representation characteristics of the mth individual in the flock to be used; e m,n Indicates the correlation value of the nth candidate indicator in the above “individual representation characteristics of the mth individual in the flock to be used” to the target disease, so that the e m,n It can show the correlation between the nth candidate indicator and the target disease, and the e m,n The value can be a random number that obeys a uniform distribution in the range of (0,1); n is a positive integer, n≤N, N is a positive integer, and N represents the number of candidate indicators; random(0,1) means extracting a random number that obeys a uniform distribution from the range of (0,1).
[0075] Based on the relevant content of S2 above, it can be known that if you want to mine the influencing factors of the target disease (for example, fatty liver, etc.) from a large amount of physical examination data, you can initialize the flock characterization data of the flock to be used so that the "flock characterization data of the flock to be used" can simulate the different impacts of all candidate indicators on the target disease, so that you can subsequently determine the influencing factors of the target disease based on the "flock characterization data of the flock to be used".
[0076] It should be noted that the embodiment of the present application does not limit the execution order of S1 and S2. For example, S1 and S2 can be executed in sequence, S2 and S1 can be executed in sequence, or S1 and S2 can be executed simultaneously.
[0077] S3: Determine the fitness value of each individual in the sheep flock to be used according to the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the individual representation feature of each individual in the sheep flock to be used.
[0078] The "fitness value of the mth individual in the flock to be used" indicates the universality of the impact of all candidate indicators represented by the mth individual on the target disease. A larger "fitness value of the mth individual in the flock to be used" indicates a higher universality of the impact of all candidate indicators represented by the mth individual on the target disease. m is a positive integer, m≤M, where M is a positive integer.
[0079] In addition, the embodiment of the present application does not limit the above-mentioned determination process of the "fitness value of the mth individual in the flock to be used". For example, it can be implemented by any of the implementation methods for determining the "fitness value of the mth individual in the flock to be used" shown below.
[0080] Based on the relevant content of S3 above, it can be seen that after obtaining a large number of physical examination samples and their actual classification labels under the target disease, as well as the individual representation characteristics of each individual in the flock to be used, these physical examination samples and their actual classification labels under the target disease can be used to verify the universality of the influence of all candidate indicators represented by these individual representation characteristics on the target disease, and obtain the fitness values of these individuals, so that the individuals with the highest universality can be determined based on these fitness values in the future.
[0081] S4: According to the fitness value of each individual in the flock of sheep to be used, a leader in the flock of sheep to be used is determined from all individuals in the flock of sheep to be used.
[0082] Here, the "leader" is used to represent the individual with the highest universality among all individuals in the flock to be used.
[0083] In addition, the embodiment of the present application does not limit the above-mentioned "leader" determination process. For example, it can be specifically as follows: first, the fitness values of all individuals in the flock to be used are subjected to a maximum analysis to obtain the highest fitness value; then, the individual with the highest fitness value in the flock to be used is determined as the leader in the flock to be used.
[0084] Based on the relevant content of S4 above, it can be known that after obtaining the fitness value of each individual in the flock to be used, the individual with the highest universality can be selected from all individuals in the flock to be used based on these fitness values and determined as the leader in the flock to be used, so that the individual representation characteristics of each individual in the flock to be used can be updated based on the leader.
[0085] S5: Determine whether the preset stop condition is met. If so, execute S7; if not, execute S6.
[0086] The above-mentioned "preset stopping condition" can be pre-set; and the embodiments of the present application do not limit the "pre-set stopping condition". For example, it can be: the number of iterations reaches a preset number threshold, or the fitness value of the leader in the to-be-used flock reaches a preset fitness threshold (for example, the preset fitness threshold is 0.8569). The preset number threshold and the preset fitness threshold can both be pre-set.
[0087] S6: Update the individual representation features of each individual in the to-be-used flock according to the fitness value of each individual in the to-be-used flock and the leader of the to-be-used flock, and return to execute S3 and subsequent steps.
[0088] In an embodiment of the present application, when it is determined that the current iteration has not yet reached the preset stop condition, it can be determined that the universality of the impact of all candidate indicators represented by the leader determined by the current iteration on the target disease is still relatively low. Therefore, in order to further improve the accuracy of the influencing factors of the target disease, the fitness value of each individual in the flock to be used and the leader in the flock to be used can be referred to, and the individual representation characteristics of each individual in the flock to be used can be updated, so that S3 and its subsequent steps can be continued to be executed based on the updated individual representation characteristics of these individuals to realize the next round of leader screening process. In this way, a leader with a relatively high universality can be gradually found through multiple rounds of iterative processes, thereby achieving the purpose of gradually finding the best influencing factors of the target disease.
[0089] It should be noted that the embodiment of the present application does not limit the updating process of the above-mentioned "individual representation characteristics of each individual in the flock to be used". For example, it can be implemented by any of the implementation methods of updating the "individual representation characteristics of each individual in the flock to be used" shown below.
[0090] S7: Determine at least one influencing factor of the target disease based on the individual representative characteristics of the leader sheep in the flock to be used.
[0091] In an embodiment of the present application, when it is determined that the current iteration has reached the preset stop condition, it can be determined that the universality of the influence of all candidate indicators represented by the leader determined in the current iteration on the target disease is relatively high, so the individual representation characteristics of the leader can be directly referenced to determine at least one influencing factor of the target disease (for example, α candidate indicators with relatively high correlation representation values in the individual representation characteristics of the leader can be determined as influencing factors of the target disease), so that these influencing factors of the target disease are more correlated with the target disease, thereby achieving the purpose of mining the influencing factors of the target disease from a large number of physical examination results. For example, the influencing factors of fatty liver obtained by screening may include the following 16 influencing factors: body mass index, serum triglyceride measurement, weight, waist circumference, serum uric acid measurement, serum gamma-glutamyl transferase measurement, serum high-density lipoprotein cholesterol measurement, serum alanine aminotransferase measurement, glucose measurement (fasting), hemoglobin concentration, red blood cell count, gender, diastolic blood pressure, hematocrit measurement, systolic blood pressure, and creatinine measurement.
[0092] It should be noted that the embodiment of the present application does not limit α, for example, α=16; and α can be preset.
[0093] Based on the relevant contents of S1 to S7 above, it can be known that for the influencing factor determination method provided in the embodiment of the present application, first, at least one physical examination sample (for example, the index value of each candidate indicator) and the actual classification label of each physical examination sample under the target disease are extracted from a large amount of physical examination data, and the individual representation characteristics of each individual in the flock to be used are initialized, so that the individual representation characteristics can represent the degree of correlation between each candidate indicator and the target disease; secondly, based on these physical examination samples and their actual classification labels, as well as the individual representation characteristics of each individual in the flock to be used, the fitness value of each individual is determined, so that the fitness value can represent the influence of the influencing factor represented by the individual on the target disease. The universality of the disease; then, according to the fitness values of these individuals, the leader in the flock to be used is determined; finally, according to the fitness values of these individuals and the leader, the individual representation characteristics of these individuals are updated, and the above step of "determining the fitness value of each individual according to these physical examination samples and their actual classification labels, and the individual representation characteristics of each individual in the flock to be used" is returned to continue to be executed until it is determined that the preset stopping condition is met, and the influencing factors of the target disease are determined according to the individual representation characteristics of the leader in the flock to be used. In this way, the optimal influencing factors of the target disease can be gradually found through an iterative process, thereby achieving the purpose of mining the influencing factors of a certain disease.
[0094] In one possible implementation, in order to further improve the search effect of influencing factors of the target disease, the present embodiment also provides another possible implementation of S3, which may specifically include S31-S32:
[0095] S31: Binarizing the individual representation features of each individual in the flock to be used to obtain factor selection description data of each individual in the flock to be used.
[0096] The "factor selection description data for the mth individual in the flock to be used" is used to describe the characteristics of the mth individual in the flock to be used using binary data. This allows the "factor selection description data for the mth individual in the flock to be used" to represent the candidate indicator group selected for the target disease, thereby enabling the mth individual to better represent the combined impact of all candidate indicators on the target disease. m is a positive integer, m≤M, and M is a positive integer.
[0097] In addition, the embodiments of the present application do not limit the above-mentioned "factor selection description data of the mth individual in the flock to be used". For example, it may include the screening results of at least one candidate indicator, so that the "factor selection description data of the mth individual in the flock to be used" can represent the impact of these candidate indicators on the target disease in a binary manner (that is, a combination of influencing factors for the target disease). It can be seen that when the above-mentioned "at least one candidate indicator" includes N candidate indicators, the "factor selection description data of the mth individual in the flock to be used" can include the screening results of N candidate indicators.
[0098] In addition, the embodiment of the present application does not limit the representation method of the above-mentioned "factor selection description data of the mth individual in the flock to be used". For example, it can be represented in the vector form shown in formula (5) so that the vector dimension of the "factor selection description data of the mth individual in the flock to be used" is equal to the number of candidate indicators.
[0099] CLS m =(c m,1 ,c m,2 ,…,c m,N ) (5)
[0100] Among them, CLS m Indicates the factor selection descriptive data of the mth individual in the flock to be used; c m,n Indicates the screening result of the nth candidate indicator in the above “factor selection description data of the mth individual in the flock to be used”, so that the c m,n It can indicate whether to select the nth candidate indicator as the influencing factor of the target disease, and if e m,nIf the preset screening conditions are met, the nth candidate indicator is selected as the influencing factor of the target disease, so the c m,n is a first value (e.g., 1), if e m,n If the preset screening conditions are not met, the nth candidate indicator will not be selected as the influencing factor of the target disease, so the c m,n is a second value (for example, 0); n is a positive integer, n≤N, N is a positive integer, and N represents the number of candidate indicators.
[0101] The above-mentioned “preset screening conditions” are used to select the influencing factors of the target disease from N candidate indicators; and the embodiment of the present application is not limited to the above-mentioned “preset screening conditions”. For example, if it is desired to determine α influencing factors for the target disease, the “preset screening conditions” can be specifically selected from SEP m Select α relatively high correlation representation values from the SEP, so that the "preset screening condition" can be realized based on the SEP m The purpose of selecting candidate indicators with higher correlation values for the target disease from N candidate indicators.
[0102] It can be seen that after obtaining the above-mentioned "individual representation characteristics of the mth individual in the flock to be used" SEP m Afterwards, the SEP can be m All elements e in m,1 、e m,2 , ..., and e m,N Sort by numerical value from small to large to get the first sorting result; if the first sorting result indicates e m,n The arrangement position of e is ≤α, then it can be determined that m,n If the preset screening conditions are met, the nth candidate indicator can be selected as the influencing factor of the target disease, so the c m,n is the first value (eg, 1); however, if the first sorting result indicates e m,n The arrangement position of e is greater than α, then it can be determined that m,n If the preset screening conditions are not met, the nth candidate indicator can be discarded, so the c m,n is a second value (for example, 0). Wherein, n is a positive integer, n≤N, N is a positive integer, and N represents the number of candidate indicators.
[0103] In fact, the above SEP m There may be orders of magnitude differences between the N included correlation degree representation values. Therefore, in order to avoid the interference caused by the order of magnitude differences, the embodiment of the present application also provides another possible implementation method for determining the above-mentioned "factor selection description data of the mth individual in the flock to be used" (such as Figure 3 As shown), it may specifically include S311-S313:
[0104] S311: normalizing the individual representation features of the mth individual in the flock to be used to obtain normalized features of the mth individual.
[0105] The above-mentioned “normalization processing” is used to eliminate the possible order of magnitude differences between multiple data; and the embodiments of the present application do not limit the implementation method of the “normalization processing”. For example, the softmax normalization method can be used for implementation.
[0106] The "normalized characteristics of the mth individual" are used to represent the characteristics of the mth individual in the flock to be used. Furthermore, this embodiment of the present application does not limit the "normalized characteristics of the mth individual" and, for example, may include the normalized impact representation value of at least one candidate indicator. Therefore, when the "at least one candidate indicator" includes N candidate indicators, the "normalized characteristics of the mth individual" may include the normalized impact representation values of the N candidate indicators.
[0107] In addition, the embodiment of the present application does not limit the representation method of the above-mentioned "normalized features of the mth individual". For example, it can be represented in the vector form shown in formulas (6)-(7) so that the vector dimension of the "normalized features of the mth individual" is equal to the number of candidate indicators.
[0108] NS m =(s m,1 ,s m,2 ,…,s m,N ) (6)
[0109]
[0110] Where, NS m represents the normalized characteristics of the mth individual; s m,n Indicates the normalized influence representation value of the nth candidate indicator in the above “normalized features of the mth individual”, so that the s m,n It can represent the normalized correlation degree of the nth candidate indicator to the target disease; n is a positive integer, n≤N, N is a positive integer, and N represents the number of candidate indicators.
[0111] S312: Determine at least one position to be selected that meets a preset selection condition and at least one position to be discarded that does not meet the preset selection condition from the normalized features of the mth individual.
[0112] The above “preset selection condition” can be preset, for example, when the above NS m After all elements in the NS are sorted from large to small according to their values, the "preset selection condition" can be that the elements ranked first α in the above NS mThe position in which the "preset selection condition" can be based on the NS m The purpose of selecting candidate indicators with a higher degree of relevance to the target disease from N candidate indicators.
[0113] It can be seen that after obtaining the above “normalized feature of the mth individual” NS m After that, you can use the NS m All elements s in m,1 、s m,2 , ..., and s m,N Sort by numerical value to get the second sorting result; if the second sorting result indicates s m,n The arrangement position of s is ≤α, then it can be determined that m,n Satisfy the preset selection conditions, so that the nth candidate indicator can be selected as the influencing factor of the target disease, so the s m,n In the above NS m The position in is marked as the position to be selected; however, if the second sorting result indicates s m,n The arrangement position of s is greater than α, then it can be determined that m,n The preset selection conditions are not met, so the nth candidate indicator can be discarded, so the s m,n In the above NS m The position in is marked as the position to be discarded. Where n is a positive integer, n≤N, N is a positive integer, and N represents the number of candidate indicators.
[0114] S313: setting each position to be selected in the normalized feature of the mth individual to a first value, and setting each position to be discarded in the normalized feature of the mth individual to a second value, to obtain factor selection description data of the mth individual in the flock to be used.
[0115] In the embodiment of the present application, after obtaining at least one position to be selected and at least one position to be discarded corresponding to the above-mentioned "normalized feature of the mth individual", each position to be selected in the "normalized feature of the mth individual" can be set to a first value (for example, 1), and each position to be discarded in the "normalized feature of the mth individual" can be set to a second value (for example, 0), and the factor selection description data of the mth individual in the flock to be used is obtained, so that the vector dimension of the "factor selection description data of the mth individual" is consistent with the vector dimension of the "normalized feature of the mth individual", so that the "factor selection description data of the mth individual" can represent the factors based on the NS m The candidate indicators with a relatively high degree of relevance to the target disease are selected from all the candidate indicators, so that the "factor selection description data of the mth individual" can represent the influence of all the candidate indicators on the target disease.
[0116] Based on the relevant content of S31 above, it can be known that after obtaining the individual representation feature of the mth individual in the flock to be used, the individual representation feature of the mth individual can be binarized (for example, Figure 3 The binarization process shown in the figure) is used to obtain the factor selection description data of the m-th individual, so that the "factor selection description data of the m-th individual" can represent the characteristics of the m-th individual in the form of binary data, thereby enabling the "factor selection description data of the m-th individual" to represent the influence of these candidate indicators on the target disease in a binary manner (that is, a combination of influencing factors for the target disease).
[0117] S32: Determine the fitness value of each individual in the sheep flock to be used according to at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the factor selection description data of each individual in the sheep flock to be used.
[0118] As an example, S32 may specifically include S321-S324:
[0119] S321: Determine at least one training data, the actual classification label of the at least one training data, at least one test data, and the actual classification label of the at least one test data based on at least one physical examination sample and the actual classification label of the at least one physical examination sample under the target disease.
[0120] The above-mentioned “training data” is used to represent the physical examination samples required for use in the model building process.
[0121] The above-mentioned “test data” is used to represent the physical examination samples required to be used in the model testing process.
[0122] It should be noted that the embodiments of the present application do not limit the number of the above-mentioned "training data" and the number of the above-mentioned "test data". For example, the sum of the number of the above-mentioned "training data" and the number of the above-mentioned "test data" is equal to the number of the above-mentioned "physical examination samples".
[0123] In addition, the embodiments of the present application do not limit the implementation of S321. For ease of understanding, the following is an illustration with examples.
[0124] As an example, when the above-mentioned "at least one physical examination sample" and the above-mentioned "actual classification label of at least one physical examination sample under the target disease" can construct 2,000 positive samples + 2,000 negative samples, the 2,000 positive samples and the 2,000 negative samples can be shuffled first to obtain 4,000 samples to be divided; then the 4,000 samples to be divided are divided according to a preset division ratio (for example, 4:1, etc.) to obtain a training data set and a test data set, so that the training data set can be used to train the model, and the test data set can be used to test the model; finally, the physical examination samples recorded for each sample to be divided in the training data set and its actual classification label under the target disease are determined as a training data and its actual classification label, and the physical examination samples recorded for each sample to be divided in the test data set and its actual classification label under the target disease are determined as a test data and its actual classification label.
[0125] It should be noted that the embodiment of the present application does not limit the representation method of the above-mentioned "samples to be divided". For example, it can be represented in the manner shown in formula (8).
[0126]
[0127] In the formula, SD r Represents the rth sample to be divided; represents the physical examination sample recorded in the rth sample to be divided; L (r) Represents the actual classification label of the physical examination sample recorded in the rth sample to be divided under the target disease; where r is a positive integer, r≤R, R is a positive integer, and R represents the number of samples to be divided.
[0128] S322: Using the factor selection description data of each individual in the to-be-used flock, perform indicator selection processing on at least one training data to obtain at least one training sample corresponding to each individual in the to-be-used flock.
[0129] The bth training sample corresponding to the mth individual is obtained by using the factor selection description data of the mth individual and performing indicator selection processing on the bth training data. Where m is a positive integer, m≤M, M is a positive integer, b is a positive integer, b≤B, B is a positive integer, and B represents the number of training data.
[0130] In addition, the embodiment of the present application does not limit the determination process of the above-mentioned "bth training sample corresponding to the mth individual". For ease of understanding, it is explained below with reference to three examples.
[0131] Example 1. The process of determining the above-mentioned "bth training sample corresponding to the mth individual" can be specifically as follows: first, the positions of each second numerical value in the above-mentioned "factor selection description data of the mth individual" are determined as positions to be masked; then the values of each position to be masked in the bth training data are replaced with zero to obtain the bth training sample corresponding to the mth individual, so that the "bth training sample corresponding to the mth individual" only retains the indicator values of each candidate indicator selected by the "factor selection description data of the mth individual".
[0132] Example 2: When the above “first value” is 1 and the above “second value” is 0, the above process of determining “the b-th training sample corresponding to the m-th individual” can also be implemented using formulas (9)-(10).
[0133]
[0134]
[0135] Where, represents the bth training sample corresponding to the mth individual; Represents the weighted index value of the nth candidate index in the above “bth training sample corresponding to the mth individual”, so that It can indicate whether the nth candidate indicator is selected as the influencing factor of the target disease in the above-mentioned "factor selection description data of the mth individual"; Represents the bth training data; represents the index value of the nth candidate index recorded in the bth training data (that is, the test result for the nth candidate index); c m,n It represents the screening result of the nth candidate indicator in the above “factor selection description data of the mth individual in the flock to be used”; n is a positive integer, n≤N, N is a positive integer, and N represents the number of candidate indicators; b is a positive integer, b≤B, B is a positive integer, and B represents the number of training data.
[0136] Example 3, the process of determining the above-mentioned "bth training sample corresponding to the mth individual" can be specifically as follows: first, the positions of each second numerical value in the above-mentioned "factor selection description data of the mth individual" are determined as positions to be deleted; then, each position to be deleted in the bth training data is deleted to obtain the bth training sample corresponding to the mth individual, so that the "bth training sample corresponding to the mth individual" is only used to record the indicator values of each candidate indicator selected by the "factor selection description data of the mth individual".
[0137] Based on the relevant content of S322 above, it can be known that after obtaining the factor selection description data of the mth individual in the flock to be used and at least one training data, the factor selection description data of the mth individual can be used to perform indicator selection processing on each training data (that is, specify which candidate indicators' indicator values can be retained and which candidate indicators' indicator values can be discarded) to obtain the various training samples corresponding to the mth individual, so that these sample data can be used subsequently to construct a classification model corresponding to the mth individual.
[0138] S323: Using at least one training sample corresponding to each individual in the to-be-used flock and at least one actual classification label of the training data, the classification model to be processed is trained to obtain a classification model corresponding to each individual in the to-be-used flock.
[0139] The aforementioned "unprocessed classification model" refers to a classification model that needs to be trained; and the embodiments of the present application are not limited to the "unprocessed classification model." For example, it can be implemented using any existing or future classification model. For another example, it can be implemented using a Gaussian kernel-based support vector machine.
[0140] The classification model corresponding to the mth individual is used to perform classification processing based on the candidate indicators selected by the "factor selection description data of the mth individual". Where m is a positive integer, m≤M, and M is a positive integer.
[0141] In addition, the embodiment of the present application does not limit the determination process of the above-mentioned "classification model corresponding to the m-th individual", for example, it may specifically include S3231-S3232:
[0142] S3231: Determine the actual classification label of the b-th training data as the actual classification label of the b-th training sample corresponding to the m-th individual, where m is a positive integer, m≤M, M is a positive integer, b is a positive integer, b≤B, B is a positive integer, and B represents the number of training data.
[0143] S3232: Using the B training samples corresponding to the m-th individual and the actual classification labels of the B training samples corresponding to the m-th individual, train the classification model to be processed to obtain the classification model corresponding to the m-th individual.
[0144] It should be noted that the embodiments of the present application do not limit the implementation of S3232. For example, it can be implemented using any existing or future classification model training method.
[0145] Based on the relevant content of S323 above, it can be known that after obtaining at least one training sample corresponding to the mth individual in the flock to be used and its actual classification label, these training samples and their actual classification labels can be used to train the classification model to be processed to obtain the classification model corresponding to the mth individual in the flock to be used, so that the "classification model corresponding to the mth individual" can have the function of performing classification processing with reference to the various candidate indicators selected by the "factor selection description data of the mth individual".
[0146] S324: Determine the fitness value of each individual in the flock of sheep to be used by using the classification model corresponding to each individual in the flock of sheep to be used, at least one test data, and the actual classification label of the at least one test data.
[0147] As an example, S324 may specifically include S3241-S3243:
[0148] S3241: Determine the mth model classification result of each test data using the classification model corresponding to the mth individual in the flock to be used, where m is a positive integer, m≤M, and M is a positive integer.
[0149] The "mth model classification result for the fth test data" refers to whether the fth test data belongs to the target disease (e.g., fatty liver disease) as predicted by the "classification model corresponding to the mth individual" above. f is a positive integer, f≤F, where F is a positive integer and represents the number of test data.
[0150] In addition, the embodiments of the present application do not limit the determination process of the above-mentioned "mth model classification result of the fth test data". For example, it can be specifically: directly input the fth test data into the classification model corresponding to the mth individual, and obtain the mth model classification result of the fth test data output by the "classification model corresponding to the mth individual".
[0151] In addition, in order to further improve the model classification effect, the embodiment of the present application also provides another possible implementation method for determining the above-mentioned "m-th model classification result of the f-th test data", which may specifically include steps 21 and 22:
[0152] Step 21: Using the factor selection description data of the mth individual in the flock to be used, perform indicator selection processing on the fth test data to obtain the fth test sample corresponding to the mth individual.
[0153] It should be noted that the implementation of step 21 is similar to the implementation of the process of determining “the b-th training sample corresponding to the m-th individual” shown in S322 above.
[0154] Step 22: Input the f-th test sample corresponding to the m-th individual into the classification model corresponding to the m-th individual, and obtain the m-th model classification result of the f-th test data output by the “classification model corresponding to the m-th individual”.
[0155] Based on the relevant content of S3241 above, it can be seen that after obtaining the classification model corresponding to the m-th individual, the "classification model corresponding to the m-th individual" can be used to classify the f-th test data to obtain the m-th model classification result of the f-th test data, so that the "m-th model classification result of the f-th test data" can indicate whether the f-th test data belongs to the target disease. Wherein, f is a positive integer, f≤F, F is a positive integer, and F represents the number of test data.
[0156] S3242: Determine classification performance characterization data of the classification model corresponding to the mth individual based on the mth model classification result of the at least one test data and the actual classification label of the at least one test data.
[0157] The above-mentioned "classification performance characterization data of the classification model corresponding to the mth individual" is used to represent the classification performance of the classification model corresponding to the mth individual; and the embodiment of the present application does not limit the "classification performance characterization data of the classification model corresponding to the mth individual", for example, it may include the F1 score of the classification model corresponding to the mth individual.
[0158] In addition, the embodiment of the present application does not limit the determination process of the above-mentioned "classification performance characterization data of the classification model corresponding to the mth individual" (that is, the implementation method of S3242). For example, it may specifically include S32421-S32424:
[0159] S32421: Determine the accuracy of the classification model corresponding to the mth individual based on the mth model classification result of the at least one test data and the actual classification label of the at least one test data.
[0160] S32422: Determine the recall rate of the classification model corresponding to the mth individual based on the mth model classification result of the at least one test data and the actual classification label of the at least one test data.
[0161] S32423: Determine the F1 score of the classification model corresponding to the mth individual using the precision of the classification model corresponding to the mth individual and the recall of the classification model corresponding to the mth individual.
[0162] S32424: Determine the F1 score of the classification model corresponding to the m-th individual as classification performance representation data of the classification model corresponding to the m-th individual.
[0163] S3243: Determine the classification performance characterization data of the classification model corresponding to the m-th individual as the fitness value of the m-th individual in the flock to be used.
[0164] In an embodiment of the present application, after obtaining the classification performance characterization data of the classification model corresponding to the mth individual, the classification performance characterization data can be directly determined as the fitness value of the mth individual, so that it can represent the influence of the candidate indicator combination selected by the above-mentioned "factor selection description data of the mth individual" on the target disease, thereby enabling the "fitness value of the mth individual" to better represent the universality of the influencing factors represented by the mth individual on the target disease.
[0165] Based on the relevant content of S324 above, it can be known that after obtaining the classification model corresponding to the mth individual in the flock to be used, the classification model corresponding to the mth individual can be tested using at least one test data and the actual classification label of the at least one test data to obtain the fitness value of the mth individual, so that the "fitness value of the mth individual" can represent the influence of the candidate indicator combination selected by the above "factor selection description data of the mth individual" on the target disease, so that the "fitness value of the mth individual" can better represent the universality of the influencing factors represented by the mth individual on the target disease, so that the influencing factors of the target disease can be determined based on the "fitness value of the mth individual" in the future.
[0166] Based on the relevant contents of S31 to S32 above, it can be known that after obtaining the individual representation characteristics of the mth individual in the flock to be used, the fitness value of the mth individual can be determined with the help of binarization processing and fitness calculation process, so that the "fitness value of the mth individual" can represent the influence of the candidate indicator combination selected by the above "factor selection description data of the mth individual" on the target disease, so that the "fitness value of the mth individual" can better represent the universality of the influence of all candidate indicators represented by the mth individual on the target disease, which is conducive to improving the mining effect of the influencing factors of the target disease.
[0167] In one possible implementation, in order to further improve the mining effect of the influencing factors of the target disease, the embodiment of the present application further provides a possible implementation of S6 above (that is, a possible implementation of the individual representation feature update process), which may specifically include S61-S66:
[0168] S61: Determine each individual in the flock of sheep to be used except the leading sheep as an ordinary sheep.
[0169] In this embodiment of the present application, after obtaining the leader sheep in the flock to be used, each individual in the flock other than the leader sheep can be determined as an ordinary sheep, so that the "ordinary sheep" can represent each individual in the flock other than the leader sheep. Thus, if the flock to be used includes M individuals, then the other M-1 individuals in the flock to be used other than the leader sheep can be determined as ordinary sheep, so that the flock to be used includes 1 leader sheep and M-1 ordinary sheep.
[0170] S62: performing pre-replacement processing on the factor selection description data of each common sheep according to the factor selection description data of the leader sheep in the flock to be used, and obtaining factor selection pre-replacement data corresponding to each common sheep.
[0171] The "pre-replacement data for factor selection corresponding to the lth common sheep" refers to the characteristics of the lth common sheep as it approaches the leader sheep. This data represents the candidate indicator group selected for the target disease when the lth common sheep approaches the leader sheep. Where l is a positive integer, l ≤ L, L is a positive integer, and L represents the number of common sheep.
[0172] In addition, the embodiment of the present application does not limit the determination process of the above-mentioned "factor selection pre-replacement data corresponding to the first ordinary sheep" (such as Figure 4 As shown), for example, it may specifically include S621-S622:
[0173] S621: Compare the factor selection description data of the leading sheep in the flock to be used with the factor selection description data of the lth ordinary sheep to obtain the comparison result corresponding to the lth ordinary sheep, so that the "comparison result corresponding to the lth ordinary sheep" is used to represent the difference between the factor selection description data of the leading sheep and the factor selection description data of the lth ordinary sheep.
[0174] S622: Determine the first replacement position and the second replacement position according to the comparison result corresponding to the lth ordinary sheep.
[0175] As an example, S622 may specifically include S6221-S6222:
[0176] S6221: If the above-mentioned "comparison result corresponding to the first common sheep" indicates that the value at the target position in the above-mentioned "factor selection description data of the leader sheep" is the first value (for example, 1), and the value at the target position in the above-mentioned "factor selection description data of the first common sheep" is the second value (for example, 0), the target position can be determined as the first replacement position (for example, Figure 4The 4th data position shown) is selected so that the "first replacement position" can represent the position in the factor selection description data of the 1st ordinary sheep that needs to be replaced with the first numerical value, so as to achieve the purpose of the 1st ordinary sheep approaching the leading sheep.
[0177] S6222: If the comparison result corresponding to the first common sheep indicates that the value at the target position in the factor selection description data of the leader sheep is the second value (e.g., 0), and the value at the target position in the factor selection description data of the first common sheep is the first value (e.g., 1), the target position can be determined as the second replacement position (e.g., Figure 4 The 8th data position shown) is selected so that the "second replacement position" can represent the position in the factor selection description data of the 1st ordinary sheep that needs to be replaced with the second value, so as to achieve the purpose of the 1st ordinary sheep approaching the leading sheep.
[0178] S623: Replace the value at the first replacement position in the factor selection description data of the lth ordinary sheep with the first value, and replace the value at the second replacement position in the factor selection description data of the lth ordinary sheep with the second value, to obtain the factor selection pre-replacement data corresponding to the lth ordinary sheep.
[0179] Based on the relevant content of S62 above, it can be known that after obtaining the lth ordinary sheep, the factor selection description data of the leader sheep in the flock to be used can be referenced to perform pre-replacement processing on the factor selection description data of the lth ordinary sheep to obtain the factor selection pre-replacement data corresponding to the lth ordinary sheep, so that the factor selection pre-replacement data corresponding to the lth ordinary sheep is closer to the factor selection description data of the leader sheep, so that the "factor selection pre-replacement data corresponding to the lth ordinary sheep" can represent the characteristics of the lth ordinary sheep when it approaches the leader sheep, thereby achieving the purpose of the lth ordinary sheep approaching the leader sheep. Wherein, l is a positive integer, l≤L, L is a positive integer, and L represents the number of ordinary sheep.
[0180] S63: Selecting pre-replacement data based on at least one physical examination sample, an actual classification label of the at least one physical examination sample under the target disease, and factors corresponding to each common sheep, and determining a post-pre-replacement fitness value corresponding to each common sheep.
[0181] The "pre-replacement fitness value corresponding to the lth common sheep" represents the effect of the candidate indicator combination selected by the "pre-replacement data for factor selection corresponding to the lth common sheep" on the target disease. This allows the "pre-replacement fitness value corresponding to the lth common sheep" to represent the universality of the effect of all candidate indicators represented by the lth common sheep after pre-replacement on the target disease. Where l is a positive integer, l ≤ L, L is a positive integer, and L represents the number of common sheep.
[0182] It should be noted that the implementation of S63 is similar to the implementation of S32 above, and for the sake of brevity, it will not be repeated here.
[0183] S64: The difference between the fitness value after pre-replacement corresponding to each ordinary sheep and the fitness value of each ordinary sheep is determined as the pre-replacement benefit value corresponding to each ordinary sheep.
[0184] The pre-replacement benefit value corresponding to the lth ordinary sheep represents the benefit generated when the lth ordinary sheep approaches the leader sheep; and the "pre-replacement benefit value corresponding to the lth ordinary sheep" can be positive or negative. Where l is a positive integer, l ≤ L, L is a positive integer, and L represents the number of ordinary sheep.
[0185] In addition, the embodiment of the present application does not limit the determination process of the above-mentioned "pre-replacement benefit value corresponding to the lth ordinary sheep". For example, it can be implemented specifically using formula (11).
[0186]
[0187] Among them, Gain l represents the pre-replacement benefit value corresponding to the lth ordinary sheep; Indicates the fitness value after pre-replacement corresponding to the lth ordinary sheep; represents the pre-replacement data of the factor selection corresponding to the lth ordinary sheep; D test Indicates "at least one test data" above; Fit(CLS l ,D test ) represents the fitness value of the lth common sheep; CLS l Represents the descriptive data of the factor selection of the lth ordinary sheep.
[0188] S65: Update the individual representation characteristics of the leading sheep according to the pre-replacement benefit values corresponding to all ordinary sheep.
[0189] As an example, S65 may specifically include S651-S653:
[0190] S651: Initialize tamp=1.
[0191] S652: Update the individual representation characteristics of the leading sheep according to the pre-replacement benefit value corresponding to the tamp-th ordinary sheep, and the first replacement position and the second replacement position corresponding to the tamp-th ordinary sheep.
[0192] As an example, S652 may specifically include S6521-S6523:
[0193] S6521: According to the pre-replacement benefit value corresponding to the tamp-th ordinary sheep and the first replacement position corresponding to the tamp-th ordinary sheep, the value at the first replacement position in the individual representation feature of the leading sheep is updated to obtain a first updated value.
[0194] The above “first replacement position corresponding to the tampth common sheep” refers to the position that changes from the second value to the first value (for example, 0→1). For example, when the factor corresponding to the tampth common sheep selects the Wth value in the pre-replacement data, tamp The value at the position is the first value (for example, 1), and the factor selection of the tap-th common sheep describes the Wth value in the data. tamp The value at the position is the second value (for example, 0), and the leader factor selects the Wth value in the description data. tamp When the value at the position is the first value, the above "first replacement position corresponding to the tamp-th ordinary sheep" is W tamp .
[0195] In addition, the embodiment of the present application does not limit the implementation of S6521. For example, it can be implemented using formula (12).
[0196]
[0197] Where, Indicates the first updated value; The leader's individual characteristics are W tamp Upper value; SEP lead =(e lead,1 ,e lead,2 ,…,e lead,N ) represents the individual characteristics of the leader; W tamp Indicates the first replacement position corresponding to the tamp-th ordinary sheep, and W tamp ∈{1,2,…,N}; Gain tamp It represents the pre-replacement benefit value corresponding to the tamp-th ordinary sheep; δ is the incentive adjustment parameter, δ>0, and δ can be set in advance.
[0198] S6522: According to the pre-replacement benefit value corresponding to the tamp-th ordinary sheep and the second replacement position corresponding to the tamp-th ordinary sheep, the value at the second replacement position in the individual representation feature of the leading sheep is updated to obtain a second updated value.
[0199] The above “second replacement position corresponding to the tampth common sheep” refers to the position that changes from the first value to the second value (for example, 1→0). For example, when the factor corresponding to the tampth common sheep selects the Pth value in the pre-replacement data, tampThe value at the position is the second value (for example, 0), and the factor selection of the tamp-th common sheep describes the P-th tamp The value at position P is the first value (for example, 1), and the leading factor selects the factor describing the Pth position in the data. tamp When the value at the position is the second value, the above "second replacement position corresponding to the tamp-th ordinary sheep" is P tamp .
[0200] In addition, the embodiment of the present application does not limit the implementation of S6522. For example, it can be implemented using formula (13).
[0201]
[0202] Where, Indicates the second updated value; The leader's individual characteristics are represented by P tamp Upper value; SEP lead =(e lead,1 ,e lead,2 ,…,e lead,N ) represents the individual characteristics of the leader; P tamp Indicates the second replacement position corresponding to the tamp-th ordinary sheep, and P tamp ∈{1,2,…,N}; Gain tamp It represents the pre-replacement benefit value corresponding to the tamp-th ordinary sheep; δ is the incentive adjustment parameter, δ>0, and δ can be set in advance.
[0203] S6523: Replace the first replacement position W in the leader's individual representation feature with the first updated value tamp The second updated value is used to replace the second replacement position P in the individual representation feature of the leader tamp The above values are used to obtain the updated individual representation features of the leader (for example, when W tamp <P tamp When , the updated individual representation characteristics of the leader can be expressed using formula (14)).
[0204]
[0205] Where SEP′ lead represents the updated individual representation characteristics of the leader; Indicates the first updated value; Indicates the second updated value.
[0206] Based on the relevant content of S652 above, it can be known that after obtaining the tamp, the pre-replacement benefit value corresponding to the tamp-th ordinary sheep, as well as the first replacement position and the second replacement position corresponding to the tamp-th ordinary sheep can be referred to to update the individual representation characteristics of the leading sheep, so as to realize the incentive mechanism generated for the leading sheep when the tamp-th ordinary sheep approaches the leading sheep.
[0207] S653: Determine whether the tamp reaches L. If so, complete the update process of the individual representation features of the leader; if not, execute S654.
[0208] In an embodiment of the present application, if the tamp reaches L, it can be determined that the process of motivating the leader sheep by L ordinary sheep has been completed, so the update process of the individual representation characteristics of the leader sheep can be ended; however, if the tamp does not reach L, it can be determined that the process of motivating the leader sheep by at least one ordinary sheep has not been executed, so S654 and its subsequent steps can be executed.
[0209] S654: Update tamp (as shown in formula (15)), and return to execute S652 and subsequent steps.
[0210] tamp′=tamp+1 (15)
[0211] Where tamp′ represents the updated tamp.
[0212] Based on the relevant content of S65 above, it can be known that after obtaining the pre-replacement profit values corresponding to all ordinary sheep, the pre-replacement profit values corresponding to all ordinary sheep can be used to represent the individual characteristics of the leading sheep, so as to realize the incentive mechanism generated for the leading sheep when all ordinary sheep approach the leading sheep.
[0213] S66: Based on the individual representation characteristics of the leader sheep, update the individual representation characteristics of each individual in the flock to be used.
[0214] In fact, after the leader of the flock of sheep to be used obtains the incentive, each ordinary sheep in the flock of sheep to be used will use the leader as a learning target to update its own representation features. Based on this, the embodiment of the present application provides a possible implementation of S66, which may specifically include S661-S662:
[0215] S661: Based on the individual representation features of the leading sheep, update the individual representation features of each common sheep (as shown in formula (16)).
[0216]
[0217] Where, represents the updated individual representation characteristics of the lth common sheep; Indicates the individual representation characteristics of the lth common sheep before updating; SEP′ lead Represents the individual representation characteristics after the leader is updated; η represents the learning rate, η>0, and η can be pre-set; where l is a positive integer, l≤L, L is a positive integer, and L represents the number of ordinary sheep.
[0218] S662: Determine updated individual representation features of each individual in the flock to be used based on the individual representation features of the leader sheep and the individual representation features of each ordinary sheep.
[0219] In the embodiment of the present application, after completing the update of the individual representation features of the leader sheep and the individual representation features of each common sheep, the individual representation features of the leader sheep (ie, SEP′ lead ), and the individual characteristics of these common sheep (i.e., ) are collected to obtain the updated individual representation features of each individual in the flock to be used (that is, ).
[0220] Based on the relevant content of S66 above, it can be known that after completing the update of the individual representation characteristics of the leader sheep, the individual representation characteristics of each ordinary sheep can be updated with reference to the individual representation characteristics of the leader sheep. In this way, the purpose of updating the individual representation characteristics of each individual in the flock to be used can be achieved, so that S3 and its subsequent steps can be continued based on the updated individual representation characteristics of each individual in the flock to be used to realize the next round of leader sheep screening process. In this way, through multiple rounds of iterative processes, a leader sheep with a relatively high degree of universality can be gradually found, so that the best influencing factors of the target disease can be gradually found.
[0221] Based on the relevant contents of S61 to S66 above, it can be known that after obtaining the fitness value of each individual in the to-be-used flock and the leader in the to-be-used flock, the leader can be incentivized by using each ordinary sheep in the to-be-used flock to obtain the updated individual representation characteristics of the leader; then, referring to the updated individual representation characteristics of the leader, the individual representation characteristics of each ordinary sheep are updated, so that the purpose of updating the individual representation characteristics of each individual in the to-be-used flock can be achieved, so that S3 and its subsequent steps can be continued based on the updated individual representation characteristics of each individual in the to-be-used flock to achieve the next round of leader screening process. It can be seen that the present application can use an incentive mechanism to quickly obtain the individual representation characteristics of each individual in the to-be-used flock, which is conducive to improving the efficiency of mining the influencing factors of the target disease.
[0222] In fact, the "leading sheep factor selection description data" described above can better indicate which candidate indicators can affect the target disease. Based on this, in order to improve the mining effect of influencing factors of the target disease, the embodiment of the present application also provides another possible implementation of S7, which can specifically include: determining at least one influencing factor of the target disease based on the leader factor selection description data in the flock to be used.
[0223] It can be seen that when it is determined that the current iteration has reached the preset stop condition, it can be determined that the universality of the influence of all candidate indicators represented by the leader determined by the current iteration on the target disease is relatively high, so the various candidate indicators selected by the leader's factor selection description data can be directly determined as the various influencing factors of the target disease (that is, the candidate indicators corresponding to each first value of the leader's factor selection description data are all determined as influencing factors of the target disease). Among them, because the above-mentioned "leader's factor selection description data" can more directly indicate which candidate indicators are selected as influencing factors of the target disease, it is possible to quickly determine the various influencing factors of the target disease based on the "leader's factor selection description data" in the future, which is conducive to improving the efficiency of mining influencing factors of the target disease.
[0224] In their research on sheep flocks, the inventors found that some individuals in the flock did not differ much, but some individuals differed greatly. Therefore, those individuals with little difference can be divided into the same population, and those individuals with relatively large differences can be divided into different populations, so that the influencing factors of the target disease can be mined and processed based on each population as a unit. This is conducive to greatly improving the efficiency of mining the influencing factors of the target disease.
[0225] Based on this, in order to further improve the mining effect of the influencing factors of the target disease, the embodiment of the present application also provides another possible implementation of the influencing factor determination method, which may specifically include steps 31 to 38:
[0226] Step 31: Obtain at least one physical examination sample and an actual classification label of the at least one physical examination sample under a target disease.
[0227] It should be noted that for the relevant content of step 31, please refer to S1 above.
[0228] Step 32: Initialize the flock characterization data of the flock to be used.
[0229] The above-mentioned “flock characterization data of the flock to be used” is used to characterize the characteristics of the flock to be used; and the “flock characterization data of the flock to be used” may include the number of individuals in the flock to be used, the number of groups of the flock to be used, and the individual representation characteristics of each individual in the flock to be used.
[0230] The "number of groups of the flock to be used" is used to indicate the number of groups in the flock to be used; and the present embodiment does not limit the "number of groups of the flock to be used." For example, it can be G (e.g., G=25). Where G is a positive integer. It should be noted that the "number of groups of the flock to be used" can be pre-set.
[0231] In addition, for the relevant content of the above “number of individuals in the flock to be used” and the relevant content of the above “individual representation characteristics of each individual in the flock to be used”, please refer to the relevant content of S2 above.
[0232] Step 33: Determine the fitness value of each individual in the sheep flock to be used according to the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the individual representation characteristics of each individual in the sheep flock to be used.
[0233] It should be noted that for the relevant content of step 33, please refer to S3 above.
[0234] Step 34: Based on the individual characteristics of each individual in the flock to be used, cluster all individuals in the flock to be used to obtain G populations, wherein each population includes at least one individual.
[0235] The above-mentioned "clustering processing" is used to classify two individuals with relatively similar individual representation characteristics in the flock to be used into the same class (that is, the same group or the same population), and classify two individuals with dissimilar individual characteristics into different classes (that is, different groups or different populations); and the embodiments of the present application do not limit the implementation method of the above-mentioned "clustering processing", for example, it can be implemented using any existing or future clustering algorithm (for example, K-means++ clustering algorithm, etc.).
[0236] The gth population is used to represent the gth cluster obtained after clustering; and the gth population includes at least one individual. Where g is a positive integer, g≤G, G is a positive integer, and G represents the number of populations (that is, the number of groups or cluster categories).
[0237] It should be noted that the embodiment of the present application does not limit the execution order between step 34 and step 33. For example, step 34 and step 33 may be executed sequentially, step 33 and step 34 may be executed sequentially, or step 34 and step 33 may be executed simultaneously. In addition, for ease of explanation, the following description uses "execution of step 33 and step 34 in sequence" as an example.
[0238] Step 35: According to the fitness value of each individual in the g-th population, determine the leader of the g-th population from all individuals in the g-th population, where g is a positive integer, g≤G, and G is a positive integer.
[0239] The above “leader in the g-th population” is used to refer to the individual with the highest universality among all individuals in the g-th population.
[0240] In addition, the embodiment of the present application does not limit the above-mentioned determination process of the "leader in the g-th population". For example, it can be specifically as follows: first, a maximum analysis is performed on the fitness values of all individuals in the g-th population to obtain the highest fitness value; then, the individual with the highest fitness value in the g-th population is determined as the leader in the g-th population (as shown in formula (17)).
[0241]
[0242]
[0243]
[0244] In the formula, Best (t,g) represents the individual representation characteristics of the leader in the g-th population at the t-th iteration; represents the individual representation characteristics of the vth individual in the gth population at the tth iteration; Denotes the factor selection descriptive data of the vth individual in the gth population at the tth iteration; D test Indicates the above “at least one test data”; represents the fitness value of the vth individual in the gth population at the tth iteration; represents the number of individuals in the g-th population at the t-th iteration; The individual representation feature represents the individual with the highest fitness value in the g-th population.
[0245] Step 36: Determine whether the preset stop condition is met. If so, execute step 38; if not, execute step 37.
[0246] It should be noted that for the relevant content of the above-mentioned “preset stop condition”, please refer to S5 above.
[0247] Step 37: Update the individual representation feature of each individual in the g-th population according to the fitness value of each individual in the g-th population and the leader in the g-th population, and return to execute step 33 and subsequent steps.
[0248] It should be noted that step 37 can be implemented using any of the implementations of S6 above, simply by replacing the "flock to be used" in any of the implementations of S6 above with the "g-th population." Thus, step 37 can be implemented using the process shown in formulas (20)-(24).
[0249]
[0250]
[0251]
[0252]
[0253]
[0254] Where, represents the pre-replacement benefit value corresponding to the jth common sheep in the gth population at the tth iteration; represents the pre-replacement fitness value corresponding to the jth common sheep in the gth population at the tth iteration; represents the factor selection pre-replacement data corresponding to the jth common sheep in the gth population under the tth iteration; D test Indicates the above “at least one test data”; represents the fitness value of the jth common sheep in the gth population at the tth iteration; Represents the factor selection description data of the jth common sheep in the gth population under the tth iteration; j is a positive integer, j≤J, J represents the number of common sheep in the gth population under the tth iteration, J=Class g -1; represents the individual representation characteristics of the leader in the g-th population at the t-th iteration; Indicates the individual representation characteristics of the tamp-th common sheep in the g-th population under the t-th iteration W tamp Upper value; W tamp represents the first replacement position of the tamp-th ordinary sheep in the g-th population under the t-th iteration, and W tamp ∈{1,2,…,J}; The pre-replacement benefit value corresponding to the tamp-th ordinary sheep in the g-th population at the t-th iteration; Indicates the first updated value corresponding to the tamp-th common sheep in the g-th population at the t-th iteration; Indicates the individual characteristics of the tamp-th ordinary sheep in the g-th population under the t-th iteration, P tamp Upper value; P tamp represents the second replacement position of the tamp-th ordinary sheep in the g-th population under the t-th iteration, and Ptamp ∈{1,2,…,J}; Indicates the second updated value of the tamp-th common sheep in the g-th population at the t-th iteration; Best (t,g)′ represents the updated individual representation characteristics of the leader in the g-th population at the t-th iteration; represents the updated individual representation characteristics of the lth common sheep in the gth population at the tth iteration; Represents the individual representation characteristics of the lth common sheep in the gth population before update in the tth iteration.
[0255] It should be noted that formulas (20)-(24) are similar to formulas (11)-(14) and (16) above, and for the sake of brevity, they will not be repeated here.
[0256] Based on the relevant content of step 37 above, it can be seen that in an embodiment of the present application, when it is determined that the current iteration has not yet reached the preset stop condition, it can be determined that the universality of the influence of all candidate indicators represented by the leaders in the G populations determined by the current iteration on the target disease is still relatively low. Therefore, in order to further improve the accuracy of the influencing factors of the target disease, the fitness value of each individual in the g-th population and the leader in the g-th population can be referred to to update the individual representation characteristics of each individual in the g-th population, so that step 33 and its subsequent steps can be continued based on the updated individual representation characteristics of all individuals in the g-th population to realize the next round of leader screening process, so that through multiple rounds of iterative processes, leaders with relatively high universality can be gradually found, so that the best influencing factors of the target disease can be gradually found.
[0257] Step 38: Determine at least one influencing factor of the target disease based on the leaders in the G populations.
[0258] It should be noted that the embodiment of the present application does not limit the implementation method of step 38. For ease of understanding, two examples are used below to illustrate.
[0259] In Example 1, step 38 may specifically include steps 381 and 382:
[0260] Step 381: Compare the fitness values of the leaders in the G populations to obtain a comparison result, so that the “comparison result” can represent the relative sizes of the fitness values of the leaders in the G populations.
[0261] Step 382: When the G populations include the population to be used, and the comparison result indicates that the fitness value of the leader in the population to be used is the highest, determine at least one influencing factor of the target disease based on the individual representation characteristics of the leader in the population to be used.
[0262] In an embodiment of the present application, if the fitness value of the leader in the population to be used is the highest, it can be determined that the universality of the influence of all candidate indicators represented by the leader in the population to be used on the target disease is the highest. Therefore, the individual representation characteristics of the leader in the population to be used can be directly referenced to determine at least one influencing factor of the target disease (for example, α candidate indicators with relatively high correlation representation values in the individual representation characteristics of the leader in the population to be used can all be determined as influencing factors of the target disease), so that these influencing factors of the target disease have a greater correlation with the target disease, thereby achieving the purpose of digging out the influencing factors of the target disease from a large number of physical examination results.
[0263] Based on the relevant contents of steps 381 to 382 above, it can be seen that after determining that the current iteration has reached the preset stop condition, at least one influencing factor of the target disease can be determined based on the individual representation characteristics of the leaders in the G populations to make the influencing factors of the target disease more accurate.
[0264] In Example 2, step 38 may specifically include steps 383 and 384:
[0265] Step 383: Compare the fitness values of the leaders in the G populations to obtain a comparison result, so that the “comparison result” can represent the relative sizes of the fitness values of the leaders in the G populations.
[0266] Step 384: When the G populations include the population to be used, and the comparison result indicates that the fitness value of the leader in the population to be used is the highest, descriptive data is selected based on the factors of the leader in the population to be used to determine at least one influencing factor of the target disease.
[0267] In an embodiment of the present application, if the fitness value of the leader in the population to be used is the highest, it can be determined that the universality of the influence of all candidate indicators represented by the leader in the population to be used on the target disease is the highest, so the various candidate indicators selected by the factor selection description data of the leader can be directly determined as the various influencing factors of the target disease (that is, the candidate indicators corresponding to the first numerical values of the factor selection description data of the leader are all determined as the influencing factors of the target disease). Among them, because the above-mentioned "factor selection description data of the leader in the population to be used" can more directly indicate which candidate indicators are selected as the influencing factors of the target disease, it is possible to quickly determine the various influencing factors of the target disease based on the "factor selection description data of the leader in the population to be used", which is conducive to improving the efficiency of mining the influencing factors of the target disease.
[0268] Based on the relevant contents of steps 383 to 384 above, it can be seen that after determining that the current iteration has reached the preset stop condition, descriptive data can be selected based on the factors of the leaders in the G populations to determine at least one influencing factor of the target disease, which is conducive to improving the determination effect of the influencing factors of the target disease.
[0269] Based on the relevant content of step 38 above, it can be known that when it is determined that the current iteration has reached the preset stop condition, it can be determined that the universality of the influence of all candidate indicators represented by the leaders in the G populations determined by the current iteration on the target disease is relatively high. Therefore, it is possible to directly refer to the individual representation characteristics of the leaders in the G populations (or, factor selection description data) to determine at least one influencing factor of the target disease, so that these influencing factors of the target disease are more correlated with the target disease, thereby achieving the purpose of digging out the influencing factors of the target disease from a large number of physical examination results.
[0270] Based on the relevant contents of steps 31 to 38 above, it can be seen that the sheep to be used can first be divided into populations to obtain multiple populations; then, for each population, the optimal search process and the incentive update process within the group can be iteratively performed to obtain the leaders in each population; finally, based on these leaders, the influencing factors of the target disease are determined. It can be seen that because the iterative search method based on grouping can effectively reduce the search range of each search, the iterative search method based on grouping can effectively improve the optimal search efficiency, thereby making the mining process of the influencing factors of the target disease implemented based on the iterative search method based on grouping converge faster, which is conducive to improving the mining efficiency of the influencing factors of the target disease.
[0271] In their research on sheep, the inventors also discovered that population diversity can be improved by leveraging the migration of individuals between different populations. Based on this, and to further improve the global search capability for influencing factors of a target disease, the present embodiment also provides another possible implementation of the influencing factor determination method. In this implementation, in addition to steps 31-38 above, the method may also include steps 39-41:
[0272] Step 39: Determine the probability of individual migration of the g-th population based on the number of individuals in the g-th population, where g is a positive integer, g≤G, and G is a positive integer.
[0273] The above “number of individuals in the g-th population” is used to indicate how many individuals there are in the g-th population.
[0274] The above-mentioned “probability of individual migration in the g-th population” is used to indicate the probability that an individual in the g-th population leaves the g-th population; and the greater the “probability of individual migration in the g-th population”, the greater the possibility that an individual in the g-th population will leave.
[0275] Moreover, the embodiment of the present application does not limit the determination process of the "probability of individual migration of the g-th population", for example, it can be implemented using formula (25).
[0276]
[0277] Where, represents the probability of individual migration of the g-th population at the t-th iteration; L is a constant that can be preset, and the embodiment of the present application does not limit L. For example, when the above-mentioned "number of individuals in the flock to be used" is 4000, L = 0.001; It represents the number of individuals in the gth population at the tth iteration; Min represents the migration threshold of individuals within the group, and Min can be set in advance.
[0278] Step 40: Determine the individual migration condition of the g-th population based on the individual migration probability of the g-th population, where g is a positive integer, g≤G, and G is a positive integer.
[0279] The above-mentioned “individual migration condition of the g-th population” is used to represent the condition that is reached when an individual in the g-th population leaves the g-th population; and the “individual migration condition of the g-th population” can be preset.
[0280] In addition, the embodiment of the present application does not limit the above-mentioned "g-th population individual migration condition", for example, it can specifically be: the individual migration characteristic data reaches a preset migration threshold. It should be noted that the relevant content of the above-mentioned "migration characteristic data" can be seen in step 41 below.
[0281] In addition, the embodiments of the present application do not limit the determination process of the above-mentioned "individual migration conditions of the g-th population". For example, it can be specifically: searching for the candidate migration conditions corresponding to the above-mentioned "individual migration probability of the g-th population" from the pre-constructed mapping relationship to be used, and determining them as the individual migration conditions of the g-th population.
[0282] The aforementioned "mapping relationship to be used" is used to record the candidate departure conditions corresponding to each candidate departure probability segment. This embodiment of the present application is not limited to this "mapping relationship to be used." For example, it may specifically include: a correspondence between the first candidate departure probability segment and the first candidate departure condition, a correspondence between the second candidate departure probability segment and the second candidate departure condition, ..., and a correspondence between the Uth candidate departure probability segment and the Uth candidate departure condition, where U is a positive integer.
[0283] It can be seen that after obtaining the individual migration probability of the g-th population, the "individual migration probability of the g-th population" can be matched with each candidate migration probability segment in the mapping relationship to be used to obtain a matching result; if the matching result indicates that the "individual migration probability of the g-th population" belongs to the u-th candidate migration probability segment in the mapping relationship to be used, then the u-th candidate migration condition corresponding to the u-th candidate migration probability segment can be directly determined as the individual migration condition of the g-th population. Wherein, u is a positive integer, u∈{1, 2, ..., U}.
[0284] Step 41: After obtaining the migration characterization data of each individual in the g-th population, if the migration characterization data of the target individual in the g-th population meets the individual migration condition of the g-th population, then delete the target individual from the g-th population and add the target individual to the population to be expanded. Where g is a positive integer, g≤G, and G is a positive integer.
[0285] Among them, the "migration representation data of the vth individual in the gth population" is used to indicate the possibility of the vth individual in the gth population leaving the gth population; and the larger the "migration representation data of the vth individual in the gth population", the more likely the vth individual is to leave the gth population.
[0286] In addition, the embodiment of the present application does not limit the determination process of the above-mentioned "migration characterization data of the vth individual in the gth population". For example, it can be specifically: randomly drawing a number between 0 and 1, and determining it as the migration characterization data of the vth individual in the gth population.
[0287] The above-mentioned "target individual" refers to an individual that exists in the g-th population and meets the individual migration conditions of the g-th population; and the embodiment of the present application does not limit the screening process of the "target individual", for example, it can be specifically: after obtaining the migration characterization data of the v-th individual in the g-th population, if the migration characterization data of the v-th individual meets the individual migration conditions of the g-th population (for example, the migration characterization data of the v-th individual reaches a preset migration threshold), then the v-th individual can be determined as the target individual; however, if the migration characterization data of the v-th individual does not meet the individual migration conditions of the g-th population (for example, the migration characterization data of the v-th individual does not reach a preset migration threshold), then the v-th individual can be discarded. Wherein, v is a positive integer, v≤the number of individuals in the g-th population.
[0288] The "population to be expanded" is the population to which the target individual is added; and this "population to be expanded" is a population different from the g-th population. Therefore, the "population to be expanded" can be determined based on the other populations in the G populations except the g-th population.
[0289] In addition, the embodiments of the present application do not limit the determination process of the above-mentioned "population to be expanded". For example, it can be specifically as follows: first, all populations except the g-th population among the G populations are determined as candidate populations; then the population to be expanded with the smallest number of individuals (that is, the smallest population) is screened out from the G-1 candidate populations, so that the target individual can be subsequently migrated from the g-th population to the population to be expanded, thereby realizing the individual migration process between different populations.
[0290] In fact, in order to further improve population diversity, the embodiment of the present application also provides another possible implementation method for determining the above-mentioned "population to be expanded", which can be specifically: first, all the other populations except the g-th population in the G populations are determined as candidate populations; then a candidate population is randomly selected from the G-1 candidate populations and determined as the population to be expanded, so that the target individual can be subsequently migrated from the g-th population to the population to be expanded, thereby realizing the individual migration process between different populations. Among them, because the population to be expanded is randomly selected from the G-1 candidate populations, the randomness of the population to be expanded is relatively large, and thus the randomness of the migration target of the target individual is relatively large, which is conducive to further improving the randomness of the individual migration process between different populations, thereby helping to improve population diversity, and further helping to further improve the global search capability of the influencing factors of the target disease.
[0291] It should be noted that the execution time of step 41 is later than the execution time of step 37. In addition, the embodiment of the present application does not limit the relationship between the execution time of steps 39-40 and the execution time of step 37. For example, they may be earlier than, later than, or equal to each other.
[0292] Based on the relevant contents of steps 31 to 41 above, it can be known that in each iteration process, not only can the individual representation characteristics of each individual in the flock to be used be updated, but the population can also be updated with the help of the migration process between different populations, which is conducive to improving the population diversity, thereby improving the global search capability of the influencing factors of the target disease, and further improving the mining efficiency and stability of the influencing factors of the target disease, so that the influencing factor determination method provided in the embodiment of the present application (such as Figure 5 As shown in the figure, the method can have the characteristics of high search efficiency, strong global optimization ability, and strong algorithm versatility, so that the influencing factor determination method can be used in the future to mine the influencing factors of any disease from a large amount of physical examination data.
[0293] Based on the relevant contents of the above-mentioned method for determining influencing factors, an embodiment of the present application further provides an influencing factor determination device, which is described below with reference to the accompanying drawings for ease of understanding.
[0294] See also Figure 6 , which is a structural diagram of an influencing factor determination device provided by an embodiment of the present application, and as Figure 6 As shown, the influencing factor determination device 600 includes:
[0295] The sample acquisition unit 601 is configured to acquire at least one physical examination sample and an actual classification label of the at least one physical examination sample under a target disease; wherein the physical examination sample includes an indicator value of at least one candidate indicator;
[0296] A flock initialization unit 602 is configured to initialize flock characterization data of the flock to be used, wherein the flock characterization data of the flock to be used includes the number of individuals in the flock to be used and individual representation characteristics of each individual in the flock to be used; the individual representation characteristics are used to indicate the degree of correlation between each candidate indicator and the target disease;
[0297] a fitness determination unit 603, configured to determine a fitness value of each individual in the to-be-used flock of sheep based on the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the individual representative characteristics of each individual in the to-be-used flock of sheep;
[0298] A leader determining unit 604 is configured to determine a leader in the flock of sheep to be used from all individuals in the flock of sheep to be used according to the fitness value of each individual in the flock of sheep to be used;
[0299] A flock updating unit 605 is configured to update the individual representation characteristics of each individual in the to-be-used flock according to the fitness value of each individual in the to-be-used flock and the leader in the to-be-used flock, and return the process to the fitness determination unit to execute the step of determining the fitness value of each individual in the to-be-used flock according to the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the individual representation characteristics of each individual in the to-be-used flock;
[0300] The factor determination unit 606 is configured to determine at least one influencing factor of the target disease based on the individual representative characteristics of the leading sheep in the flock to be used after determining that the preset stop condition is met.
[0301] In a possible implementation, the fitness determination unit 603 includes:
[0302] a binarization processing subunit, configured to perform binarization processing on the individual representation features of each individual in the sheep flock to be used, to obtain factor selection description data of each individual in the sheep flock to be used; wherein the factor selection description data is used to represent the candidate indicator group selected for the target disease;
[0303] The first determination subunit is used to determine the fitness value of each individual in the flock to be used based on the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the factor selection description data of each individual in the flock to be used.
[0304] In one possible embodiment, the number of individuals in the flock to be used is M;
[0305] The binarization processing subunit is specifically used to: perform normalization processing on the individual representation characteristics of the mth individual to obtain the normalized characteristics of the mth individual; determine at least one to-be-selected position that meets a preset selection condition and at least one to-be-discarded position that does not meet the preset selection condition from the normalized characteristics of the mth individual; set each to-be-selected position in the normalized characteristics of the mth individual to a first value, and set each to-be-discarded position in the normalized characteristics of the mth individual to a second value, to obtain the factor selection description data of the mth individual in the flock to be used; m is a positive integer, m≤M, and M is a positive integer.
[0306] In a possible implementation manner, the first determining subunit includes:
[0307] a second determining subunit, configured to determine, based on the at least one physical examination sample and the actual classification label of the at least one physical examination sample under the target disease, at least one training data, the actual classification label of the at least one training data, at least one test data, and the actual classification label of the at least one test data;
[0308] an indicator selection subunit, configured to use the factor selection description data of each individual in the to-be-used flock to perform indicator selection processing on the at least one training data to obtain at least one training sample corresponding to each individual in the to-be-used flock;
[0309] a model training subunit, configured to train a classification model to be processed using at least one training sample corresponding to each individual in the to-be-used flock and an actual classification label of the at least one training data, to obtain a classification model corresponding to each individual in the to-be-used flock;
[0310] The third determining subunit is configured to determine the fitness value of each individual in the to-be-used flock by using the classification model corresponding to each individual in the to-be-used flock, the at least one test data, and the actual classification label of the at least one test data.
[0311] In one possible embodiment, the number of individuals in the flock to be used is M;
[0312] The third determination subunit is specifically used to: determine the mth model classification result of each test data using the classification model corresponding to the mth individual; determine the classification performance characterization data of the classification model corresponding to the mth individual based on the mth model classification result of the at least one test data and the actual classification label of the at least one test data; determine the classification performance characterization data of the classification model corresponding to the mth individual as the fitness value of the mth individual in the flock to be used; wherein m is a positive integer, m≤M, and M is a positive integer.
[0313] In a possible implementation, the flock updating unit 605 includes:
[0314] a fourth determining subunit, configured to determine each individual in the flock of sheep to be used except the leading sheep as an ordinary sheep;
[0315] a pre-replacement processing subunit, configured to perform pre-replacement processing on the factor selection description data of each of the ordinary sheep according to the factor selection description data of the leader sheep, to obtain factor selection pre-replacement data corresponding to each of the ordinary sheep;
[0316] a fifth determining subunit, configured to select pre-replacement data based on the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the factors corresponding to each of the common sheep, and determine a post-pre-replacement fitness value corresponding to each of the common sheep;
[0317] a sixth determining subunit, configured to determine a difference between the fitness value after pre-replacement corresponding to each of the ordinary sheep and the fitness value of each of the ordinary sheep as a pre-replacement benefit value corresponding to each of the ordinary sheep;
[0318] A first updating subunit is configured to update the individual representation characteristics of the leading sheep according to the pre-replacement benefit values corresponding to all ordinary sheep;
[0319] The second updating subunit is configured to update the individual representation characteristics of each individual in the to-be-used flock according to the individual representation characteristics of the leader sheep.
[0320] In a possible implementation, the factor determination unit 606 is specifically configured to: after determining that a preset stop condition is met, determine at least one influencing factor of the target disease based on the factor selection description data of the leader in the flock to be used.
[0321] In a possible implementation, the leader determination unit 604 is specifically configured to determine the individual with the highest fitness value in the flock to be used as the leader in the flock to be used.
[0322] In a possible implementation manner, the flock characterization data of the to-be-used flock further includes the number of groups of the to-be-used flock; the number of groups is G;
[0323] The influencing factor determination device 600 further includes:
[0324] a population division unit, configured to cluster all individuals in the flock of sheep to be used according to the individual representation characteristics of each individual in the flock of sheep to be used, to obtain G populations; wherein each population includes at least one individual;
[0325] The leader determination unit 604 is specifically configured to determine the individual with the highest fitness value in the g-th population as the leader in the g-th population; wherein g is a positive integer, g≤G, and G is a positive integer.
[0326] In one possible embodiment, the flock updating unit 605 is specifically used to: update the individual representation characteristics of each individual in the g-th population according to the fitness value of each individual in the g-th population and the leader in the g-th population, and return to the fitness determination unit to execute the step of determining the fitness value of each individual in the flock to be used according to the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the individual representation characteristics of each individual in the flock to be used.
[0327] In one possible implementation, the factor determination unit 606 is specifically used to: after determining that a preset stop condition is reached, compare the fitness values of the leaders in the G populations to obtain a comparison result; if the comparison result indicates that the fitness value of the leader in the population to be used is the highest, then determine at least one influencing factor of the target disease based on the individual representation characteristics of the leader in the population to be used; wherein the G populations include the population to be used.
[0328] In a possible implementation, the influencing factor determination device 600 further includes:
[0329] An individual migration unit is used to determine the probability of individual migration of the g-th population based on the number of individuals in the g-th population; determine the individual migration conditions of the g-th population based on the individual migration probability of the g-th population; after obtaining the migration characterization data of each individual in the g-th population, if the migration characterization data of the target individual in the g-th population meets the individual migration conditions of the g-th population, delete the target individual from the g-th population and add the target individual to the population to be expanded; wherein, the population to be expanded is determined based on other populations in the G populations except the g-th population.
[0330] In a possible implementation, the influencing factor determination device 600 further includes:
[0331] A first determining unit is configured to determine all populations except the g-th population among the G populations as candidate populations; and to select a population to be expanded with a minimum number of individuals from the G-1 candidate populations;
[0332] or,
[0333] The second determining unit is configured to determine all populations except the g-th population among the G populations as candidate populations; and randomly select a candidate population from the G-1 candidate populations to determine as the population to be expanded.
[0334] Based on the relevant content of the influencing factor determination device 600, it can be known that for the influencing factor determination device 600, first, at least one physical examination sample (for example, the index value of each candidate index) and the actual classification label of each physical examination sample under the target disease are extracted from a large amount of physical examination data, and the individual representation characteristics of each individual in the flock to be used are initialized, so that the individual representation characteristics can represent the degree of correlation between each candidate index and the target disease; secondly, based on these physical examination samples and their actual classification labels, as well as the individual representation characteristics of each individual in the flock to be used, the fitness value of each individual is determined, so that the fitness value can represent the influence of the influencing factor represented by the individual on the target disease. The universality of the disease; then, according to the fitness values of these individuals, the leader in the flock to be used is determined; finally, according to the fitness values of these individuals and the leader, the individual representation characteristics of these individuals are updated, and the above step of "determining the fitness value of each individual according to these physical examination samples and their actual classification labels, and the individual representation characteristics of each individual in the flock to be used" is returned to continue to be executed until it is determined that the preset stopping condition is met, and the influencing factors of the target disease are determined according to the individual representation characteristics of the leader in the flock to be used. In this way, the optimal influencing factors of the target disease can be gradually found through an iterative process, thereby achieving the purpose of mining the influencing factors of a certain disease.
[0335] In addition, the above-mentioned "influencing factor determination device 600" also improves the search efficiency, global optimization capability, and mining versatility of the influencing factor determination device 600 by clustering the entire flock, using binary representation to represent the selected factors, using Softmax to normalize each individual factor, using incentive mechanism to quickly update the optimal individual parameters in the optimization group, using migration mechanism to improve global search capability, maintaining population diversity, and iteratively executing optimization algorithms to gradually find the best relevant factors. Therefore, the influencing factor determination device 600 has the characteristics of high search efficiency, strong global optimization capability, strong versatility, etc., so that the influencing factor determination device 600 has high mining speed and stability.
[0336] In addition, an embodiment of the present application also provides an influencing factor determination device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements any implementation of the influencing factor determination method provided in the embodiment of the present application.
[0337] In addition, an embodiment of the present application also provides a computer-readable storage medium, which stores instructions. When the instructions are executed on a terminal device, the terminal device executes any implementation of the influencing factor determination method provided in the embodiment of the present application.
[0338] In addition, an embodiment of the present application further provides a computer program product. When the computer program product is run on a terminal device, the terminal device executes any implementation of the influencing factor determination method provided in the embodiment of the present application.
[0339] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0340] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0341] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0342] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0343] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for determining influencing factors, characterized in that: The method comprises: Obtaining at least one physical examination sample and an actual classification label of the at least one physical examination sample under a target disease; wherein the physical examination sample includes an indicator value of at least one candidate indicator; Initializing flock characterization data of the flock to be used; wherein the flock characterization data of the flock to be used includes the number of individuals in the flock to be used and individual representation characteristics of each individual in the flock to be used; the individual representation characteristics are used to represent the degree of correlation between each of the candidate indicators and the target disease; Determining a fitness value of each individual in the to-be-used flock of sheep according to the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the individual representative characteristics of each individual in the to-be-used flock of sheep; Determining a leader in the flock of sheep to be used from all individuals in the flock of sheep to be used according to the fitness value of each individual in the flock of sheep to be used; According to the fitness value of each individual in the to-be-used flock of sheep and the leading sheep in the to-be-used flock, the individual representation characteristics of each individual in the to-be-used flock are updated, and the step of determining the fitness value of each individual in the to-be-used flock according to the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the individual representation characteristics of each individual in the to-be-used flock is continued; until it is determined that a preset stopping condition is met, and at least one influencing factor of the target disease is determined according to the individual representation characteristics of the leading sheep in the to-be-used flock; The determining of the fitness value of each individual in the to-be-used flock of sheep based on the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the individual representation characteristics of each individual in the to-be-used flock of sheep comprises: binarizing the individual representation characteristics of each individual in the to-be-used flock of sheep to obtain factor selection description data of each individual in the to-be-used flock of sheep; wherein the factor selection description data is used to represent the candidate indicator group selected for the target disease; determining the fitness value of each individual in the to-be-used flock of sheep based on the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the factor selection description data of each individual in the to-be-used flock of sheep; The method of updating the individual representation characteristics of each individual in the flock of sheep to be used according to the fitness value of each individual in the flock of sheep to be used and the leader sheep in the flock of sheep to be used includes: determining each individual in the flock of sheep to be used except the leader sheep as an ordinary sheep; performing pre-replacement processing on the factor selection description data of each ordinary sheep according to the factor selection description data of the leader sheep to obtain the factor selection pre-replacement data corresponding to each ordinary sheep; determining the pre-replacement fitness value corresponding to each ordinary sheep according to the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the factor selection pre-replacement data corresponding to each ordinary sheep; determining the difference between the pre-replacement fitness value corresponding to each ordinary sheep and the fitness value of each ordinary sheep as the pre-replacement benefit value corresponding to each ordinary sheep; updating the individual representation characteristics of the leader sheep according to the pre-replacement benefit values corresponding to all ordinary sheep; and updating the individual representation characteristics of each individual in the flock of sheep to be used according to the individual representation characteristics of the leader sheep.
2. The method according to claim 1, characterized in that The number of individuals in the flock to be used is M; The process of determining the factor selection description data of the mth individual in the flock to be used includes: Normalizing the individual representation feature of the m-th individual to obtain the normalized feature of the m-th individual; m is a positive integer, m≤M, and M is a positive integer; Determining at least one position to be selected that meets a preset selection condition and at least one position to be discarded that does not meet the preset selection condition from the normalized features of the mth individual; Each position to be selected in the normalized feature of the mth individual is set to a first value, and each position to be discarded in the normalized feature of the mth individual is set to a second value, to obtain factor selection description data of the mth individual in the flock to be used.
3. The method according to claim 1, characterized in that The determining of the fitness value of each individual in the to-be-used flock of sheep based on the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the factor selection description data of each individual in the to-be-used flock of sheep includes: Determining, based on the at least one physical examination sample and the actual classification label of the at least one physical examination sample under the target disease, at least one training data, the actual classification label of the at least one training data, at least one test data, and the actual classification label of the at least one test data; Using the factor selection description data of each individual in the to-be-used flock, perform indicator selection processing on the at least one training data to obtain at least one training sample corresponding to each individual in the to-be-used flock; Using at least one training sample corresponding to each individual in the to-be-used flock and an actual classification label of the at least one training data, the classification model to be processed is trained to obtain a classification model corresponding to each individual in the to-be-used flock; The fitness value of each individual in the to-be-used flock is determined by using the classification model corresponding to each individual in the to-be-used flock, the at least one test data, and the actual classification label of the at least one test data.
4. The method according to claim 3, characterized in that The number of individuals in the flock to be used is M; The process of determining the fitness value of the mth individual in the flock to be used includes: Determine the mth model classification result of each of the test data using the classification model corresponding to the mth individual; wherein m is a positive integer, m≤M, and M is a positive integer; Determining classification performance representation data of the classification model corresponding to the mth individual according to the mth model classification result of the at least one test data and the actual classification label of the at least one test data; The classification performance characterization data of the classification model corresponding to the m-th individual is determined as the fitness value of the m-th individual in the flock to be used.
5. The method according to claim 1, characterized in that The step of determining at least one influencing factor of the target disease based on the individual characteristics of the leading sheep in the flock to be used includes: Descriptive data of factors of the leader in the flock to be used are selected to determine at least one influencing factor of the target disease.
6. The method according to claim 1, characterized in that The step of determining a leader of the flock of sheep to be used from all individuals in the flock of sheep to be used according to the fitness value of each individual in the flock of sheep to be used comprises: The individual with the highest fitness value in the flock to be used is determined as the leader of the flock to be used.
7. The method according to any one of claims 1 to 6, characterized in that The flock characterization data of the flock to be used further includes the number of groups of the flock to be used; the number of groups is G; The method further comprises: Clustering all individuals in the flock of sheep to be used according to the individual representation characteristics of each individual in the flock of sheep to be used to obtain G populations; wherein each population includes at least one individual; The step of determining a leader of the flock of sheep to be used from all individuals in the flock of sheep to be used according to the fitness value of each individual in the flock of sheep to be used comprises: The individual with the highest fitness value in the g-th population is determined as the leader of the g-th population; wherein g is a positive integer, g≤G, and G is a positive integer.
8. The method according to claim 7, characterized in that The updating of the individual representation characteristics of each individual in the to-be-used flock of sheep according to the fitness value of each individual in the to-be-used flock of sheep and the leader of the to-be-used flock of sheep includes: The individual representation feature of each individual in the g-th population is updated according to the fitness value of each individual in the g-th population and the leader of the g-th population.
9. The method according to claim 7, characterized in that The step of determining at least one influencing factor of the target disease based on the individual characteristics of the leading sheep in the flock to be used includes: Compare the fitness values of the leaders in the G populations and obtain the comparison results; If the comparison result indicates that the fitness value of the leader in the population to be used is the highest, then at least one influencing factor of the target disease is determined based on the individual representation characteristics of the leader in the population to be used; wherein the G populations include the population to be used.
10. The method according to claim 7, characterized in that After updating the individual representation feature of each individual in the to-be-used flock of sheep according to the fitness value of each individual in the to-be-used flock of sheep and the leader of the to-be-used flock of sheep, the method further includes: Determining the probability of individual migration of the g-th population according to the number of individuals in the g-th population; Determining individual migration conditions of the g-th population according to the individual migration probability of the g-th population; After obtaining the migration characterization data of each individual in the g-th population, if the migration characterization data of the target individual in the g-th population meets the individual migration condition of the g-th population, the target individual is deleted from the g-th population and added to the population to be expanded; wherein, the population to be expanded is determined based on other populations in the G populations except the g-th population.
11. The method according to claim 10, characterized in that The method further comprises: Determine all populations except the g-th population in the G populations as candidate populations; and select a population to be expanded with the minimum number of individuals from the G-1 candidate populations; or, The method further comprises: All populations except the g-th population in the G populations are determined as candidate populations; and a candidate population is randomly selected from the G-1 candidate populations to be determined as the population to be expanded.
12. A device for determining influencing factors, characterized in that: include: A sample acquisition unit, configured to acquire at least one physical examination sample and an actual classification label of the at least one physical examination sample under a target disease; wherein the physical examination sample includes an indicator value of at least one candidate indicator; A flock initialization unit, configured to initialize flock characterization data of the flock to be used; wherein the flock characterization data of the flock to be used includes the number of individuals in the flock to be used and individual representation characteristics of each individual in the flock to be used; the individual representation characteristics are used to indicate the degree of correlation between each candidate indicator and the target disease; a fitness determination unit, configured to determine a fitness value of each individual in the to-be-used flock of sheep based on the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the individual representative characteristics of each individual in the to-be-used flock of sheep; a leader sheep determining unit, configured to determine a leader sheep in the sheep flock to be used from all individuals in the sheep flock to be used according to the fitness value of each individual in the sheep flock to be used; a flock updating unit, configured to update the individual representation characteristics of each individual in the to-be-used flock according to the fitness value of each individual in the to-be-used flock and the leader in the to-be-used flock, and return the step of determining the fitness value of each individual in the to-be-used flock according to the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the individual representation characteristics of each individual in the to-be-used flock; a factor determination unit, configured to determine at least one influencing factor of the target disease based on individual characteristics of the leading sheep in the flock to be used after determining that a preset stop condition has been reached; The fitness determination unit is specifically configured to: perform binarization processing on the individual representation features of each individual in the to-be-used flock to obtain factor selection description data of each individual in the to-be-used flock; wherein the factor selection description data is used to represent the candidate indicator group selected for the target disease; and determine the fitness value of each individual in the to-be-used flock based on the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the factor selection description data of each individual in the to-be-used flock; The flock updating unit is specifically used to: determine each individual in the flock to be used except the leading sheep as an ordinary sheep; perform pre-replacement processing on the factor selection description data of each ordinary sheep according to the factor selection description data of the leading sheep, and obtain the factor selection pre-replacement data corresponding to each ordinary sheep; determine the pre-replacement fitness value corresponding to each ordinary sheep according to the at least one physical examination sample, the actual classification label of the at least one physical examination sample under the target disease, and the factor selection pre-replacement data corresponding to each ordinary sheep; determine the difference between the pre-replacement fitness value corresponding to each ordinary sheep and the fitness value of each ordinary sheep as the pre-replacement benefit value corresponding to each ordinary sheep; update the individual representation characteristics of the leading sheep according to the pre-replacement benefit values corresponding to all ordinary sheep; and update the individual representation characteristics of each individual in the flock to be used according to the individual representation characteristics of the leading sheep.
13. An influencing factor determination device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for determining an influencing factor according to any one of claims 1 to 11 is implemented.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device executes the influencing factor determination method according to any one of claims 1 to 11.
Citation Information
Patent Citations
A multi-population niche genetic method for feature selection
CN109242100A
Effective hybrid feature selection method based on improved binary krill swarm algorithm and information gain algorithm
CN110837884A