A method and system for identifying sparse positive samples based on hybrid detection and statistical ranking
By using a hybrid detection and statistical ranking method, negative criteria and relaxation parameters are set, a hybrid pool is constructed, and multi-level ranking is performed. This solves the problem of high false negative rate in sparse positive sample detection, achieves a balance between high recall and low false positive rate, and improves the robustness and accuracy of detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2026-02-06
- Publication Date
- 2026-06-02
AI Technical Summary
In the detection of sparse positive samples, existing pooled detection technologies suffer from insufficient robustness in negative determination and overly conservative positive determination strategies, resulting in high false negative rates and making it difficult to achieve a balance between high recall and low false positive rates with a fixed testing budget.
A method based on hybrid detection and statistical ranking is adopted. By setting adjustable negative criteria and relaxation parameters, a hybrid pool is constructed, the frequency of sample occurrence is counted, and multi-level ranking and individual validation are performed to ensure high recall and control false positive rate.
It significantly reduces the risk of missed detections due to false negative results, achieving a balance between high recall and low false positive rate under a fixed testing budget, and improving the robustness and accuracy of the test.
Smart Images

Figure CN122135864A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics and large-scale screening technology, specifically relating to a sparse positive sample identification method and system based on hybrid detection and statistical ranking. By introducing negative determination conditions and relaxation parameters, the robustness and accuracy of positive sample identification are significantly improved under a fixed detection budget. Background Technology
[0002] In scenarios involving large-scale sparse positive sample detection, the number of positive samples Much smaller than the total number of samples ( Traditional single-sample testing methods require individual testing of each sample, resulting in numerous tests and high reagent consumption, making it difficult to meet the demands of high-throughput screening. Pooled testing technology, by combining multiple samples for testing, can significantly reduce the number of tests and lower testing costs, and has become a core solution for detecting sparse positive samples. However, existing pooled testing technologies still have the following key drawbacks, severely limiting their testing performance and application scope.
[0003] First, the robustness of negative determination is insufficient, leading to a high risk of missed detections. Current technologies generally adopt the rule of "a sample is considered negative if it appears in any negative pool," without fully considering the unavoidable noise interference during actual testing (such as fluctuations in the sensitivity of test reagents and sample processing errors). In complex testing scenarios such as low viral load and low concentration of target analytes, false negative results from a single test can easily lead to the incorrect exclusion of real positive samples, resulting in missed detections.
[0004] Secondly, the positive detection strategy is too conservative, resulting in a high false negative rate. Existing methods typically set strict judgment conditions in the decoding stage to pursue low false positive rates and high accuracy. However, in screening scenarios with strict requirements for "false negative rates," this conservative strategy often lacks effective error tolerance and proactive recall mechanisms, leading to the omission of some genuine positive samples and making it difficult to achieve a good balance between sensitivity and specificity.
[0005] Therefore, under the condition of a fixed detection budget, how to design a decoding method that can ensure high recall, effectively control false detections, and be robust to detection noise has become a key technical problem that urgently needs to be solved in this field. Summary of the Invention
[0006] To address the problems in the prior art, this invention proposes a sparse positive sample identification method and system based on hybrid detection and statistical ranking.
[0007] The technical solution adopted in this invention is as follows:
[0008] In a first aspect, the present invention discloses a sparse positive sample identification method based on hybrid detection and statistical ranking, comprising the following steps:
[0009] S1: Collection Each sample to be tested is numbered, and the number of positive samples does not exceed [number missing]. Prepare an equal amount of each sample. share;
[0010] S2: Set a clear negative determination threshold relaxation parameters Set the blending matrix Used to guide the sample mixing process; defines the detection result vector. Used to record the state of the mixing pool (negative or positive); defines the cumulative negative pool count vector. This is used to record the number of times each sample appears in the negative pool; a cumulative positive pool count vector is defined. This is used to record the number of times each sample appears in the positive pool;
[0011] S3: Based on the preset mixing matrix Construct a hybrid pool; detect Each mixing cell is used to output the detection results. ;
[0012] S4: Based on the test results, count the number of times each sample appears in the negative pool, and then determine the threshold. Exclude clearly negative samples to generate a candidate sample set. ; Calculate the number of times each candidate sample appears in the positive pool;
[0013] S5: Sort the samples in a multi-level manner based on the frequency of their appearance in the positive pool, and select the top... 1 sample as the estimated positive set ;
[0014] S6: Estimating the positive set The samples in the dataset are individually tested and verified, and the final identification results are output.
[0015] Secondly, the present invention provides a sparse positive sample identification system for implementing the above-described method, comprising: a sample preprocessing module, a parameter setting module, a mixed detection module, a negative exclusion and frequency statistics module, a sorting and screening module, and a verification output module; the sample preprocessing module is used to number the collected samples and prepare an equal quantity of each sample. The parameter setting module specifies the negative threshold for determination. relaxation parameters Mixed matrix Initialize the detection result vector Cumulative negative pool count vector Cumulative positive pool count vector The mixed detection module is used to construct a mixed pool and perform detection, and output the detection results; the negative exclusion and frequency statistics module is used to count the number of times each sample appears in the negative pool, exclude clearly negative samples according to the judgment threshold, generate a candidate sample set, and count the number of times each candidate sample appears in the positive pool; the sorting and screening module is used to sort the candidate samples in multiple levels and output the estimated positive set according to preset rules; the verification output module is used to perform individual detection and verification on the estimated positive set and output the final positive sample identification results.
[0016] Thirdly, the present invention discloses an electronic device, including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the sparse positive sample identification method based on hybrid detection and statistical ranking.
[0017] Fourthly, the present invention discloses a machine-readable storage medium storing machine-executable instructions, which, when called and executed by a processor, are used to implement the sparse positive sample identification method based on hybrid detection and statistical sorting.
[0018] Compared with the prior art, the beneficial effects of the present invention include:
[0019] This invention overcomes the limitations of the traditional rule that "a sample is considered negative if it appears in any negative pool" by setting adjustable negative determination conditions. It can more robustly handle noise and uncertainty in the detection process and significantly reduce the risk of real positive samples being incorrectly excluded due to false negative results.
[0020] This invention introduces a relaxation parameter in the positive screening stage, establishing a flexible fault-tolerance mechanism. This mechanism allows the method to prioritize high recall for positive samples within a fixed detection budget, while controlling the number of false positives within an expected range, thus achieving a better balance between recall and precision.
[0021] The decoding framework of this invention does not rely on a specific hybrid matrix design and is compatible with various hybrid strategies such as random and deterministic approaches. Its key parameters (such as explicit negative determination threshold and relaxation amount) can be flexibly configured according to different application scenarios and detection performance requirements, demonstrating broad applicability and good system adaptability. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating a sparse positive sample identification method based on hybrid detection and statistical ranking, as described in this invention.
[0023] Figure 2 For embodiments of the present invention with different relaxation parameters Different number of mixed detections The corresponding detection success rate chart.
[0024] Figure 3 This is a schematic diagram of the sparse positive sample identification system of the present invention. Detailed Implementation
[0025] The present invention will be further described and illustrated below with reference to specific embodiments. The technical features of each embodiment of the present invention can be combined accordingly, provided that there is no mutual conflict.
[0026] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.
[0027] like Figure 1 The diagram shown is a flowchart of a sparse positive sample identification method according to a specific embodiment of the present invention. This embodiment starts from... Samples were collected from 100 subjects, of which the number of positive samples was [number missing]. Typically, in order to ensure the number of mixed detections Fewer, number of positive samples It should be smaller than The specific process of this invention embodiment is as follows:
[0028] from Virus samples were collected from all subjects, ensuring that each sample was not contaminated by external factors during the collection process. The viral sample number of each subject was [number missing]. Prepare an equal amount of each sample. Each portion contains the same viral load and is used for subsequent mixing operations.
[0029] Select a hybrid strategy and hybrid matrix based on the scenario requirements. Various forms can be used, such as sparse random matrices and deterministic binary matrices; a clearly defined negative determination threshold should be set. , The value should be no less than ,in Representation matrix The The number of non-zero elements in the column; setting the relaxation parameter. Define the detection result vector. Used to record the state of the mixing pool (negative or positive), initially set to all zeros; define a cumulative negative pool count vector. This is used to record the number of times each sample appears in the negative pool, initially set to all zeros; a cumulative positive pool count vector is defined. This is used to record the number of times each sample appears in the positive pool, initially set to all zeros.
[0030] According to the mixing matrix Build A mixing pool, in which the matrix Each row corresponds to a pooling pool, and each column corresponds to a sample. Indicates the first The sample was added to the first... In each mixing pool; for Each mixing cell is tested, and the test results are output. ,in Indicates the first One of the mixed pools tested positive. Indicates the first The mixed pool was negative.
[0031] For each sample Count the number of times it appears in all negative pools. Initialize the candidate sample set For each sample If its occurrence count in all negative pools If the sample is negative, it is excluded from the candidate sample set. For each candidate sample Count the number of times it appears in all positive pools. This is used as the scoring criterion for suspected positive samples.
[0032] For candidate sample set The samples are sorted in multiple levels: based on the cumulative number of times the sample appears in the positive pool. Sort the samples in descending order, prioritizing those with the highest suspected positive rate; if multiple samples... If they are the same, then it is based on the cumulative number of times the sample appears in the negative pool. Sort the samples in ascending order, prioritizing those that appear less frequently in the negative pool and have a low probability of being suspected negative; if the samples... and If all are the same, sort them in ascending order by sample number to ensure the uniqueness of the sorting result.
[0033] Select the first from the sorted sample sequence 1 sample as the estimated positive set ,in Indicates the size of the candidate sample set; if Then the candidate sample set All samples were included in the estimated positive set.
[0034] For estimating the positive set All samples are tested individually, and the samples with positive test results are output as the final positive samples.
[0035] To verify the feasibility and effectiveness of the technical solution proposed in this invention, simulation experiments were conducted. The experiments simulated... A total of 100 samples to be tested (referred to as sample numbers 1 to 5000) were randomly selected from among them. One sample was positive, and the rest were negative. The viral load in the positive samples followed a specific pattern. The viral load of negative samples was set to 0, ensuring a uniform distribution of viral load on the surface; a clear threshold for determining a negative result was established. Mixed matrix Let be a Bernoulli random matrix, where the elements are... by The probability is 1. The probability is 0. The detection result of each mixing cell can be expressed as:
[0036] (1)
[0037] In formula (1): Indicates the first The test results (negative or positive) of each mixing pool. Representation matrix The The number of non-zero elements in a row. express The viral load of each sample to be tested is represented by the LOD (Limit of Detection) of the test reagent. A negative result is achieved when the viral load in the pool is less than the LOD. This is done to meet the required number of pooled tests. In this case, the results of the mixed cell detection can be utilized. and mixture matrix All positive samples are accurately identified with a probability approaching 1, where This represents hyperparameters.
[0038] Figure 2 These are experimental results from embodiments of the present invention, providing different relaxation parameters. Different number of mixed detections The corresponding detection success rate, where the success rate represents the number of times that all positive samples can be accurately identified when the above process is repeated 500 times. Experimental results show that, under the same number of tests, using the relaxation parameter ( The decoding strategy using relaxation parameters () achieved a significantly higher detection success rate than that without relaxation parameters (). This is a baseline method. Therefore, the positive sample identification method of this invention can achieve higher identification reliability with a fixed detection budget.
[0039] Figure 3 This is a schematic diagram of the sparse positive sample identification system of the present invention. The positive sample identification system of the present invention is used to implement the sample identification method in the foregoing embodiments, and includes:
[0040] Sample preprocessing module: Numbers the collected samples and prepares an equal amount of sample preprocessing. share;
[0041] Parameter setting module: Set the explicit negative determination threshold relaxation parameters Mixed matrix Initialize the detection result vector Cumulative negative pool count vector Cumulative positive pool count vector ;
[0042] Hybrid Detection Module: Constructs a hybrid pool according to a preset hybrid matrix, performs detection, and outputs the detection results;
[0043] Negative exclusion and frequency statistics module: Counts the number of times each sample appears in the negative pool, excludes clearly negative samples based on the judgment threshold, and generates a candidate sample set; counts the number of times each candidate sample appears in the positive pool;
[0044] The sorting and filtering module sorts the samples in multiple levels based on the frequency of their appearance in the positive pool and outputs an estimated positive set according to preset rules.
[0045] Verification output module: Performs separate detection and verification on the estimated positive set, and outputs the final positive sample identification results.
[0046] This invention also provides an electronic device, including a memory and a processor;
[0047] The memory is used to store computer programs;
[0048] The processor is configured to implement the above-described sparse positive sample identification method based on hybrid detection and statistical ranking when executing the computer program.
[0049] This invention also provides a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements the above-described sparse positive sample identification method based on hybrid detection and statistical ranking.
[0050] The computer-readable storage medium can be an internal storage unit of any data processing device described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device of any data processing device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of any data processing device. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.
[0051] Obviously, the embodiments and accompanying drawings described above are merely some examples of this application. Those skilled in the art can apply this application to other similar situations based on these drawings without any creative effort. Furthermore, it is understood that although the work done in this development process may be complex and lengthy, for those skilled in the art, certain design, manufacturing, or production modifications made based on the technical content disclosed in this application are merely conventional technical means and should not be considered as insufficient disclosure of this application. Several modifications and improvements can be made without departing from the concept of this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.
Claims
1. A sparse positive sample identification method based on hybrid detection and statistical ranking, characterized in that: Includes the following steps: S1: Collection Each sample to be tested is numbered, and the number of positive samples does not exceed [number missing]. Prepare an equal amount of each sample share; S2: Set a clear negative determination threshold relaxation parameters Set the blending matrix Used to guide the sample mixing process; defines the detection result vector. Used to record the state of the mixing pool; Define the cumulative negative pool count vector This is used to record the number of times each sample appears in the negative pool; Define the cumulative positive pool count vector This is used to record the number of times each sample appears in the positive pool; S3: Based on the preset mixing matrix Build a hybrid pool; Detection Each mixing cell is used to output the detection results. ; S4: Based on the test results, count the number of times each sample appears in the negative pool, and then determine the threshold. Exclude clearly negative samples to generate a candidate sample set. ; Calculate the number of times each candidate sample appears in the positive pool; S5: Sort the samples in a multi-level manner based on the frequency of their appearance in the positive pool, and select the top... 1 sample as the estimated positive set ; S6: Estimating the positive set The samples in the dataset are individually tested and verified, and the final identification results are output.
2. The sparse positive sample identification method based on hybrid detection and statistical ranking according to claim 1, characterized in that: Step S1 specifically includes: S11: Collection A sample set is defined as a set of samples to be tested. The number of positive samples Less than ; S12: Prepare an equal amount of each sample Each portion contains the same viral load and is used for subsequent mixing operations.
3. The sparse positive sample identification method based on hybrid detection and statistical ranking according to claim 1, characterized in that: In step S2, the mixing matrix Use sparse random matrices or deterministic binary matrices; clearly define the negative determination threshold. , ,in Representation matrix The The number of non-zero elements in the column; relaxation parameter .
4. The sparse positive sample identification method based on hybrid detection and statistical ranking according to claim 1, characterized in that: Step S3 specifically includes: S31: According to the preset mixing matrix Build A mixing pool, in which the matrix Each row corresponds to a pooling pool, and each column corresponds to a sample. Indicates the first The sample was added to the first... In a mixing pool; S32: Yes Each mixing cell is used for detection, and the detection results are output. ,in Indicates the first A pool must be positive, meaning it contains at least one positive sample. Indicates the first Each pool is negative, meaning that theoretically there are no positive samples in the pool.
5. The sparse positive sample identification method based on hybrid detection and statistical ranking according to claim 1, characterized in that: Step S4 specifically includes: S41: For each sample Count the number of times it appears in all negative pools. ,in The value ranges from 0 to the number of times the sample participates in the mixing process, with the sign... This indicates an indicator function that responds to any condition or event. , exist The value is 1 if the condition is met, and 0 otherwise. S42: Initialize the candidate sample set For each sample If its occurrence count in all negative pools If the sample is negative, it is excluded from the candidate sample set, and the candidate sample set is updated. ; S43: For each candidate sample Count the number of times it appears in all positive pools. This was used as the basis for the suspected positive result of the sample.
6. The sparse positive sample identification method based on hybrid detection and statistical ranking according to claim 5, characterized in that: Step S5 specifically includes: S51: For the candidate sample set The samples in the data are sorted in multiple levels: Based on the cumulative number of times the sample appears in the positive pool Sort the samples in descending order and prioritize those with the highest degree of suspected positivity. If multiple samples If they are the same, then it is based on the cumulative number of times the sample appears in the negative pool. Sort the samples in ascending order and prioritize those that appear less frequently in the negative pool and have a low probability of being suspected negative. If the sample and If all are the same, sort them in ascending order by sample number to ensure the uniqueness of the sorting results; S52: Select the first [sample] from the sorted sample sequence 1 sample as the estimated positive set ,in Indicates the size of the candidate sample set; if Then the candidate sample set All samples were included in the estimated positive set.
7. The sparse positive sample identification method based on hybrid detection and statistical ranking according to claim 6, characterized in that: Step S6 specifically includes: S61: Estimating the positive set All samples are tested individually, and the samples with positive test results are output as the final positive samples.
8. A sparse positive sample identification system implementing the method of any one of claims 1-7, characterized in that, include: Sample preprocessing module: Numbers the collected samples and prepares an equal amount of sample preprocessing. share; Parameter setting module: Set the explicit negative determination threshold relaxation parameters Mixed matrix Initialize the detection result vector Cumulative negative pool count vector Cumulative positive pool count vector ; Hybrid Detection Module: Constructs a hybrid pool according to a preset hybrid matrix, performs detection, and outputs the detection results; Negative exclusion and frequency statistics module: Counts the number of times each sample appears in the negative pool, excludes clearly negative samples based on the judgment threshold, and generates a candidate sample set; counts the number of times each candidate sample appears in the positive pool; The sorting and filtering module sorts the samples in multiple levels based on the frequency of their appearance in the positive pool and outputs an estimated positive set according to preset rules. Verification output module: Performs separate detection and verification on the estimated positive set, and outputs the final positive sample identification results.
9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor, the processor executing the machine-executable instructions to implement the sparse positive sample identification method based on hybrid detection and statistical ranking as described in any one of claims 1 to 7.
10. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions, which, when called and executed by a processor, are used to implement the sparse positive sample identification method based on hybrid detection and statistical ranking as described in any one of claims 1 to 7.