Sample retrace method, device, equipment, storage medium and program

Through the training of the flyback model screening and scoring, combined with the accurate scanning of the sample identification platform, the problem of low efficiency in malicious sample discovery in the existing technology is solved, and efficient sample flyback and threat level updates are achieved.

CN114462474BActive Publication Date: 2025-08-19QI AN XIN TECHNOLOGY GROUP INC +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111574824.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-21
Publication Date
2025-08-19
Estimated Expiration
2041-12-21

AI Technical Summary

Technical Problem

In the prior art, antivirus software is difficult to scan and identify massive historical samples in a short time, resulting in low efficiency in discovering malicious samples and inability to find hidden malicious samples in a timely and accurate manner.

Method used

The trained retrieval model is used to score the retrieval samples, filter out high-probability samples with possible changes in the scan results, and accurately scan using the sample identification platform, and combine machine learning models for model performance evaluation and update.

Benefits of technology

Through screening and accurate scanning, the burden on the sample identification platform is reduced, the retrieval efficiency and malicious sample detection rate are improved, and the threat level is timely updated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114462474B_ABST
    Figure CN114462474B_ABST
Patent Text Reader

Abstract

The present invention provides a sample retrace method, apparatus, device, storage medium, and program. The method includes: scanning multiple samples to be retraced based on a trained retrace model to obtain first scores for the multiple samples to be retraced; the first scores represent the probability of a change in the scan results of the samples to be retraced; obtaining at least one first sample based on the first scores of the multiple samples to be retraced; the probability of a change in the scan results corresponding to the first score of the at least one first sample being higher than the probability of a change in the scan results corresponding to the first scores of other samples in the multiple samples to be retraced except the first sample; and scanning the at least one first sample using a sample identification platform to obtain an identification result for the at least one first sample. The method of the embodiment of the present invention has high retrace efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a sample retrace method, apparatus, device, storage medium and program. Background Art

[0002] With the development of network technology, the types of malware are increasing, and some are even able to evade detection by antivirus software. Despite the continuous updates of antivirus software, facing a large number of suspicious historical samples, it is difficult to complete the scanning and identification of all samples in a short period of time, and it is even more difficult to identify all malicious samples in a timely manner.

[0003] Currently, a common approach is to scan recently popular samples to identify hidden malicious samples. However, this method cannot accurately and timely identify all hidden malicious samples from a large number of samples, resulting in low scanning efficiency. Summary of the Invention

[0004] In response to the problems in the prior art, embodiments of the present invention provide a sample retrace method, apparatus, device, storage medium, and program.

[0005] Specifically, the embodiments of the present invention provide the following technical solutions:

[0006] In a first aspect, an embodiment of the present invention provides a sample retrace method, comprising:

[0007] Scanning a plurality of samples to be retraced based on the trained retrace model to obtain first scores for the plurality of samples to be retraced; the first scores represent probabilities that scanning results of the samples to be retraced will change;

[0008] acquiring at least one first sample based on the first scores of the plurality of samples to be retraced; wherein a probability of a change in a scanning result corresponding to the first score of the at least one first sample being higher than a probability of a change in the scanning result corresponding to the first scores of other samples in the plurality of samples to be retraced except the first sample;

[0009] The at least one first sample is scanned using a sample identification platform to obtain an identification result of the at least one first sample.

[0010] Furthermore, after obtaining the identification result of the at least one first sample, the method further includes:

[0011] The performance of the retrace model is determined based on the first score and the identification result of the at least one first sample.

[0012] Furthermore, the method further comprises:

[0013] Obtain multiple samples to be tested;

[0014] Scanning the plurality of samples to be tested using at least one candidate retrace model and the retrace model respectively, and obtaining a second score of each of the samples to be tested corresponding to each of the candidate retrace model and the retrace model;

[0015] Whether to update the retrace model is determined according to each candidate retrace model and the second score of each to-be-tested sample corresponding to the retrace model.

[0016] Furthermore, the performance of each of the candidate retrace models and the retrace model is related to the second score of each of the samples to be tested corresponding to each of the candidate retrace models and the second score of each of the samples to be tested corresponding to the retrace model, and determining whether to update the retrace model according to the second score of each of the candidate retrace models and the samples to be tested corresponding to the retrace model includes:

[0017] determining the performance of each candidate retrace model and each retrace model according to the second scores of each of the to-be-tested samples corresponding to each candidate retrace model and each retrace model;

[0018] If there is a target candidate retrace model among the candidate retrace models, the retrace model is updated to the target candidate retrace model; the target candidate retrace model has the best performance among the at least one candidate retrace model, and the performance of the target candidate retrace model is better than that of the retrace model.

[0019] Furthermore, determining the performance of each candidate retrace model and each retrace model according to the second scores of each of the samples to be tested corresponding to the candidate retrace model and the retrace model, respectively, includes:

[0020] obtaining, based on each candidate retrace model and each second score of each to-be-tested sample corresponding to the retrace model, at least one second sample corresponding to each candidate retrace model and the retrace model; wherein a probability of a change in a scanning result corresponding to the second score of the at least one second sample being higher than a probability of a change in the scanning result corresponding to the second scores of other samples among the plurality of to-be-tested samples except the second sample;

[0021] The performance of each of the candidate retrace models and the retrace model is determined according to the second score of the at least one second sample corresponding to each of the candidate retrace models and the retrace model.

[0022] Furthermore, determining the performance of each of the candidate retrace models and the retrace model includes:

[0023] Scanning each of the candidate retrace models and the at least one second sample corresponding to the retrace model using a scanning and identification platform to obtain identification results of each of the candidate retrace models and the at least one second sample corresponding to the retrace model;

[0024] The performance of each of the candidate retrace models and the retrace model is determined according to the second score and the identification result of the at least one second sample corresponding to each of the candidate retrace models and the retrace model.

[0025] Furthermore, the method further comprises:

[0026] Establishing an initial machine learning model and obtaining training data; the training data includes historical identification records of multiple scanned historical samples and attribute information of the historical samples;

[0027] Based on the training data, the machine learning model is trained to obtain the retrace model.

[0028] In a second aspect, an embodiment of the present invention further provides a sample retrace device, comprising:

[0029] a scanning module, configured to scan a plurality of samples to be retraced based on the trained retrace model, and obtain a first score for the plurality of samples to be retraced; the first score represents a probability that a scanning result of the samples to be retraced changes;

[0030] an acquisition module, configured to acquire at least one first sample based on the first scores of the plurality of samples to be retraced, wherein a probability of a change in a scanning result corresponding to the first score of the at least one first sample being higher than a probability of a change in the scanning result corresponding to the first scores of other samples in the plurality of samples to be retraced except the first sample;

[0031] The scanning module is further configured to scan the at least one first sample using the sample identification platform to obtain an identification result of the at least one first sample.

[0032] In a third aspect, an embodiment of the present invention further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the sample retrace method described in the first aspect are implemented.

[0033] In a fourth aspect, an embodiment of the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the sample retrace method described in the first aspect.

[0034] In a fifth aspect, an embodiment of the present invention further provides a computer program product having executable instructions stored thereon, which, when executed by a processor, enables the processor to implement the steps of the sample retrace method described in the first aspect.

[0035] The sample retracement method, apparatus, equipment, storage medium and program provided in the embodiments of the present invention scan multiple samples to be retraced based on a trained retracement model to obtain a first score for the samples to be retracement; the first score represents the probability that the scanning result of the sample to be retracement changes; and then the first sample with a high probability of changing the scanning result corresponding to the first score is screened out, that is, the retracement sample is screened using the trained model, which greatly reduces the burden on the sample identification platform; finally, the sample identification platform is used to scan the at least one first sample to obtain the identification result of the at least one first sample, so that the sample identification platform can perform sample retrace with a small retracement workload, thereby improving the retrace efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0037] Figure 1 A schematic flow chart of an embodiment of a sample retrace method provided by the present invention;

[0038] Figure 2 A schematic flow chart of another embodiment of the sample retrace method provided by the present invention;

[0039] Figure 3 1 is a structural diagram of an embodiment of a sample retrace device provided by the present invention;

[0040] Figure 4 It is a structural diagram of an embodiment of an electronic device provided by the present invention. DETAILED DESCRIPTION

[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0042] First, the terms and application scenarios involved in the embodiments of the present invention are introduced:

[0043] Retrace sample: software that has been scanned and identified by antivirus software and needs to be scanned and identified again.

[0044] Sample Threat Level: This is a numerical value assigned by the sample identification platform based on the antivirus software's identification results. The higher the level, the greater the threat posed by the sample.

[0045] Due to the advancement of network security technology, malware is becoming increasingly diverse. Furthermore, because antivirus software cannot be updated in a timely manner, viruses can evade detection by camouflaging themselves with malware. Therefore, frequent scanning of historical samples is necessary. When the antivirus module version is updated or the virus database is upgraded, the virus identification rules may change. In this case, the scan results of historical samples may also change. Samples originally identified as common may now be classified as viruses under the new virus identification rules, and vice versa. Therefore, it is necessary to rescan historical samples for which scan results have already been obtained, that is, to obtain the scan results under the new rules.

[0046] Related technologies, limited by the computing power of sample identification platforms, typically cannot perform a comprehensive review of all historical samples in a short period of time. Typically, only historical samples that appeared shortly before a rule change or those that were popular within a short period of time are reviewed. These methods are unable to accurately and timely identify all hidden malicious samples from a massive amount of samples, resulting in low scanning efficiency.

[0047] The sample retrace method of the embodiment of the present invention uses the trained model to screen the retrace samples, which greatly reduces the burden on the sample identification platform, allowing the sample identification platform to perform sample retrace with a small retrace workload, thereby improving the retrace efficiency.

[0048] The following combination Figure 1-Figure 4 The technical solution of the present invention is described in detail with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0049] Figure 1 FIG. 1 is a flow chart of an embodiment of a sample retrace method provided by an embodiment of the present invention. Figure 1 As shown, the method provided by the embodiment of the present invention includes:

[0050] Step 101 : Scan multiple samples to be retraced based on the trained retrace model to obtain first scores of the multiple samples to be retraced; the first scores represent the probability that the scanning results of the samples to be retraced will change.

[0051] Specifically, multiple samples to be retraced are first obtained from the historical sample set and scanned using the trained retrace model. For any sample to be retraced, the scan result generated by the retrace may or may not be the same as the previously detected scan result. Therefore, the first score represents the probability of a change in the scan result. A higher probability indicates a higher probability that the scan result of the sample will be different from the previously detected scan result.

[0052] The retracement model can be used to determine in advance whether the test results of the sample have changed. If the first score is high, it means that there is a high possibility of change. Therefore, it is necessary to retrace the sample identification platform again.

[0053] For example, the scanning results may include the threat level of historical samples. If the threat level obtained by scanning back any historical sample is higher than the threat level detected previously, it means that a new virus has been discovered; if the threat level obtained by scanning back any historical sample is lower than the threat level detected previously, it means that a false alarm has been discovered.

[0054] The first score is positively correlated with the probability of a change in the sample's scan result. The greater the probability, the greater the first score. For example, if the previous scan result of the sample showed a low threat level, if the current scan finds a high threat level, the first score will be high.

[0055] The input parameter of the retrace model is the sample to be retraced, and the output parameter includes: a first score.

[0056] For example, the sample identification platform may identify samples with a hazard level of 10, 20, 30, 50, 60, or 70. Samples with a hazard level of 10 or 20 are identified as common samples and added to the whitelist; samples with a hazard level of 50, 60, or 70 are identified as viruses and added to the blacklist.

[0057] When it is necessary to rescan historical samples that have been identified as viruses, the samples to be rescanned can be historical samples with a hazard level of 50, 60, or 70.

[0058] When rescanning historical samples that were previously identified as common samples, the samples to be rescanned may be historical samples with hazard levels of 10 or 20.

[0059] Step 102: Based on the first scores of the plurality of samples to be retraced, obtain at least one first sample; the probability that the scanning result corresponding to the first score of the at least one first sample changes is higher than the probability that the scanning result corresponding to the first scores of the other samples in the plurality of samples to be retraced changes except the first sample.

[0060] Specifically, some first samples with the highest probability of changes in the scanning results corresponding to the first score are filtered out. For example, each sample to be retraceed is arranged in descending or ascending order according to the probability of changes in the scanning results corresponding to the first score, and at least one first sample with the highest probability of changes in the scanning results corresponding to the first score is taken.

[0061] Through the above two steps, the retrace model can be used to screen out the most suspicious samples that require attention, thereby narrowing the processing scope of the retrace task. Samples with changed scan results can then be rescanned more accurately.

[0062] The first sample may be referred to as a formal retracement sample, that is, a sample that needs to be identified using a sample identification platform.

[0063] Step 103: Scan the at least one first sample using the sample identification platform to obtain an identification result of the at least one first sample.

[0064] Specifically, the sample identification platform is used to scan at least one first sample to obtain an identification result of the at least one first sample.

[0065] The sample identification platform is, for example, anti-virus software, and the identification result may be the current threat level of the first sample.

[0066] In this embodiment, the combination of the retrace model and the sample identification platform significantly improves the detection rate of malicious samples and ensures the timely update of the sample threat level to a certain extent.

[0067] The method of this embodiment scans multiple samples to be retraced based on the trained retracement model to obtain a first score for the samples to be retraced; the first score represents the probability that the scanning result of the sample to be retraced will change; and then the first sample with a high probability of changing the scanning result corresponding to the first score is screened out, that is, the trained model is used to implement the screening of retracement samples, which greatly reduces the burden on the sample identification platform; finally, the sample identification platform is used to scan the at least one first sample to obtain the identification result of the at least one first sample, so that the sample identification platform can perform sample retracement with a small retracement workload, thereby improving the retracement efficiency.

[0068] Optionally, the following operations may be performed after step 103:

[0069] The performance of the retrace model is determined based on the first score and the identification result of the at least one first sample.

[0070] Specifically, the performance of the retrace model can be evaluated based on the identification results. If the first score and the identification result have a high degree of match, the performance of the retrace model is better. If the first score and the identification result have a low degree of match, the performance of the retrace model is worse, so as to monitor the performance changes of the retrace model at any time.

[0071] Optionally, before step 101, the model may be trained to obtain a retrace model, which may be achieved in the following manner:

[0072] Establish an initial machine learning model and obtain training data; the training data includes historical identification records of multiple scanned historical samples and attribute information of the historical samples;

[0073] Based on the training data, the machine learning model is trained to obtain a retracement model.

[0074] Specifically, such as Figure 2 As shown, the historical identification records and attribute information of the scanned historical samples are used to construct a training dataset and complete the training of the machine learning model. The historical samples can be suspicious samples.

[0075] Among them, the attribute information includes, for example, inherent information such as the size and file type of historical samples, and the historical identification records include identification results (such as threat level) and version information of the antivirus software and virus database used during identification.

[0076] Machine learning models may include, for example, neural network models, logistic regression models, support vector machines, random forests, and the like.

[0077] In the above implementation, by training the machine learning model, the screened samples to be retraced are made more accurate, thereby improving the retrace efficiency and accuracy.

[0078] In one embodiment, the method further comprises:

[0079] Obtain multiple samples to be tested;

[0080] Scanning a plurality of samples to be tested using at least one candidate retrace model and a retrace model respectively, and obtaining a second score of each sample to be tested corresponding to each candidate retrace model and retrace model;

[0081] Whether to update the retrace model is determined according to each candidate retrace model and the second score of each to-be-tested sample corresponding to the retrace model.

[0082] Optionally, the performance of each candidate retrace model and the retrace model is related to the second score of each corresponding sample to be tested, such as Figure 2As shown, the performance of each candidate retrace model and the retrace model can be determined respectively according to the second scores of each to-be-tested sample corresponding to each candidate retrace model and the retrace model;

[0083] If there is a target candidate retrace model among the candidate retrace models, the retrace model is updated to the target candidate retrace model; the target candidate retrace model has the best performance among at least one candidate retrace model, and the performance of the target candidate retrace model is better than that of the retrace model.

[0084] Specifically, since the sample distribution may change over time, the retracement model needs to be updated in order to ensure that the performance of the model always maintains a high performance.

[0085] A test sample set is constructed using scanned historical samples. The test sample set includes multiple samples to be tested. The test sample set is used to test some selected new models (such as candidate retrace models) and old models (i.e., retrace models). That is, each candidate retrace model and retrace model scans multiple samples to be tested respectively to obtain a second score for each sample to be tested.

[0086] Based on the second scores of each sample to be tested obtained by scanning, the performance of each candidate retrace model and the retrace model is determined respectively; if there is a candidate retrace model with better performance than the retrace model among the candidate retrace models, the retrace model is updated and the retrace model is updated to the target candidate retrace model with the best performance among the candidate retrace models.

[0087] In the above embodiment, in order to maintain the performance of the retrace model at a high level, the retrace model is updated when there are other models with better performance.

[0088] In one embodiment, the performance of the candidate retrace models and the retrace models can be specifically achieved in the following manner:

[0089] Obtaining, based on each candidate retrace model and each second score of each to-be-tested sample corresponding to the retrace model, at least one second sample corresponding to each candidate retrace model and the retrace model; wherein a probability of a change in a scanning result corresponding to the second score of the at least one second sample being higher than a probability of a change in the scanning result corresponding to the second score of other samples, other than the second sample, among the plurality of to-be-tested samples;

[0090] The performance of each candidate retrace model and the retrace model is determined according to the second score of the at least one second sample corresponding to each candidate retrace model and the retrace model.

[0091] Further, each candidate retrace model and the at least one second sample corresponding to the retrace model are scanned using a scanning and identification platform to obtain identification results of each candidate retrace model and the at least one second sample corresponding to the retrace model;

[0092] The performance of each candidate retrace model and the retrace model is determined according to the second score and the identification result of the at least one second sample corresponding to each candidate retrace model and the retrace model.

[0093] Specifically, based on the second scores of multiple samples to be tested corresponding to each candidate retrace model, second samples with a higher probability of change in the scanning results corresponding to the second scores of each candidate retrace model can be obtained; based on the second scores of multiple samples to be tested corresponding to the retrace model, second samples with a higher probability of change in the scanning results corresponding to the second scores of the retrace model can be obtained. For example, multiple samples to be tested can be arranged in descending or ascending order based on the probability of change in the scanning results corresponding to the second scores. For example, MapReduce and heap sorting technology can be used to screen out some second samples with the highest probability of change in the scanning results corresponding to each candidate retrace model and the scores corresponding to the retrace model. MapReduce and heap sorting technology can quickly screen out samples with a higher probability of change in the scanning results corresponding to the scores.

[0094] Optionally, for the second samples with a higher probability of change in the scanning results corresponding to the second scores of the screened models, data preprocessing such as deduplication and aggregation can be performed before scanning using the sample identification platform.

[0095] The second score and the probability of a corresponding scan result changing may be positively correlated, that is, the greater the second score, the greater the probability.

[0096] For example, the top 20 samples with the highest scores are used as the second samples.

[0097] The second sample may be referred to as a test retracement sample, that is, a sample that needs to be identified by using a sample identification platform during the process of testing the model.

[0098] Furthermore, each candidate retrace model and at least one second sample corresponding to the retrace model are scanned using a scanning and identification platform to obtain identification results of each candidate retrace model and at least one second sample corresponding to the retrace model. Finally, the performance of each candidate retrace model and the retrace model is determined based on the second scores and identification results of the multiple second samples corresponding to each candidate retrace model and the retrace model. For example, a higher degree of match between the second score of the second sample and the identification result indicates better performance.

[0099] Among them, MapReduce is a computing model, framework and platform for parallel processing of big data.

[0100] The candidate retracement model may be obtained by training an established machine learning model with a training sample set.

[0101] In the above implementation, the performance of each model is evaluated through the second scores of multiple test samples and the identification results of the screened second samples with higher second scores, and the implementation scheme is simple.

[0102] The sample retrace device provided by the present invention is described below. The sample retrace device described below and the sample retrace method described above can be referred to in correspondence with each other.

[0103] Figure 3 FIG. 1 is a schematic structural diagram of an embodiment of a sample retracement device provided by the present invention. Figure 3 As shown, the sample retrace device provided in this embodiment includes:

[0104] The scanning module 210 is configured to scan a plurality of samples to be retraced based on the trained retrace model to obtain a first score for the plurality of samples to be retraced; the first score indicates a probability that a scanning result of the samples to be retraced changes;

[0105] An acquisition module 220 is configured to acquire at least one first sample based on the first scores of the plurality of samples to be retraced, wherein a probability of a change in a scanning result corresponding to the first score of the at least one first sample is higher than a probability of a change in the scanning result corresponding to the first scores of other samples in the plurality of samples to be retraced except the first sample;

[0106] The scanning module 210 is further configured to scan the at least one first sample using the sample identification platform to obtain an identification result of the at least one first sample.

[0107] Optionally, it also includes:

[0108] A processing module is configured to determine the performance of the retrace model based on the first score and the identification result of the at least one first sample.

[0109] Optionally, the acquisition module 220 is further configured to acquire a plurality of samples to be tested;

[0110] The scanning module 210 is further configured to scan the plurality of samples to be tested using at least one candidate retrace model and the retrace model, respectively, to obtain a second score of each sample to be tested corresponding to each candidate retrace model and the retrace model;

[0111] The processing module is further configured to determine whether to update the retrace model according to each candidate retrace model and the second score of each sample to be tested corresponding to the retrace model.

[0112] Optionally, the performance of each candidate retrace model and the retrace model is related to the second score of each corresponding sample to be tested, and the processing module is specifically configured to:

[0113] Determining the performance of each candidate retrace model and each retrace model according to the second scores of each to-be-tested sample corresponding to each candidate retrace model and each retrace model;

[0114] If there is a target candidate retrace model among the candidate retrace models, the retrace model is updated to the target candidate retrace model; the target candidate retrace model has the best performance among the at least one candidate retrace model, and the performance of the target candidate retrace model is better than that of the retrace model.

[0115] Optionally, the processing module is specifically configured to:

[0116] obtaining, based on each candidate retrace model and each second score of each to-be-tested sample corresponding to the retrace model, at least one second sample corresponding to each candidate retrace model and the retrace model; wherein a probability of a change in a scanning result corresponding to the second score of the at least one second sample being higher than a probability of a change in the scanning result corresponding to the second scores of other samples among the plurality of to-be-tested samples except the second sample;

[0117] The performance of each of the candidate retrace models and the retrace model is determined according to the second score of the at least one second sample corresponding to each of the candidate retrace models and the retrace model.

[0118] Optionally, the processing module is specifically configured to:

[0119] Scanning each of the candidate retrace models and the at least one second sample corresponding to the retrace model using a scanning and identification platform to obtain identification results of each of the candidate retrace models and the at least one second sample corresponding to the retrace model;

[0120] The performance of each of the candidate retrace models and the retrace model is determined according to the second score and the identification result of the at least one second sample corresponding to each of the candidate retrace models and the retrace model.

[0121] Optionally, the processing module is further configured to:

[0122] Establishing an initial machine learning model and obtaining training data; the training data includes historical identification records of multiple scanned historical samples and attribute information of the historical samples;

[0123] Based on the training data, the machine learning model is trained to obtain the retrace model

[0124] Based on the training data, the machine learning model is trained to obtain the retrace model.

[0125] The device of the embodiment of the present invention is used to execute the method in any of the aforementioned method embodiments. Its implementation principle and technical effects are similar and will not be repeated here.

[0126] Here is an example:

[0127] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call logic instructions in the memory 830 to execute the following method, including: scanning a plurality of samples to be retraced based on a trained retrace model to obtain first scores for the plurality of samples to be retraced; the first scores represent the probability of a change in the scan results of the plurality of samples to be retraced; obtaining at least one first sample based on the first scores of the plurality of samples to be retraced; the probability of a change in the scan results corresponding to the first score of the at least one first sample being higher than the probability of a change in the scan results corresponding to the first scores of other samples in the plurality of samples to be retraced except the first sample; and scanning the at least one first sample using a sample identification platform to obtain an identification result of the at least one first sample.

[0128] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0129] On the other hand, an embodiment of the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the sample retracement method provided in the above embodiments is implemented to execute the following steps: scanning a plurality of samples to be retraced based on a trained retracement model to obtain a first score of the plurality of samples to be retraced; the first score represents a probability that the scanning result of the sample to be retraced changes; obtaining at least one first sample based on the first scores of the plurality of samples to be retraced; the probability that the scanning result corresponding to the first score of the at least one first sample changes is higher than the probability that the scanning result corresponding to the first score of other samples in the plurality of samples to be retraced except the first sample changes; and scanning the at least one first sample using a sample identification platform to obtain an identification result of the at least one first sample.

[0130] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0131] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A sample retrace method, characterized in that: include: Scanning a plurality of samples to be retraced based on the trained retrace model to obtain first scores of the plurality of samples to be retraced; The first score represents a probability that the scan result of the sample to be retraced changes relative to a previously detected scan result of the sample to be retraced; acquiring at least one first sample based on the first scores of the plurality of samples to be retraced; wherein a probability of a change in a scanning result corresponding to the first score of the at least one first sample being higher than a probability of a change in the scanning result corresponding to the first scores of other samples in the plurality of samples to be retraced except the first sample; Scanning the at least one first sample using a sample identification platform to obtain an identification result of the at least one first sample; The method further comprises: Obtain multiple samples to be tested; Scanning the plurality of samples to be tested using at least one candidate retrace model and the retrace model respectively, and obtaining a second score of each of the samples to be tested corresponding to each of the candidate retrace model and the retrace model; obtaining, based on each candidate retrace model and each second score of each to-be-tested sample corresponding to the retrace model, at least one second sample corresponding to each candidate retrace model and the retrace model; wherein a probability of a change in a scanning result corresponding to the second score of the at least one second sample being higher than a probability of a change in the scanning result corresponding to the second scores of other samples among the plurality of to-be-tested samples except the second sample; Scanning each of the candidate retrace models and the at least one second sample corresponding to the retrace model using a scanning and identification platform to obtain identification results of each of the candidate retrace models and the at least one second sample corresponding to the retrace model; determining a degree of match between each candidate retrace model and the second score and the identification result of at least one second sample corresponding to the retrace model, respectively; and determining performance of each candidate retrace model and the retrace model based on the respective degrees of match, wherein the degree of match is positively correlated with the performance; If there is a target candidate retrace model among the candidate retrace models, the retrace model is updated to the target candidate retrace model; the target candidate retrace model has the best performance among the at least one candidate retrace model, and the performance of the target candidate retrace model is better than that of the retrace model.

2. The sample retrace method according to claim 1, wherein: After obtaining the identification result of the at least one first sample, the method further includes: The performance of the retrace model is determined based on the first score and the identification result of the at least one first sample.

3. The sample retrace method according to claim 1 or 2, characterized in that: The method further comprises: Establishing an initial machine learning model and obtaining training data; the training data includes historical identification records of multiple scanned historical samples and attribute information of the historical samples; Based on the training data, the machine learning model is trained to obtain the retrace model.

4. A sample retrace device, characterized in that: include: A scanning module, configured to scan a plurality of samples to be retraced based on a trained retrace model to obtain first scores of the plurality of samples to be retraced; The first score represents a probability that the scan result of the sample to be retraced changes relative to a previously detected scan result of the sample to be retraced; an acquisition module, configured to acquire at least one first sample in the scan result based on the first scores of the plurality of samples to be retraced; a probability that the scanning result corresponding to the first score of the at least one first sample changes is higher than a probability that the scanning results corresponding to the first scores of other samples among the plurality of samples to be retraced, except the first sample, change; The scanning module is further configured to scan the at least one first sample using the sample identification platform to obtain an identification result of the at least one first sample; The device further comprises: The scanning module is further configured to scan a plurality of samples to be tested using at least one candidate retrace model and the retrace model, respectively, to obtain a second score of each sample to be tested corresponding to each candidate retrace model and the retrace model; The processing module is also used to: obtaining, based on each candidate retrace model and each second score of each to-be-tested sample corresponding to the retrace model, at least one second sample corresponding to each candidate retrace model and the retrace model; wherein a probability of a change in a scanning result corresponding to the second score of the at least one second sample being higher than a probability of a change in the scanning result corresponding to the second scores of other samples among the plurality of to-be-tested samples except the second sample; Scanning each of the candidate retrace models and the at least one second sample corresponding to the retrace model using a scanning and identification platform to obtain identification results of each of the candidate retrace models and the at least one second sample corresponding to the retrace model; determining a degree of match between each candidate retrace model and the second score and the identification result of at least one second sample corresponding to the retrace model, respectively; and determining performance of each candidate retrace model and the retrace model based on the respective degrees of match, wherein the degree of match is positively correlated with the performance; If there is a target candidate retrace model among the candidate retrace models, the retrace model is updated to the target candidate retrace model; the target candidate retrace model has the best performance among the at least one candidate retrace model, and the performance of the target candidate retrace model is better than that of the retrace model.

5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the sample retrace method according to any one of claims 1 to 3 are implemented.

6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the sample retrace method according to any one of claims 1 to 3 are implemented.

7. A computer program product having executable instructions stored thereon, characterized in that: When the instruction is executed by a processor, the processor implements the steps of the sample retrace method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Method and device and system for scanning files

    CN102982278A

  • Data processing method, device and equipment and storage medium

    CN110751227A

  • Method and device for determining to-be-compared medical data

    CN112185575A