The invention relates to a text-based
software source code multistage
feature generation method, and belongs to the technical field of
software code reuse detection. According to the method, the Hash values of the file and the effective code are calculated through the Hash
algorithm, the file
fingerprint and the code
fingerprint are formed and used for file or code-level
code reuse detection, the detection efficiency is high, and the detection result is accurate; effective code hash values are calculated through a hash
algorithm, code fingerprints are formed and used for
code reuse detection of the code level, the influence of code modification
modes such as adding or deleting annotations, spaces, table making characters, car returning and line feed codes on code reuse detection in the code reuse process is eliminated, and the
detection rate of code reuse detection is increased; according to the method, the features of the file, the effective code and the code block are calculated through the Hash
algorithm, the
software source code with the complex content is converted into a set of Hash values with the
fixed length, the code reuse detection efficiency is improved, and efficient code reuse detection in a large-scale data scene is achieved.