Repeat code scanning method, device and electronic device
By grouping based on class functions in the application package and comparing codes within the group, the problem of excessive time spent on repeated code scanning in large Internet companies projects is solved, and efficient repeated code scanning and recognition is achieved.
Patent Information
- Application Number
- CN202210426102.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-04-21
AI Technical Summary
In object-oriented programming, the same code is copied to each project due to componentized splitting, resulting in a large number of duplicate codes with similar or identical functions in the App package, increasing user download traffic and development and maintenance costs. The full scanning of the prior art takes too long, especially in large Internet company projects.
By getting all the classes loaded at the target application package when it runs, and classifying classes of the same function into the same group based on the class's functions. Then, class-to-two code comparison is performed in each group, and classes with duplicate codes are selected for marking. This solution combines project engineering characteristics to divide classes based on functional dimensions, narrows the scope of comparison, and only compares within the group.
The speed of repeated code scanning is greatly improved, and the optimization time complexity is about one-by-one of the conventional scheme, which significantly reduces scanning time and improves the hit rate of repeated codes.
Smart Images

Figure CN114706589B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of software development, and in particular to a repeated code scanning method, device and electronic equipment. Background Art
[0002] With the popularity of object-oriented programming, that is, the componentized and modularized development mode of App, the original large functions are divided into multiple small modules, and different modules are managed by different teams. In the development process, the same code is often copied to each project because the basic components are not promptly sunk after the componentization split, resulting in a lot of duplicate code with similar / same functions in the App package. This problem not only causes users to spend more traffic when downloading the App, but also for developers, as the complexity of the system gradually increases, the cost of maintenance is also getting higher and higher. Based on this, code repetitive scanning is a relatively important step in the software development process. The conventional idea of scanning duplicate code in the prior art is to divide the modules into modules according to a certain size, and compare each character content of the two modules word by word to see if they are completely equal. When judging whether the current class has duplicate code, the scanning range is all other classes except the current class itself, and its time complexity is O(n) = (n2-n) / 2, where n is the number of classes. The number of project classes of large Internet companies ranges from tens of thousands to hundreds of thousands, and full scanning takes too long. Summary of the invention
[0003] In view of this, embodiments of the present invention provide a method, device and electronic device for scanning a repetitive code, thereby improving the scanning speed of the repetitive code.
[0004] According to the first aspect, an embodiment of the present invention provides a method for scanning for duplicate code, the method comprising: obtaining all classes loaded when the target application package is running; based on the functions of each class, dividing the classes representing the same functions into the same group; performing code comparison on the classes in each group pair by pair, and screening out classes with duplicate code for duplicate class marking.
[0005] Optionally, based on the functions of each class, classes representing the same functions are divided into the same group, including: grouping each class based on page display function, local display function, network function and tool class.
[0006] Optionally, the codes of the classes in each group are compared pairwise, and the classes with duplicate codes are screened for marking as duplicate classes, including: calculating the string similarity of two classes participating in the comparison in the current group; based on the string similarity of the two classes, determining whether the two classes participating in the comparison contain duplicate codes; if it is determined that the two classes participating in the comparison contain duplicate codes, marking the two classes participating in the comparison in the current group as duplicate classes.
[0007] Optionally, calculating the string similarity of two classes participating in the comparison in the current group includes: respectively calculating the string similarity of class names, field names and method names of the two classes participating in the comparison in the current group.
[0008] Optionally, the determining whether the two classes involved in the comparison contain duplicate code based on the string similarity of the two classes includes: determining whether the string similarity of the class names of the two classes involved in the comparison is greater than a preset threshold for the class names; if it is greater than the preset threshold for the class names, determining whether the string similarity of the field names of the two classes is greater than a preset threshold for the fields; if it is greater than the preset threshold for the fields, determining whether the string similarity of the method names of the two classes is greater than a preset threshold for the methods; if it is greater than the preset threshold for the methods, generating a comprehensive similarity based on the string similarity of the class names, the string similarity of the field names, and the string similarity of the method names; determining whether the comprehensive similarity is greater than a comprehensive preset threshold; if the comprehensive similarity is greater than the comprehensive preset threshold, determining that the two classes involved in the comparison contain duplicate code.
[0009] Optionally, the step of calculating the string similarity of the class names or field names or method names of two classes participating in the comparison in the current group includes: obtaining all the characters in the class names or field names or method names of the two classes participating in the comparison in the current group to generate a word bag; based on whether the characters in the class names or field names or method names of the two classes participating in the comparison in the current group are included in the word bag, generating a one-to-one corresponding comparison vector for the two classes; calculating the vector similarity between the two comparison vectors, and determining the string similarity of the class names or field names or method names of the two classes participating in the comparison in the current group based on the vector similarity.
[0010] Optionally, obtaining all classes loaded when the target application package is running includes: inserting probes at the locations of each class in the target application package; running the target application package, and obtaining all classes loaded when the target application package is running according to a response result of the probe.
[0011] According to the second aspect, an embodiment of the present invention provides a duplicate code scanning device, which includes: a class extraction module, used to obtain all classes loaded when the target application package is running; a class grouping module, used to divide classes representing the same functions into the same group based on the functions of each class; a duplicate scanning module, used to compare the codes of the classes in each group pair by pair, and screen out classes with duplicate codes for duplicate class marking.
[0012] According to the third aspect, an embodiment of the present invention provides an electronic device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the method described in the first aspect or any optional implementation manner of the first aspect by executing the computer instructions.
[0013] According to a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method described in the first aspect or any optional implementation manner of the first aspect.
[0014] The technical solution provided by this application has the following advantages:
[0015] The technical solution provided by the present application obtains all classes loaded when the target application package is running, and then divides the classes representing the same functions into the same group based on the functions of each class. The codes of the classes in each group are compared in pairs, and the classes with duplicate codes are screened for duplicate class marking. In combination with the characteristics of the project engineering, this solution divides the classes into dimensions based on functions, narrows the comparison scope, and only compares within the group. Classes in different groups must not be duplicates of each other. The optimized time complexity is O(n)=(n2-nk) / 2k, where k is the number of groups and n is the number of all classes. The time complexity is about one kth of the conventional solution, which greatly improves the scanning speed of duplicate codes.
[0016] In addition, the duplication judgment criterion was changed from two classes being completely equal to similarity judgment, and similarity judgment was performed by combining multiple names in the class, which improved the duplication hit rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present invention in any way. In the accompanying drawings:
[0018] Figure 1 A schematic diagram showing the steps of a repeated code scanning method in one embodiment of the present invention is shown;
[0019] Figure 2 A schematic flow chart of a repeated code scanning method in one embodiment of the present invention is shown;
[0020] Figure 3 A schematic diagram of a process for calculating string similarity in one embodiment of the present invention is shown;
[0021] Figure 4 A schematic structural diagram of a repetitive code scanning device in one embodiment of the present invention is shown;
[0022] Figure 5 A schematic structural diagram of an electronic device in one embodiment of the present invention is shown. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0024] See also Figure 1 and Figure 2 In one embodiment, a repeated code scanning method specifically includes the following steps:
[0025] Step S101: Acquire all classes loaded when the target application package is running.
[0026] Step S102: Based on the functions of each class, classes representing the same functions are divided into the same group.
[0027] Step S103: Compare the codes of the classes in each group pair by pair, and select the classes with duplicate codes for duplicate class marking.
[0028] Specifically, the process of scanning the repeated code in the target application package is different from the prior art of comparing each class character by character. In this embodiment, all classes loaded when the target application package is running are first obtained. Then, combined with the engineering characteristics of projects such as Android / Java, ios, and python, the class is divided into dimensions according to function based on the Class class in object-oriented programming, and the comparison range is narrowed. Classes with different functions are divided into different groups to ensure that the classes between different groups are absolutely different. Since the classes in the group have similar functions, the third-party components called in the group or the original developed code may have a high degree of similarity. Each group is only compared within the group. Among them, the classes in different groups will not be repeated with each other, which can be achieved by dividing the class path, class naming and other information in the project. Finally, the classes in each group are compared with each other, and the classes with repeated code are screened for repeated class marking. Through the above steps, the time complexity of the optimized repeated code scanning is O(n) = (n2-nk) / 2k, where k is the number of groups and n is the number of all classes. The time complexity is about one k of the conventional solution. For example, when n=100, the number of comparisons in the conventional solution is 4950. When k is 10 groups, the number of comparisons in the optimized solution is only 450, which greatly improves the scanning speed of repeated codes.
[0029] Specifically, in this embodiment, step S101 includes the following steps:
[0030] Step 1: Insert probes at the locations of each class in the target application package.
[0031] Step 2: Run the target application package, and obtain all classes loaded when the target application package is running according to the response result of the probe.
[0032] Specifically, this embodiment obtains all classes loaded when the target application package is running through the instrumentation method, inserts probes into the end of each class in the target application package in advance, and then runs the target application package. The classes loaded when the target application package is running will be captured by the pre-inserted probes after loading, and then all classes loaded when the target application package is running can be accurately and conveniently obtained based on the results of the probe response.
[0033] Specifically, in one embodiment, the above step S102 specifically includes the following steps:
[0034] Step 3: Group each class based on page display function, local display function, network function and tool class. Specifically, in this embodiment, the class functions are divided into classes for providing page drawing functions, classes for providing local display functions for displaying partial screen content, classes for receiving / sending network requests, tool classes for processing static method sets of some common logic, and other functional classes other than the above functions, so that classes belonging to different groups will not be repeated, and there is no need to perform repeated comparisons. Only the classes within the group need to be compared, which improves the efficiency of repeated code scanning.
[0035] Specifically, in one embodiment, the current target application package is a project developed based on Android / Java, so grouping classes based on page display functions includes the following steps:
[0036] 1. Group the class files ending with Activity and Fragment into one group;
[0037] 2. Group the class files ending with Presenter and Contract in the MVP architecture into one group;
[0038] 3. Group the class files ending with ViewModel in the MVVM architecture into a group.
[0039] Grouping classes based on local display function involves the following steps:
[0040] 4. Group the class files ending with View, Dialog, Adapter, etc.
[0041] Grouping classes based on network capabilities involves the following steps:
[0042] 5. Group class files ending with Bean, Request, etc.
[0043] After that, the step of grouping based on the tool class is to group the class files ending with Util, Utils, etc. into one group, and finally group the class files that do not belong to any of the above groups into one group.
[0044] Among them, Activity is a container used to provide page drawing in Android, which can be generally understood as a page in App is an Activity container. Fragment is a display container for a single screen or part of a screen in Android. It is displayed in Activity, and its content range is smaller than Activity, but larger than View. View is a display container for partial screen content in Android. It is displayed in Activity or Fragment, and its content range is smaller than Fragment, and it can only display partial screen content. Dialog is a dialog box in Android, similar to view. The module responsible for sending / receiving network requests in Java or Android includes the entity Req class responsible for sending network content, the entity Resp class responsible for receiving request content, the Request class responsible for initiating network requests, and the Parser class responsible for parsing network content. The tool class is defined in Java or Android, and is a class of static method collections used to process some common logic.
[0045] Based on the above grouping, in Android / Java development projects, it is ensured that classes in different groups will not be repeated, laying the foundation for reducing the number of repeated code scans in subsequent Android / Java projects.
[0046] Specifically, in one embodiment, the above step S103 specifically includes the following steps:
[0047] Step 4: Calculate the string similarity between the two classes involved in the comparison in the current group.
[0048] Step 5: Based on the string similarity of the two classes, determine whether the two classes involved in the comparison contain duplicate codes.
[0049] Step 6: If it is determined that the two classes involved in the comparison contain duplicate codes, the two classes involved in the comparison in the current group are marked as duplicate classes.
[0050] Specifically, in this embodiment, by introducing the concept of similarity, the judgment standard of duplication is changed from two classes being completely equal to a similarity greater than a preset threshold, and similar main contents can be identified as duplication, thereby improving the hit rate of duplication. In this embodiment, the specific measurement method of similarity is obtained by obtaining the similarity of character strings, so as to be more appropriate to the character representation of the code. For example, the "cosine similarity" of two strings of characters is calculated, and the following will use this algorithm as an example to specifically illustrate this solution.
[0051] Specifically, in one embodiment, the above step 4 specifically includes the following steps:
[0052] Step 7: Calculate the string similarity of the class name, field name and method name of the two classes participating in the comparison in the current group respectively.
[0053] Specifically, in Java-based object-oriented programming, all classes usually include one or more fields and one or more methods. Assume that the class name of a class is student, the field field refers to the conditions that need to be defined in the class, such as "age", "gender", "name", etc., and the method method refers to the specific method executed in the class, such as "output age". In addition, there are modifiers and so on. This embodiment achieves the purpose of achieving the most accurate comparison with the minimum number of comparisons by comparing the string similarities of the class names, field names, and method names of the two classes one by one, thereby improving the efficiency of repeated code scanning without losing accuracy. In addition, if the comparison time permits, in order to further improve the comparison accuracy, the comparison steps of parameters and return values can also be added.
[0054] Specifically, in this embodiment, the above step five specifically includes the following steps:
[0055] Step 8: Determine whether the string similarity of the class names of the two classes involved in the comparison is greater than a preset threshold of the class name.
[0056] Step 9: If it is greater than the preset threshold of the class name, determine whether the string similarity of the field names of the two classes is greater than the preset threshold of the field.
[0057] Step 10: If it is greater than the field preset threshold, determine whether the string similarity of the method names of the two classes is greater than the method preset threshold.
[0058] Step 11: If it is greater than the preset threshold of the method, a comprehensive similarity is generated based on the string similarity of the class name, the string similarity of the field name, and the string similarity of the method name.
[0059] Step 12: Determine whether the comprehensive similarity is greater than the comprehensive preset threshold.
[0060] Step 13: If the comprehensive similarity is greater than the comprehensive preset threshold, it is determined that the two classes involved in the comparison contain duplicate codes.
[0061] Specifically, in this embodiment, comparison thresholds are set for class name comparison, field name comparison, and method name comparison, respectively. First, the string similarity of the two class names is compared. If the string similarity of the two class names is greater than the preset threshold of the class name (e.g., 80%), the next step of comparison is entered, otherwise the comparison is directly jumped out, and it is considered that the two classes currently involved in the comparison are not the same class. In the next step, the string similarity of the field name is compared with the preset threshold of the field. Similarly, if the string similarity of the field names of the two classes is greater than the preset threshold, the next step of comparison is entered again, otherwise the comparison is directly jumped out, and it is considered that the two classes currently involved in the comparison are not the same class. After that, the string similarity of the method name is compared with the preset threshold of the method. When the string similarity of the method name is greater than the preset threshold of the method, the final comprehensive preset threshold comparison is performed. First, based on the string similarity of the class name, the string similarity of the field name and the string similarity of the method name, a comprehensive similarity is generated. For example, in this embodiment, a weighted summation method can be used to calculate the comprehensive similarity. The weights of the class name similarity, the field name similarity and the method name similarity are respectively assigned to 0.5, 0.3 and 0.2. Then, the comprehensive similarity = 0.5* string similarity of the class name + 0.3* string similarity of the field name + 0.2* string similarity of the method name. Then, the comprehensive similarity is compared with the comprehensive preset threshold. For example, the comprehensive preset threshold is 50%. If the comprehensive similarity is greater than 50%, it is considered that the two classes contain a large amount of duplicate code, so that the two classes are determined to be duplicate classes, and the two classes participating in the comparison in the current group are marked as duplicate classes. In this embodiment, through multiple threshold comparison, if the comparison process of the two classes does not meet any of the above conditions, they are not considered to be duplicate classes, thereby avoiding misjudgment; once the two classes participating in the comparison meet all the above conditions, the probability that the classes participating in the comparison are duplicate classes increases significantly, thereby further improving the accuracy of code duplication scanning.
[0062] Specifically, in one embodiment, the specific steps of calculating the string similarity of the class names of two classes involved in the comparison, or calculating the string similarity of the field names, or calculating the string similarity of the method names are as follows:
[0063] Step 14: Get all the characters in the class names, field names or method names of the two classes involved in the comparison in the current group, and generate a word bag.
[0064] Step 15: Based on whether the characters in the class names, field names, or method names of the two classes participating in the comparison in the current group are contained in the word bag, a one-to-one comparison vector is generated for the two classes.
[0065] Step 16: Calculate the vector similarity between the two comparison vectors, and determine the string similarity of the class names or field names or method names of the two classes participating in the comparison in the current group based on the vector similarity.
[0066] Specifically, if Figure 3 As shown, in this embodiment, firstly, all characters in the class names, field names or method names of the two classes involved in the comparison in the current group are obtained to generate word bags. Taking the field name as an example, assuming that the two field names are:
[0067] String1 = "Route"
[0068] String2 = "RouterInfo"
[0069] Extract all characters to generate word bag as {'R', 'o', 'u', 't', 'e', 'r', 'I', 'n', 'f'}.
[0070] Afterwards, by determining whether the characters in the class names, field names or method names of the two classes participating in the comparison in the current group are included in the word bag, a one-to-one comparison vector for the two classes is generated. For example: if 0, 1 is used to determine whether an element is in the word bag (the present invention is only used as an example and is not limited to this), strings 1 and 2 can be converted into: StringA = [111110000], StringB = [111111111]. Finally, by calculating the similarity of the two vectors, the string similarity of the two names can be obtained through the vector similarity. For example, the value of the similarity of the above two vectors calculated based on the cosine similarity is 0.555. Through the processing steps of this embodiment, the abstract string can be converted into a processable digital vector, thereby calculating the accurate string similarity.
[0071] Through the above steps, the technical solution provided by this application obtains all classes loaded when the target application package is running, and then divides the classes representing the same functions into the same group based on the functions of each class. The codes of the classes in each group are compared two by two, and the classes with duplicate codes are screened for duplicate class marking. This solution combines the characteristics of the project engineering, divides the classes into dimensions based on functions, narrows the comparison scope, and only compares within the group. Classes in different groups must not be repeated with each other. The optimized time complexity is O(n) = (n2-nk) / 2k, where k is the number of groups and n is the number of all classes. The time complexity is about one kth of the conventional solution, which greatly improves the scanning speed of duplicate codes.
[0072] In addition, the duplication judgment criterion was changed from two classes being completely equal to similarity judgment, and similarity judgment was performed by combining multiple names in the class, which improved the duplication hit rate.
[0073] like Figure 4 As shown, this embodiment also provides a repeated code scanning device, the device comprising:
[0074] The class extraction module 101 is used to obtain all classes loaded when the target application package is running. For details, please refer to the relevant description of step S101 in the above method embodiment, which will not be repeated here.
[0075] The class grouping module 102 is used to group classes representing the same functions into the same group based on the functions of each class. For details, please refer to the relevant description of step S102 in the above method embodiment, which will not be repeated here.
[0076] The repeat scanning module 103 is used to compare the codes of the classes in each group, screen the classes with repeated codes and mark them as repeated classes. For details, please refer to the relevant description of step S103 in the above method embodiment, which will not be repeated here.
[0077] The repeated code scanning device provided in the embodiment of the present invention is used to execute the repeated code scanning method provided in the above embodiment. Its implementation method and principle are the same. For details, please refer to the relevant description of the above method embodiment and will not be repeated here.
[0078] Through the collaborative cooperation of the above-mentioned components, the technical solution provided by this application obtains all classes loaded when the target application package is running, and then divides the classes representing the same functions into the same group based on the functions of each class. The codes of the classes in each group are compared two by two, and the classes with duplicate codes are screened for duplicate class marking. This solution combines the characteristics of the project engineering, divides the classes into dimensions based on functions, narrows the comparison scope, and only compares within the group. Classes in different groups must not be repeated with each other. The optimized time complexity is O(n) = (n2-nk) / 2k, where k is the number of groups and n is the number of all classes. The time complexity is about one kth of the conventional solution, which greatly improves the scanning speed of duplicate codes.
[0079] In addition, the duplication judgment criterion was changed from two classes being completely equal to similarity judgment, and similarity judgment was performed by combining multiple names in the class, which improved the duplication hit rate.
[0080] Figure 5 An electronic device according to an embodiment of the present invention is shown, the device includes a processor 901 and a memory 902, which can be connected via a bus or other means. Figure 5 The example of connecting through bus is taken in the following.
[0081] The processor 901 may be a central processing unit (CPU). The processor 901 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.
[0082] The memory 902 is a non-transitory computer-readable storage medium that can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as program instructions / modules corresponding to the methods in the above method embodiments. The processor 901 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory 902, that is, implementing the methods in the above method embodiments.
[0083] The memory 902 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required by at least one function; the data storage area may store data created by the processor 901, etc. In addition, the memory 902 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 902 may optionally include a memory remotely arranged relative to the processor 901, and these remote memories may be connected to the processor 901 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0084] One or more modules are stored in the memory 902 , and when executed by the processor 901 , the method in the above method embodiment is executed.
[0085] The specific details of the above electronic device can be understood by referring to the corresponding descriptions and effects in the above method embodiments, and will not be repeated here.
[0086] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the implemented program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above-mentioned types of memory.
[0087] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A repeated code scanning method, It is characterized in that The method comprises: Get all classes loaded when the target application package is running; Based on the functions of each class, the classes representing the same functions are grouped into the same group; Compare the codes of the classes in each group separately, and select the classes with duplicate codes and mark them as duplicate classes; The code comparison is performed on each class in each group, and classes with duplicate codes are screened for duplicate class marking, including: respectively calculating the string similarity of the class name, field name and method name of the two classes participating in the comparison in the current group; generating the comprehensive similarity of the two classes based on the string similarity of the class name, the string similarity of the field name and the string similarity of the method name of the two classes; based on the comprehensive similarity of the two classes, determining whether the two classes participating in the comparison contain duplicate codes; if it is determined that the two classes participating in the comparison contain duplicate codes, marking the two classes participating in the comparison in the current group as duplicate classes; The step of calculating the string similarity of the class names or field names or method names of two classes participating in the comparison in the current group includes: obtaining all characters in the class names or field names or method names of the two classes participating in the comparison in the current group to generate a word bag; based on whether the characters in the class names or field names or method names of the two classes participating in the comparison in the current group are included in the word bag, generating a one-to-one corresponding comparison vector for the two classes; calculating the vector similarity between the two comparison vectors, and determining the string similarity of the class names or field names or method names of the two classes participating in the comparison in the current group based on the vector similarity.
2. The method according to claim 1, It is characterized in that Based on the functions of each class, the classes representing the same functions are divided into the same group, including: Each category is grouped based on page display function, local display function, network function and tool category.
3. The method according to claim 1, It is characterized in that The step of determining whether the two classes involved in the comparison contain duplicate codes based on the comprehensive similarity of the two classes includes: Determine whether the string similarity of the class names of the two classes involved in the comparison is greater than a preset threshold of the class name; If it is greater than the class name preset threshold, then determine whether the string similarity of the field names of the two classes is greater than the field preset threshold; If it is greater than the preset threshold of the field, then determine whether the string similarity of the method names of the two classes is greater than the preset threshold of the method; If it is greater than a preset threshold of the method, a comprehensive similarity is generated based on the string similarity of the class name, the string similarity of the field name, and the string similarity of the method name; Determine whether the comprehensive similarity is greater than a comprehensive preset threshold; If the comprehensive similarity is greater than the comprehensive preset threshold, it is determined that the two classes involved in the comparison contain duplicate codes.
4. The method according to claim 1, It is characterized in that The method of obtaining all classes loaded by the target application package at runtime includes: Inserting probes at the locations of each class in the target application package; The target application package is run, and all classes loaded when the target application package is run are obtained according to the response result of the probe.
5. A repeat code scanning device, It is characterized in that The device comprises: The class extraction module is used to obtain all classes loaded by the target application package at runtime; A class grouping module is used to group classes representing the same functions into the same group based on the functions of each class; A repeated scanning module is used to compare the codes of the classes in each group in pairs, and screen the classes with repeated codes for repeated class marking; the code comparison of the classes in each group in pairs, and screening the classes with repeated codes for repeated class marking, includes: calculating the string similarity of the class names, field names and method names of the two classes participating in the comparison in the current group; generating the comprehensive similarity of the two classes based on the string similarity of the class names, the string similarity of the field names and the string similarity of the method names of the two classes; based on the comprehensive similarity of the two classes, determining whether the two classes participating in the comparison contain repeated codes; if it is determined that the two classes participating in the comparison contain repeated codes, marking the two classes participating in the comparison in the current group as repeated classes; The step of calculating the string similarity of the class names or field names or method names of two classes participating in the comparison in the current group includes: obtaining all characters in the class names or field names or method names of the two classes participating in the comparison in the current group to generate a word bag; based on whether the characters in the class names or field names or method names of the two classes participating in the comparison in the current group are included in the word bag, generating a one-to-one corresponding comparison vector for the two classes; calculating the vector similarity between the two comparison vectors, and determining the string similarity of the class names or field names or method names of the two classes participating in the comparison in the current group based on the vector similarity.
6. An electronic device, It is characterized in that include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 4 by executing the computer instructions.
7. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Method and apparatus for detecting pirated application program, computer device and storage medium
CN109446753A
Program similarity detection method and device
CN110297750A