Similarity detection method and device, electronic equipment and storage medium
By vectorizing the code files to be tested and matching similarity in the preset vector database, the problems of low stability and code security of open source code derivative software are solved, and high-precision homologous code detection is achieved.
Patent Information
- Application Number
- CN202311603480.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-05-30
AI Technical Summary
The existing open source code derived from low software stability, code base redundancy and vulnerability transmission have affected code management and code security.
By vectorizing the code file to be tested, its complex relationships and representations are extracted, and a vector with a similar value greater than the preset threshold is determined in the preset vector database, and the similarity between the code file and the vector database is calculated to improve the accuracy of homologous code detection.
This method can effectively improve the accuracy of homologous code detection, identify homologous code directories, code snippets and syntax trees in the code base, and improve the accuracy of code detection.
Smart Images

Figure CN120066509A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and more specifically, to a method, apparatus, electronic device, and storage medium for detecting similarity in the field of computers. Background Art
[0002] Currently, the open-source model has become an important software development model. Open-source technologies are used in almost all enterprise software systems. More and more enterprises and developers improve or develop and expand based on open-source code. The introduction of open-source code has also improved the efficiency and reduced the cost for enterprises.
[0003] However, problems such as low stability of software derived from existing open-source code, redundancy of code libraries, and vulnerability transmission have become increasingly prominent, affecting code management and code security. Summary of the Invention
[0004] The present application provides a method, apparatus, electronic device, and storage medium for detecting similarity, which can effectively improve the accuracy of detecting homologous code.
[0005] In a first aspect, a method for detecting similarity is provided, characterized in that the method includes: obtaining at least one code file to be tested; performing vectorization processing on the at least one code file to be tested to obtain at least one first target vector; determining at least one second target vector in a preset vector database whose similarity value with the at least one first target vector is greater than a preset threshold; and determining the similarity between the code file to be tested and the preset vector database based on the at least one first target vector and the at least one second target vector.
[0006] In the above technical solution, first, at least one code file to be tested is obtained, then vectorization processing is performed on the at least one code file to be tested to obtain at least one first target vector, and then at least one second target vector whose similarity value with the at least one first target vector is greater than a preset threshold is determined in the preset vector database. Furthermore, based on the at least one first target vector and the at least one second target vector, the similarity between the code file to be tested and the preset vector database is determined. By vectorizing at least one code file to be tested, complex relationships and representations between codes can be extracted, and at least one second target vector whose similarity value with the at least one first target vector is greater than a preset threshold can be quickly determined in the preset vector database. Finally, the similarity between the code file to be tested and the preset vector database can be more accurately determined through the similar vectors of the at least one first target vector. The above method uses vector similarity matching technology and can effectively improve the accuracy of detecting homologous code.
[0007] In combination with the first aspect, in some implementations of the first aspect, when there are multiple code files to be tested, the first target vector includes a first target directory vector, a first target code vector, and a first target syntax tree vector; the first target directory vector is obtained by vectorizing the repository directory information corresponding to the multiple code files to be tested; the first target code vector is obtained by vectorizing the code snippets of the at least one code file to be tested, and the first target syntax tree vector is obtained by vectorizing the syntax trees of the at least one code file to be tested; when there is one code file to be tested, the first target vector includes the first target code vector and the first target syntax tree vector.
[0008] In combination with the first aspect and the above implementation, in some implementations of the first aspect, the method further includes: obtaining the at least one code file to be tested; extracting syntax trees from the at least one code file to be tested to obtain the syntax trees of the at least one code file to be tested; and performing a segmentation process on the at least one code file to be tested based on the syntax trees of the at least one code file to be tested and a preset length to obtain code snippets of the at least one code file to be tested.
[0009] In combination with the first aspect and the above implementation, in some implementations of the first aspect, determining at least one second target vector in the preset vector database whose similarity value with the at least one first target vector is greater than a preset threshold includes: when there are multiple code files to be tested, respectively determining in the preset vector database a second target directory vector, a second target code vector, and a second target syntax tree vector whose similarity values with the first target directory vector, the first target code vector, and the first target syntax tree vector are greater than the preset threshold; when there is one code file to be tested, respectively determining in the preset vector database a second target code vector and a second target syntax tree vector whose similarity values with the first target code vector and the first target syntax tree vector are greater than the preset threshold.
[0010] Combined with the first aspect and the above implementation manners, in some implementation manners of the first aspect, determining the similarity between the code file to be tested and the preset vector database based on the at least one first target vector and the at least one second target vector includes: obtaining the length of each first target vector and the length of each second target vector; determining a target similarity vector between each first target vector and each second target vector, and determining a target similarity value corresponding to the target similarity vector; wherein, the target similarity vector is the vector in the second target vectors with the highest similarity value to the first target vector; obtaining a similarity score of each first target vector and each second target vector based on the length of each first target vector, the length of each second target vector, and the target similarity value corresponding to the target similarity vector; and obtaining the similarity between the code file to be tested and the preset vector database based on at least one of the similarity scores.
[0011] In the above technical solution, when there are multiple code files to be detected, that is, when detecting a code repository, the directory structure of the entire code repository is extracted, and similar directories are detected in the preset vector library, so that homologous code directories can be identified, and the similarity between the code file to be tested and the preset vector library in terms of directory information can be determined; code snippets are extracted, and after vectorizing the code snippets, similar code snippets are detected in the preset vector library, and the similarity between the code file to be tested and the preset vector library in terms of code snippets can be determined; the syntax tree of the code file to be tested can also be extracted, and after vectorizing the syntax tree, similar syntax trees are detected in the preset vector library, and the similarity between the code file to be tested and the preset vector library in terms of code structure, that is, the syntax tree, can be determined. By extracting multiple features of the code file to be tested and performing multi-dimensional similarity calculations on the code file to be tested and the preset vector database, the accuracy of similarity detection can be improved.
[0012] Combined with the first aspect and the above implementation manners, in some implementation manners of the first aspect, obtaining the similarity score of each first target vector and each second target vector based on the length of each first target vector, the length of each second target vector, and the target similarity value corresponding to the target similarity vector includes: when there are multiple code files to be tested, determining a first similarity score between the first target directory vector and the second target directory vector, a second similarity score between the first target code vector and the second target code vector, and a third similarity score between the first target syntax tree vector and the second target syntax tree vector; when there is one code file to be tested, determining the second similarity score between the first target code vector and the second target code vector and the third similarity score between the first target syntax tree vector and the second target syntax tree vector.
[0013] Combined with the first aspect and the above implementation manners, in some implementation manners of the first aspect, obtaining the similarity between the code file to be tested and the preset vector database based on at least one of the similarity scores includes: when there are multiple code files to be tested, fusing the first similarity score, the second similarity score, and the third similarity score to obtain the similarity between the code files to be tested and the preset vector database; when there is one code file to be tested, fusing the second similarity score and the third similarity score to obtain the similarity between the code file to be tested and the preset vector database.
[0014] In the above technical solution, after obtaining the similarities between the code file to be tested and the preset vector library in multiple dimensions, these similarities can be fused, which can not only take into account the similarity of code segments, but also take into account the similarity of directory information and syntax trees, and can further improve the accuracy of code detection.
[0015] In a second aspect, a similarity detection device is provided. The device includes: an acquisition module, configured to acquire at least one code file to be tested; a processing module, configured to perform vectorization processing on the at least one code file to be tested to obtain at least one first target vector; a first determination module, configured to determine at least one second target vector in a preset vector database, where a similarity value between the at least one second target vector and the at least one first target vector is greater than a preset threshold; and a second determination module, configured to determine the similarity between the code file to be tested and the preset vector database based on the at least one first target vector and the at least one second target vector.
[0016] Combined with the second aspect, in some implementation manners of the second aspect, when there are multiple code files to be tested, the first target vector includes a first target directory vector, a first target code vector, and a first target syntax tree vector; the first target directory vector is obtained by vectorizing the repository directory information corresponding to the multiple code files to be tested; the first target code vector is obtained by vectorizing the code segments of the at least one code file to be tested, and the first target syntax tree vector is obtained by vectorizing the syntax trees of the at least one code file to be tested; when there is one code file to be tested, the first target vector includes the first target code vector and the first target syntax tree vector.
[0017] Combined with the second aspect and the above implementation manners, in some implementation manners of the second aspect, the device further includes a segmentation processing module, and the segmentation processing module is specifically configured to: acquire the at least one code file to be tested; extract syntax trees from the at least one code file to be tested to obtain the syntax trees of the at least one code file to be tested; and perform segmentation processing on the at least one code file to be tested based on the syntax trees of the at least one code file to be tested and a preset length to obtain code segments of the at least one code file to be tested.
[0018] Combined with the second aspect and the above implementation manners, in some implementation manners of the second aspect, the first determination module is specifically configured to: when there are multiple to-be-tested code files, respectively determine, in the preset vector database, a second target directory vector, a second target code vector, and a second target syntax tree vector whose similarity values with the first target directory vector, the first target code vector, and the first target syntax tree vector are greater than the preset threshold; when there is one to-be-tested code file, respectively determine, in the preset vector database, a second target code vector and a second target syntax tree vector whose similarity values with the first target code vector and the first target syntax tree vector are greater than the preset threshold.
[0019] Combined with the second aspect and the above implementation manners, in some implementation manners of the second aspect, the second determination module is specifically configured to: obtain the length of each first target vector and the length of each second target vector; determine a target similarity vector between each first target vector and each second target vector, and determine a target similarity value corresponding to the target similarity vector; wherein, the target similarity vector is the vector in the second target vectors with the highest similarity value to the first target vector; based on the length of each first target vector, the length of each second target vector, and the target similarity value corresponding to the target similarity vector, obtain a similarity score of each first target vector and each second target vector; based on at least one of the similarity scores, obtain the similarity between the to-be-tested code file and the preset vector database.
[0020] Combined with the second aspect and the above implementation manners, in some implementation manners of the second aspect, the second determination module includes a first generation unit, and the first generation unit is specifically configured to: when there are multiple to-be-tested code files, determine a first similarity score between the first target directory vector and the second target directory vector, a second similarity score between the first target code vector and the second target code vector, and a third similarity score between the first target syntax tree vector and the second target syntax tree vector; when there is one to-be-tested code file, determine a second similarity score between the first target code vector and the second target code vector and a third similarity score between the first target syntax tree vector and the second target syntax tree vector.
[0021] Combined with the second aspect and the above implementation manners, in some implementation manners of the second aspect, the second determination module further includes a second generation unit, and the second generation unit is specifically configured to: when there are multiple to-be-tested code files, fuse the first similarity score, the second similarity score, and the third similarity score to obtain the similarity between the to-be-tested code file and the preset vector database; when there is one to-be-tested code file, fuse the second similarity score and the third similarity score to obtain the similarity between the to-be-tested code file and the preset vector database.
[0022] In a third aspect, an electronic device is provided, including a memory and a processor. The memory is used to store executable program code, and the processor is used to call and run the executable program code from the memory, so that the electronic device executes the similarity detection method in the above first aspect and any possible implementation of the first aspect.
[0023] In a fourth aspect, a computer program product is provided, which includes: computer program code. When the computer program code runs on a computer, the computer is enabled to execute the similarity detection method in the above first aspect and any possible implementation of the first aspect.
[0024] In a fifth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer program code. When the computer program code runs on a computer, the computer is enabled to execute the similarity detection method in the above first aspect and any possible implementation of the first aspect. Description of the Drawings
[0025] Figure 1 is a schematic flowchart of a similarity detection method provided by an embodiment of the present application;
[0026] Figure 2 is a schematic diagram of the overall architecture of a similarity detection method provided by an embodiment of the present application;
[0027] Figure 3 is a schematic diagram of the structure of a similarity detection device provided by an embodiment of the present application;
[0028] Figure 4 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed Embodiments
[0029] Next, the technical solutions in the present application will be clearly and elaborately described in conjunction with the drawings. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B: "and / or" in the text is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality" means two or more than two.
[0030] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.
[0031] At present, the open-source model has become an important software development model. Open-source technologies are used in almost all enterprise software systems. More and more enterprises and developers improve or develop and expand based on open-source code. The introduction of open-source code has also improved the efficiency and reduced the costs of enterprises.
[0032] However, problems such as low software stability, redundant code libraries, and vulnerability transmission derived from existing open-source code have become increasingly prominent, affecting code management and code security. For example, a software A is developed based on open-source code B. However, when there is a problem with the open-source code B, it is necessary to determine which codes in the corresponding source code file of the software A are problematic, that is, to determine the code segments similar to the open-source code B. The current code detection methods mainly include abstract syntax tree (AST) subtree matching, Program Dependence Graph matching, code hash search, etc. However, these methods have problems such as low detection efficiency or low accuracy.
[0033] To solve the above technical problems, an embodiment of the present application provides a method for detecting similarity. The execution subject of this method can be an electronic device, which can be a tablet computer, a notebook computer, an ultra-mobile personal computer (UMPC), etc. The embodiment of the present application does not impose any restrictions on the specific type of the electronic device.
[0034] Figure 1 It is a schematic flowchart of a method for detecting similarity provided by an embodiment of the present application.
[0035] Exemplarily, as Figure 1 shown, the method 100 includes:
[0036] S101, obtain at least one code file to be detected.
[0037] S102, perform vectorization processing on the at least one code file to be detected to obtain at least one first target vector.
[0038] S103, determine at least one second target vector in a preset vector database whose similarity value with the at least one first target vector is greater than a preset threshold.
[0039] S104, based on the at least one first target vector and the at least one second target vector, determine the similarity between the code file to be detected and the preset vector database.
[0040] A method for detecting similarity provided by an embodiment of the present application first obtains at least one code file to be tested, then performs vectorization processing on the at least one code file to be tested to obtain at least one first target vector, and then determines at least one second target vector in a preset vector database whose similarity value with the at least one first target vector is greater than a preset threshold. Furthermore, based on the at least one first target vector and the at least one second target vector, the similarity between the code file to be tested and the preset vector database is determined. By vectorizing at least one code file to be tested, complex relationships and representations between codes can be extracted, and at least one second target vector whose similarity value with the at least one first target vector is greater than a preset threshold can be quickly determined in the preset vector database. Finally, the similarity between the code file to be tested and the preset vector database can be more accurately determined through the similar vectors of the at least one first target vector. The above method uses vector similarity matching technology and can effectively improve the accuracy of detecting homologous codes.
[0041] The following specifically describes the implementation manners of each step in the Figure 1 illustrated embodiment:
[0042] Regarding S101 above, it can be understood that the obtained code file to be tested can be one code file to be tested or multiple code files to be tested. It should be understood that when there are multiple code files to be tested, the multiple code files to be tested can form a code repository, that is, the code repository contains multiple code files to be tested.
[0043] In some embodiments, when one code file to be tested is obtained, similarity detection is performed on the single code file to be tested; when multiple code files to be tested are obtained, similarity detection is performed on the code repository formed by the multiple code files to be tested.
[0044] Regarding S102 above, it can be understood that the at least one code file to be tested can be vectorized through a preset vectorization model.
[0045] Exemplarily, the preset vectorization model can be the Sentence Transformers model. Sentence-transformers is an open-source toolkit for single-language, cross-language sentence and text, and image embeddings. It can also specifically be the all-mpnet-base-v2 model in the Sentence Transformers model. all-mpnet-base-v2 can be used for tasks such as clustering or semantic search.
[0046] In some embodiments, when multiple code files to be tested, i.e., a code repository, are obtained, vectorization processing is performed on the code repository. Specifically, vectorization processing is performed on the repository directory information, code snippets, and code syntax trees corresponding to the code repository respectively to obtain multiple first target vectors.
[0047] It can be understood that a vectorization model is used to perform vectorization processing on the repository directory information, code snippets, and code syntax trees respectively to obtain a first target directory vector, a first target code vector, and a first target syntax tree vector.
[0048] In other embodiments, when a single code file to be tested is obtained, vectorization processing is performed on the code file to be tested. Specifically, vectorization processing is performed on the code snippets and code syntax trees corresponding to the code file to be tested respectively to obtain multiple first target vectors.
[0049] It can be understood that a vectorization model is used to perform vectorization processing on the code snippets and code syntax trees respectively to obtain a first target code vector and a first target syntax tree vector.
[0050] It should be understood that since the vectorization model usually has a limit on the code length, if the code length in the code file exceeds the preset length, the code needs to be segmented. However, if the code is directly segmented according to the preset length, the syntax structure of the code may be interrupted. Therefore, when segmenting the code, not only the code length but also the syntax structure of the code need to be considered.
[0051] In a possible implementation, the method further includes: obtaining the at least one code file to be tested; extracting the syntax tree of the at least one code file to be tested to obtain the syntax tree of the at least one code file to be tested; and based on the syntax tree of the at least one code file to be tested and the preset length, performing segmentation processing on the at least one code file to be tested to obtain the code snippets of the at least one code file to be tested.
[0052] It can be understood that the above-mentioned extraction of the syntax tree of the code file to be tested can be specifically achieved by first performing lexical analysis on the source code through a compiler to decompose the source code into a series of tokens, and then performing syntax analysis on these tokens to combine them into higher-level structures such as expressions, declarations, and control structures according to the syntax rules of the programming language, and combining these structures to obtain the syntax tree. It should be understood that tokens are the smallest syntax units in the source code, such as keywords, identifiers, operators, literals, etc.
[0053] The above-mentioned preset length can be set according to the length limit of the vectorization model. For example, if the maximum code length that the vectorization model can input is 512 tokens, then the preset length is set to 512 tokens.
[0054] After extracting the syntax tree of the above code file to be tested, based on the syntax tree structure of the code file to be tested, the code in the code file to be tested can be segmented into code segments of a preset length.
[0055] Regarding the above S103, it can be understood that the above preset vector database may include various vector databases, such as a directory vector library, a code vector library, and a syntax tree vector library, etc.
[0056] Specifically, after obtaining at least one first target vector, according to the vector type of the first target vector, a second target vector similar to the first target vector is determined in different vector databases.
[0057] In a possible implementation, determining at least one second target vector in the preset vector database whose similarity value with the at least one first target vector is greater than a preset threshold includes: when there are multiple code files to be tested, respectively determining in the preset vector database a second target directory vector, a second target code vector, and a second target syntax tree vector whose similarity values with the first target directory vector, the first target code vector, and the first target syntax tree vector are greater than the preset threshold.
[0058] It can be understood that the above preset threshold can be custom-set according to the actual situation, for example, set to 0.6.
[0059] Specifically, when there are multiple code files to be tested, the first target directory vector is compared with the directory vector library in the preset vector database, and based on a preset vector retrieval algorithm, the cosine similarity between the first target directory vector and multiple directory vectors in the directory vector library is retrieved. If the cosine similarity between the directory vector in the directory vector library and the first target directory vector is greater than the preset threshold, then the directory vector is determined as the second target directory vector.
[0060] The first target code vector is compared with the code vector library in the preset vector database, and based on a preset vector retrieval algorithm, the cosine similarity between the first target code vector and multiple code vectors in the code vector library is retrieved. If the cosine similarity between the code vector in the code vector library and the first target code vector is greater than the preset threshold, then the code vector is determined as the second target code vector.
[0061] The first target syntax tree vector is compared with the syntax tree vector library in the preset vector database, and the cosine similarity between the first target syntax tree vector and multiple syntax tree vectors in the syntax tree vector library is retrieved based on a preset vector retrieval algorithm. If the cosine similarity between the syntax tree vector in the syntax tree vector library and the first target syntax tree vector is greater than a preset threshold, the syntax tree vector is determined as the second target syntax tree vector.
[0062] In one possible implementation, determining in a preset vector database at least one second target vector whose similarity value with the at least one first target vector is greater than a preset threshold includes: when there is one code file to be tested, determining in the preset vector database a second target code vector and a second target syntax tree vector whose similarity values with the first target code vector and the first target syntax tree vector are greater than the preset threshold, respectively.
[0063] Specifically, similar to the previous text, when the above-mentioned code file to be tested is one, the cosine similarity between the first target code vector and multiple code vectors in the code vector library is retrieved based on the preset vector retrieval algorithm. If the cosine similarity between the code vector in the code vector library and the first target code vector is greater than a preset threshold, the code vector is determined as the second target code vector.
[0064] Based on a preset vector retrieval algorithm, the cosine similarity between the first target syntax tree vector and multiple syntax tree vectors in the syntax tree vector library is retrieved. If the cosine similarity between the syntax tree vector in the syntax tree vector library and the first target syntax tree vector is greater than a preset threshold, the syntax tree vector is determined as the second target syntax tree vector.
[0065] It is understandable that, since the vectors in the preset vector database that are greater than the preset threshold are determined as the second target vectors, there may be a situation where one first target vector corresponds to multiple second target vectors. When there are multiple second target vectors, the target vector that is most similar to the first target vector can be determined from the multiple second target vectors.
[0066] The above retrieval algorithm can be a graph-based vector retrieval (Hierarchical Navigable Small Word, HNSW) algorithm, which retrieves at least one second target vector whose similarity value with the at least one first target vector is greater than a preset threshold value such as 0.6 in a preset vector library through the HNSW algorithm.
[0067] In a possible implementation manner, determining the similarity between the code file to be tested and the preset vector database based on the at least one first target vector and the at least one second target vector includes: obtaining the length of each first target vector and the length of each second target vector; determining a target similarity vector between each first target vector and each second target vector, and determining a target similarity value corresponding to the target similarity vector; wherein, the target similarity vector is the vector in the second target vectors with the highest similarity value to the first target vector; based on the length of each first target vector, the length of each second target vector, and the target similarity value corresponding to the target similarity vector, obtaining a similarity score for each first target vector and each second target vector; and based on at least one of the similarity scores, obtaining the similarity between the code file to be tested and the preset vector database.
[0068] It can be understood that the lengths of the above-mentioned first target vector and second target vector can be understood as the number of target vectors. As described above, one first target vector can correspond to multiple second target vectors. The most similar target similarity vector can be determined among the multiple second target vectors, and the target similarity value between the first target vector and the target similarity vector can be obtained. It can be understood that if there are multiple first target vectors, there will be multiple target similarity values.
[0069] In some embodiments, when there are multiple code files to be tested, that is, when testing a code repository, and the first target vectors corresponding to the code repository to be tested include a first target directory vector, a first target code vector, and a first target syntax tree vector, the second target directory vector, the second target code vector, and the second target syntax tree vector determined by the above method all contain multiple target similarity vectors.
[0070] Furthermore, the similarity between the code repository to be tested and the preset vector library in different dimensions can be determined through the second target directory vector, the second target code vector, and the second target syntax tree vector respectively. For example, according to the second target directory vector, a code repository similar to the code repository to be tested in terms of directory information can be determined; according to the second target code vector, the similarity between the code repository to be tested and the preset vector library in terms of code segments can be determined; and according to the second target syntax tree vector, the similarity between the code repository to be tested and the preset vector library in terms of syntax trees can be determined.
[0071] In the above method, when there are multiple code files to be detected, that is, when detecting a code repository, the directory structure of the entire code repository is extracted, and similar directories are detected in the preset vector library, so that homologous code directories can be identified, and the similarity between the code files to be tested and the preset vector library in terms of directory information can be determined; code snippets are extracted, and after vectorizing the code snippets, similar code snippets are detected in the preset vector library, and the similarity between the code files to be tested and the preset vector library in terms of code snippets can be determined; the syntax tree of the code files to be tested can also be extracted, and after vectorizing the syntax tree, similar syntax trees are detected in the preset vector library, and the similarity between the code files to be tested and the preset vector library in terms of code structure, that is, the syntax tree, can be determined. By extracting multiple features of the code files to be tested and calculating the similarity in multiple dimensions between the code files to be tested and the preset vector library, the accuracy of similarity detection can be improved.
[0072] In a possible implementation, obtaining the similarity score between each first target vector and each second target vector based on the length of each first target vector, the length of each second target vector, and the target similarity value corresponding to the target similarity vector includes: when there are multiple code files to be detected, determining a first similarity score between the first target directory vector and the second target directory vector, a second similarity score between the first target code vector and the second target code vector, and a third similarity score between the first target syntax tree vector and the second target syntax tree vector; when there is one code file to be detected, determining a second similarity score between the first target code vector and the second target code vector and a third similarity score between the first target syntax tree vector and the second target syntax tree vector.
[0073] It can be understood that the above second target vector is a vector in the preset vector library. Calculating the similarity score between the first target vector and the second target vector is actually calculating the similarity score between the code file to be tested and the preset vector library. As described above, the preset vector library may include a directory vector library, a code vector library, and a syntax tree vector library. The above first similarity score may be the similarity score between the code file to be tested and the directory vector library, the above second similarity score may be the similarity score between the code file to be tested and the code vector library, and the above third similarity score may be the similarity score between the code file to be tested and the syntax tree vector library.
[0074] Specifically, when there are multiple code files to be detected, that is, when detecting a code repository, the first similarity score S1 between the code repository and the preset vector library, that is, the directory vector library, in terms of directory information can be determined by the following method:
[0075] Vectorize the directory information of code repository A to obtain a first target directory vector, which contains n directory vectors. Retrieve from the directory vector library a second target directory vector that is similar (cosine similarity greater than 0.6) to these n directory vectors. There are a total of m such second target directory vectors. If there are y similar directory vectors corresponding to a directory vector X among the n directory vectors, determine the directory vector Y that is most similar to the directory vector X among these y similar directory vectors, and determine the cosine similarity, i.e., the similarity value, between vector X and vector Y. Similarly, determine the directory vector that is most similar to each of the n directory vectors and obtain n similarity values. Sum up these similarity values to get N. Calculate the first similarity score S1 between the directory vector of code repository A and the preset vector library: the sum N of the n similarity values divided by the sum (m + n) of the number of the first target directory vector and the second target directory vector minus N, i.e., S1 = N / ((m + n) - N).
[0076] Similarly, through the above calculation method, the similarity S2 between the code repository in terms of code snippets and the code vector library, and the similarity S3 between the code repository in terms of syntax trees and the syntax tree vector library can be obtained.
[0077] In some other embodiments, if there is one code file to be tested, and the first target vector corresponding to the code file to be tested includes a first target code vector and a first target syntax tree vector, then among the second target code vector and the second target syntax tree vector determined by the above method, both contain multiple target similar vectors.
[0078] Specifically, the similarity S2 between the code file to be tested in terms of code snippets and the code vector library, and the similarity S3 between the code repository in terms of syntax trees and the syntax tree vector library can be determined in the manner shown above.
[0079] Further, after obtaining the above first similarity score, the second similarity score, and the third similarity score, the similarity between the code file to be tested and the preset vector database can be calculated based on the first similarity score, the second similarity score, and the third similarity score.
[0080] In a possible implementation, obtaining the similarity between the code file to be tested and the preset vector database based on at least one of the similarity scores includes: when there are multiple code files to be tested, fusing the first similarity score, the second similarity score, and the third similarity score to obtain the similarity between the code file to be tested and the preset vector database; when there is one code file to be tested, fusing the second similarity score and the third similarity score to obtain the similarity between the code file to be tested and the preset vector database.
[0081] It can be understood that the fusion of the above first similarity score, the second similarity score, and the third similarity score, or the fusion of the second similarity score and the third similarity score, can be a weighted sum. Specifically, the similarity of each dimension can be weighted and summed according to the degree of importance.
[0082] It should be understood that for the code file to be tested, the degrees of importance of the similarities between the code file to be tested and the preset vector library in different dimensions are different. For example, if the core of the code file to be tested is the code snippet, then the similarity between the code file to be tested and the preset vector library in terms of the code snippet is the most important. Based on this, when determining the similarity between the code file to be tested and the preset vector database, the similarities of each dimension are weighted and summed according to the degree of importance.
[0083] Exemplarily, the weights between the above first similarity score, second similarity score, and third similarity score can be set according to the degrees of importance of the similarities between the code file to be tested and the preset vector library in different dimensions. For example, for multiple code files to be tested, that is, the code repository, the most important is the similarity with the code snippets in the preset vector library. Then the weight of the second similarity score corresponding to the code snippet is set to 0.7. The second most important is the similarity with the directory information in the preset vector library. Then the weight of the first similarity score corresponding to the directory information is set to 0.2. Finally, it is the similarity with the syntax tree in the preset vector library. The weight of the third similarity score corresponding to the syntax tree is set to 0.1. That is, when there are multiple code files to be tested, the similarity between the code file to be tested and the preset vector library is: 0.2 * first similarity score + 0.7 * second similarity score + 0.1 * third similarity score.
[0084] Similarly, for a single code file to be tested, the most important is the similarity with the code snippets in the preset vector library. Then the weight of the second similarity score corresponding to the code snippet is set to 0.7. The second most important is the similarity with the syntax tree in the preset vector library. The weight of the third similarity score corresponding to the syntax tree is set to 0.3. That is, when there is a single code file to be tested, the similarity between the code file to be tested and the preset vector library is: 0.7 * second similarity score + 0.3 * third similarity score.
[0085] After obtaining the similarities between the code file to be tested and the preset vector library in multiple dimensions by the above method, these similarities can be weighted and calculated, which can not only consider the similarity of code snippets, but also consider the similarities of directory information and syntax tree. And by setting the weight ratio according to the degrees of importance of the similarities between the code file to be tested and the preset vector library in different dimensions, the accuracy of code detection can be further improved.
[0086] Based on Figure 1The similarity detection method shown, in some embodiments, vectorizes a code repository A to be detected in three dimensions, that is, vectorizes the repository directory, code files, and syntax tree of the code repository A respectively to obtain a first directory vector, a first code vector, and a first syntax tree vector. Compare the obtained first directory vector, first code vector, and first syntax tree vector with the corresponding directory vector library, code vector library, and syntax tree vector library respectively, that is, perform retrieval in the directory vector library, code vector library, and syntax tree vector library respectively, and determine vectors with a cosine similarity higher than 0.6 in these vector libraries, which can be respectively called the second directory vector, the second code vector, and the second syntax tree vector. Based on the second directory vector, second code vector, and second syntax tree vector determined above, determine the code repositories similar to the repository directory, code files, and syntax tree of the code repository A.
[0087] For example, perform similarity detection on the code repository A: First, obtain the directory structure information of the code repository A, vectorize the obtained directory structure information to obtain 100 directory vectors, and then perform retrieval in the directory vector library to obtain 300 vectors similar to the 100 directory vectors. For example, the code vector M is most similar to the code vector N, and the similarity value between M and N is 1. Similarly, obtain the directory vectors in the directory vector library that are most similar to the 100 directory vectors, and obtain the similarity values between the 100 directory vectors and their most similar directory vectors, a total of 100 similarity values. Assume that the sum of the 100 similarity values is Q. The number of directory vectors of the code repository A is 100, and 300 vectors similar to the directory vectors of the code repository A are retrieved in the directory vector library. The first similarity score of the code repository A in terms of directory information with respect to the preset vector library, that is, the directory vector library, is: Q / (100 + 300 - Q).
[0088] Obtain multiple code files of the code repository A, extract the syntax trees of the multiple code files and perform segmentation to obtain multiple code segments, vectorize the multiple code segments to obtain 300 code vectors, and then perform retrieval in the code vector library to obtain 600 vectors similar to the 300 code vectors. For example, the code vector X is most similar to the code vector Y, and the similarity value between X and Y is 0.9. Similarly, obtain the code vectors in the code vector library that are most similar to the 300 code vectors, and obtain the similarity values between the 300 code vectors and their most similar code vectors, a total of 300 similarity values. Assume that the sum of the 300 similarity values is Z. The number of code vectors in the code repository A is 300, and 600 vectors similar to the code vectors of the code repository A are retrieved in the code vector library. The second similarity score of the code repository A in terms of code segments with respect to the preset vector library, that is, the code vector library, is: Z / (300 + 600 - Z).
[0089] The syntax trees of multiple code files in code repository A are vectorized to obtain 200 syntax tree vectors, and then retrieved in the syntax tree vector library to obtain 300 vectors similar to the 200 syntax tree vectors. For example, the syntax tree vector Q is the most similar to a syntax tree vector P, and the similarity value between P and Q is 1. Similarly, the syntax tree vectors in the syntax tree vector library that are the most similar to the 200 syntax tree vectors are obtained, and the similarity values between the 200 syntax tree vectors and their most similar syntax tree vectors are obtained, a total of 200 similarity values. Suppose the sum of the 200 similarity values is L. The number of syntax tree vectors in code repository A is 200, and 300 vectors similar to the syntax tree vectors of code repository A are retrieved in the syntax tree vector library. The third similarity score of code repository A with respect to the preset vector library, i.e., the syntax tree vector library, in terms of syntax trees is: L / (200 + 300 - L).
[0090] Finally, the first similarity score, the second similarity score, and the third similarity score are added according to their weights to obtain the similarity between code repository A and the preset vector library. For example, if the weights of the above directory information, code snippets, and syntax trees are 0.1, 0.7, and 0.2 respectively, then the similarity between code repository A and the preset vector library is: 0.1 * (Q / (100 + 300 - Q)) + 0.7 * (Z / (300 + 600 - Z)) + 0.2 * (L / (200 + 300 - L)).
[0091] It can be understood that when determining the labels of the code files in the preset vector library, the method of clustering can be used to quickly determine them. By determining the labels of the code files in the preset vector library, it is possible to classify and paste labels on each code file in the code repository, which is convenient for subsequent repository search and management.
[0092] In some embodiments, code vectors corresponding to multiple code snippets in multiple code files in the preset vector database are obtained; based on a preset clustering model, the multiple code vectors are clustered to obtain multiple labels corresponding to the multiple code vectors; the proportion value of each label in the multiple labels is determined, and the label with the highest proportion value is determined as the first target label of the multiple code files; wherein, the first target label is used to search for the multiple code files in the preset vector database.
[0093] It can be understood that when it is necessary to determine the labels of multiple code files in the preset vector library, the multiple code vectors corresponding to the multiple code files can be clustered based on a preset clustering model, and the first target label of the multiple code files can be accurately determined.
[0094] Exemplarily, the above-mentioned preset clustering model can be a Gaussian Mixture Model (GMM), and the code vectors can also be clustered by clustering algorithms such as k-means algorithm and bi-kmeans algorithm.
[0095] It can be understood that the code vectors for clustering are source code vectors, and the labels obtained by the above clustering can be the labels involved in training the clustering model. For example, the labels obtained during training include front-end code, server-side code, algorithm code, etc. Inputting the code vectors into the preset clustering model can obtain the labels of the code vectors.
[0096] Exemplarily, inputting 10 code vectors corresponding to a code file into the preset clustering model, the labels of 7 code vectors are front-end code, the labels of 2 code vectors are server-side code, and the label of 1 code vector is server-side code. Based on this, it can be determined that the label with the highest proportion is front-end code, so the target label of this code file is front-end code.
[0097] In some embodiments, the category label of the code file to be tested can also be determined based on the similarity between the code file to be tested and the preset vector library. First, obtain the similarity between the code file to be tested and the preset vector database; determine the first target label corresponding to the code file with the highest similarity to the code file to be tested in the preset vector database as the second target label of the code file to be tested; where the second target label is used to mark the category of the code file to be tested.
[0098] It can be understood that for a new code file to be tested, the similarity of the code file to be tested can be detected first, the target code file most similar to the code file to be tested in the preset vector library can be determined, and the first target label of the target code file can be determined as the second target label of the code file to be tested.
[0099] The above method can quickly and accurately determine the label corresponding to the code file to be tested when determining the labels corresponding to a large number of code files to be tested by clustering the code vectors through a preset clustering model; for a new code file to be tested, determining the label according to the similarity between the code file to be tested and the preset vector library can ensure that code files with the same source or similar codes have the same category label, which is convenient for managing code files.
[0100] Figure 2 It is the overall architecture diagram of a similarity detection method provided by an embodiment of the present application.
[0101] Exemplarily, such as Figure 2As shown in the figure, the overall solution of this similarity detection method can include two parts:
[0102] The first part: Based on the already constructed code vector library, perform similarity retrieval on the new code file or code repository, calculate the similarity scores respectively from three aspects: the repository directory, the code file, and the code syntax tree. The score of the similar code is jointly determined by the similarity of the code content and the syntax tree, and the similarity of the repository is jointly determined by these three to determine the similarity of the code repository.
[0103] Exemplarily, as Figure 2 shown, divide the code repository according to three dimensions, and vectorize the repository directory, code file, and code syntax tree of this code repository through the Embedding model. Among them, after the code file needs to be code-sliced, the code segments are then input into the Embedding model for vectorization. Specifically, how to slice the code file can refer to the method in the previous text and will not be elaborated here. At the same time, it is also necessary to perform syntax tree segmentation on the code syntax tree, and then input the syntax tree into the Embedding model for vectorization. Then, the vectorized directory vector, code vector, and syntax tree vector can be respectively calculated for similarity with the directory vector library, code vector library, and syntax tree vector library to obtain the first similarity score score1, the second similarity score score2, and the third similarity score score3. Finally, based on the first similarity score score1, the second similarity score score2, and the third similarity score score3, the code repository similar to this code repository can be determined, and at the same time, based on the second similarity score score2 and the third similarity score score3, the code file similar to the code file in this code repository can be determined.
[0104] The second part: Label a large number of existing code repositories in the vector library, cluster them through the clustering model, and paste the same category label on the code repositories in the same cluster. For the new code repository, use the label of the code repository most similar to it as the predicted category label.
[0105] Exemplarily, as Figure 2 shown, the category labels of a large number of code files or code repositories in the preset vector library can be determined through the clustering model. For the new code repository, the label of this new code repository can be determined according to the label of the similar code repository of this code repository.
[0106] Figure 3 is a schematic structural diagram of a similarity detection device provided by an embodiment of the present application.
[0107] Exemplarily, as Figure 3 shown, the device 300 includes:
[0108] An acquisition module 301, configured to acquire at least one code file to be tested.
[0109] A processing module 302, configured to perform vectorization processing on the at least one code file to be tested to obtain at least one first target vector.
[0110] A first determination module 303, configured to determine at least one second target vector in a preset vector database, where the similarity value between the at least one second target vector and the at least one first target vector is greater than a preset threshold.
[0111] A second determination module 304, configured to determine the similarity between the code file to be tested and the preset vector database based on the at least one first target vector and the at least one second target vector.
[0112] In a possible implementation, when there are multiple code files to be tested, the first target vector includes a first target directory vector, a first target code vector, and a first target syntax tree vector; the first target directory vector is obtained by vectorizing the repository directory information corresponding to the multiple code files to be tested; the first target code vector is obtained by vectorizing the code segments of the at least one code file to be tested, and the first target syntax tree vector is obtained by vectorizing the syntax trees of the at least one code file to be tested; when there is one code file to be tested, the first target vector includes the first target code vector and the first target syntax tree vector.
[0113] Optionally, the apparatus further includes a segmentation processing module, and the segmentation processing module is specifically configured to: acquire the at least one code file to be tested; extract syntax trees from the at least one code file to be tested to obtain the syntax trees of the at least one code file to be tested; and perform segmentation processing on the at least one code file to be tested based on the syntax trees of the at least one code file to be tested and a preset length to obtain code segments of the at least one code file to be tested.
[0114] In a possible implementation, the first determination module is specifically configured to: when there are multiple code files to be tested, determine a second target directory vector, a second target code vector, and a second target syntax tree vector in the preset vector database, where the similarity values between the second target directory vector, the second target code vector, and the second target syntax tree vector and the first target directory vector, the first target code vector, and the first target syntax tree vector are greater than the preset threshold, respectively; when there is one code file to be tested, determine a second target code vector and a second target syntax tree vector in the preset vector database, where the similarity values between the second target code vector and the second target syntax tree vector and the first target code vector and the first target syntax tree vector are greater than the preset threshold, respectively.
[0115] In a possible implementation, the second determination module is specifically configured to: obtain the length of each first target vector and the length of each second target vector; determine a target similarity vector between each first target vector and each second target vector, and determine a target similarity value corresponding to the target similarity vector; wherein, the target similarity vector is the vector with the highest similarity value to the first target vector among the second target vectors; based on the length of each first target vector, the length of each second target vector, and the target similarity value corresponding to the target similarity vector, obtain a similarity score for each first target vector and each second target vector; and based on at least one of the similarity scores, obtain the similarity between the code file to be tested and the preset vector database.
[0116] In a possible implementation, the second determination module includes a first generation unit, and the first generation unit is specifically configured to: when there are multiple code files to be tested, determine a first similarity score between the first target directory vector and the second target directory vector, a second similarity score between the first target code vector and the second target code vector, and a third similarity score between the first target syntax tree vector and the second target syntax tree vector; when there is one code file to be tested, determine a second similarity score between the first target code vector and the second target code vector and a third similarity score between the first target syntax tree vector and the second target syntax tree vector.
[0117] In a possible implementation, the second determination module further includes a second generation unit, and the second generation unit is specifically configured to: when there are multiple code files to be tested, fuse the first similarity score, the second similarity score, and the third similarity score to obtain the similarity between the code file to be tested and the preset vector database; when there is one code file to be tested, fuse the second similarity score and the third similarity score to obtain the similarity between the code file to be tested and the preset vector database.
[0118] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application.
[0119] Exemplarily, as Figure 4 shown, the electronic device 400 includes: a memory 401 and a processor 402. Among them, an executable program code 4011 is stored in the memory 401, and the processor 402 is configured to call and execute the executable program code 4011 to execute a similarity detection method.
[0120] In addition, an embodiment of the present application also protects a device, which may include a memory and a processor. Among them, an executable program code is stored in the memory, and the processor is configured to call and execute the executable program code to execute a similarity detection method provided by an embodiment of the present application.
[0121] In this embodiment, the device can be divided into functional modules according to the above method examples. For example, each functional module can be corresponded, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.
[0122] In the case of dividing each functional module corresponding to each function, the device can also include an acquisition module, a processing module, a first determination module, a second determination module, etc. It should be noted that all relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be repeated here.
[0123] It should be understood that the device provided in this embodiment is used to execute the above method for detecting a similarity, so the same effect as the above implementation method can be achieved.
[0124] In the case of adopting an integrated unit, the device can include a processing module and a storage module. Among them, when the device is applied to an electronic device, the processing module can be used to control and manage the actions of the electronic device. The storage module can be used to support the electronic device to execute mutual program codes, etc.
[0125] Among them, the processing module can be a processor or a controller, which can implement or execute various exemplary logical blocks, modules, and circuits described in combination with the disclosure of this application. The processor can also be a combination that realizes computing functions, such as including a combination of one or more microprocessors, a combination of digital signal processing (DSP) and a microprocessor, etc. The storage module can be a memory.
[0126] In addition, the device provided in the embodiment of this application can specifically be a chip, a component, or a module. The chip can include a connected processor and a memory; among them, the memory is used to store instructions, and when the processor calls and executes the instructions, the chip can execute the method for detecting a similarity provided in the above embodiment.
[0127] This embodiment also provides a computer-readable storage medium, in which computer program code is stored. When the computer program code runs on a computer, it causes the computer to execute the above related method steps to implement the method for detecting a similarity provided in the above embodiment.
[0128] This embodiment also provides a computer program product. When the computer program product runs on a computer, it causes the computer to execute the above-related steps to implement a similarity detection method provided by the above embodiment.
[0129] Among them, the device, computer-readable storage medium, computer program product, or chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.
[0130] Through the description of the above embodiments, those skilled in the art can understand that for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0131] In the embodiments provided in this application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0132] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A method for detecting similarity, characterized in that, the method includes: obtaining at least one code file to be tested; vectorizing the at least one code file to be tested to obtain at least one first target vector; determining at least one second target vector in a preset vector database whose similarity value with the at least one first target vector is greater than a preset threshold; determining the similarity between the code file to be tested and the preset vector database based on the at least one first target vector and the at least one second target vector.
2. The method according to claim 1, characterized in that, when there are multiple code files to be tested, the first target vector includes a first target directory vector, a first target code vector, and a first target syntax tree vector; the first target directory vector is obtained by vectorizing the repository directory information corresponding to the multiple code files to be tested; the first target code vector is obtained by vectorizing the code segments of the at least one code file to be tested, and the first target syntax tree vector is obtained by vectorizing the syntax tree of the at least one code file to be tested; when there is one code file to be tested, the first target vector includes the first target code vector and the first target syntax tree vector.
3. The method according to claim 2, characterized in that, the method further includes: obtaining the at least one code file to be tested; extracting the syntax tree of the at least one code file to be tested to obtain the syntax tree of the at least one code file to be tested; performing a segmentation process on the at least one code file to be tested based on the syntax tree of the at least one code file to be tested and a preset length to obtain the code segments of the at least one code file to be tested.
4. The method according to claim 2, characterized in that, the determining at least one second target vector in a preset vector database whose similarity value with the at least one first target vector is greater than a preset threshold includes: when there are multiple code files to be tested, respectively determining a second target directory vector, a second target code vector, and a second target syntax tree vector in the preset vector database whose similarity values with the first target directory vector, the first target code vector, and the first target syntax tree vector are greater than the preset threshold; when there is one code file to be tested, respectively determining a second target code vector and a second target syntax tree vector in the preset vector database whose similarity values with the first target code vector and the first target syntax tree vector are greater than the preset threshold.
5. The method according to claim 1, characterized in that, the determining the similarity between the code file to be tested and the preset vector database based on the at least one first target vector and the at least one second target vector includes: obtaining the length of each first target vector and the length of each second target vector; Determine the target similarity vector between each of the first target vectors and each of the second target vectors, and determine the target similarity value corresponding to the target similarity vector; wherein, the target similarity vector is the vector in the second target vectors with the highest similarity value to the first target vector; Based on the length of each of the first target vectors, the length of each of the second target vectors, and the target similarity value corresponding to the target similarity vector, obtain the similarity score between each of the first target vectors and each of the second target vectors; Based on at least one of the similarity scores, obtain the similarity between the code file to be tested and the preset vector database.
6. The method according to claim 5, wherein, The step of obtaining the similarity score between each of the first target vectors and each of the second target vectors based on the length of each of the first target vectors, the length of each of the second target vectors, and the target similarity value corresponding to the target similarity vector includes: When there are multiple code files to be tested, determine the first similarity score between the first target directory vector and the second target directory vector, the second similarity score between the first target code vector and the second target code vector, and the third similarity score between the first target syntax tree vector and the second target syntax tree vector; When there is one code file to be tested, determine the second similarity score between the first target code vector and the second target code vector, and the third similarity score between the first target syntax tree vector and the second target syntax tree vector.
7. The method according to claim 6, wherein, The step of obtaining the similarity between the code file to be tested and the preset vector database based on at least one of the similarity scores includes: When there are multiple code files to be tested, fuse the first similarity score, the second similarity score, and the third similarity score to obtain the similarity between the code file to be tested and the preset vector database; When there is one code file to be tested, fuse the second similarity score and the third similarity score to obtain the similarity between the code file to be tested and the preset vector database.
8. A similarity detection device, wherein, The device includes: An acquisition module, configured to acquire at least one code file to be tested; A processing module, configured to perform vectorization processing on the at least one code file to be tested to obtain at least one first target vector; A first determination module, configured to determine at least one second target vector in the preset vector database whose similarity value with the at least one first target vector is greater than a preset threshold; A second determination module, configured to determine the similarity between the code file to be tested and the preset vector database based on the at least one first target vector and the at least one second target vector.
9. An electronic device, wherein, The electronic device includes: A memory, configured to store executable program code; A processor, configured to call and run the executable program code from the memory, so that the electronic device executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 is implemented.