System for generating software code right authentication certificate, verification system and method

By combining AI semantic and structural fingerprint technology with developer signatures and trusted timestamps, we generate authentication credentials, solving the problem of traditional methods being difficult to identify advanced plagiarism and code ownership, and achieving efficient plagiarism detection and copyright protection.

CN120611364AActive Publication Date: 2025-09-09BEIJING UNITED TRUST TECH SERVICE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511079946.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-09-09
Estimated Expiration
2045-08-04

AI Technical Summary

Technical Problem

Traditional code plagiarism detection methods are difficult to identify advanced "skin-changing" plagiarism, and the code creation time and original ownership are difficult to prove conclusively, affecting the protection of developers' rights and interests.

Method used

AI semantic analysis is used to generate binary semantic fingerprints and structural fingerprints. Combined with the developer's signature and trusted timestamp, the code DNA core package is constructed. The timestamp is anchored through the trusted timestamp service platform to generate the right authentication certificate.

Benefits of technology

It improves the accuracy and reliability of plagiarism detection, provides tamper-proof timestamp evidence, ensures code originality, and protects the rights of developers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611364A_ABST
    Figure CN120611364A_ABST
Patent Text Reader

Abstract

The invention discloses a system for generating a software code right authentication certificate and a verification system and method, and the system comprises an AI semantic analysis module which is used for analyzing a target code through an AI code large model, and generating a high-dimensional semantic embedding vector; the semantic fingerprint generation module is used for hashing the vector to obtain a binary semantic fingerprint; the extraction module is used for extracting a plurality of structural features from the abstract syntax tree; the structure fingerprint generation module is used for hashing the plurality of structure features to obtain a structure fingerprint; the core package construction module is used for constructing a code DNA core package; the signature module is used for encrypting the hash value of the core packet by using a private key; the trusted timestamp acquisition module is used for carrying out overall hash to obtain a timestamp authentication hash value, and obtaining a trusted timestamp certificate according to the timestamp authentication hash value; and the code DNA core package, the developer signature and the trusted timestamp certificate jointly serve as a right authentication certificate. According to the method, the accuracy of code plagiarism detection can be improved, and reliable copyright protection is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of code copyright protection, and in particular to a system and method for generating software code ownership authentication credentials, and a verification system and method for software code ownership authentication credentials. Background Art

[0002] Traditional code plagiarism detection methods, such as those based on text matching or simple hashing, primarily focus on the surface form of the code. Advanced "skin-changing" plagiarism, however, involves replacing variable and function names, fine-tuning the code's logical structure, and rewriting textual expressions. This creates surface differences from the original code, while preserving the core content and structure. This makes it difficult for traditional methods to accurately identify this type of plagiarism, making it unable to meet the growing demand for code copyright protection.

[0003] Furthermore, during the software development process, it's often difficult to definitively prove the creation time and originality of code. Developers may face the possibility of their code being plagiarized, but lacking effective time stamps and unalterable evidence, they are unable to definitively prove their code's originality and creation time. This puts developers at a disadvantage in copyright disputes, hindering the motivation and rights protection of code creators. Summary of the Invention

[0004] In view of this, embodiments of the present application provide a system and method for generating software code ownership authentication credentials, and a system and method for verifying software code ownership authentication credentials, to solve at least one of the above technical problems.

[0005] In a first aspect, an embodiment of the present application provides a system for generating a software code authentication certificate, comprising: AI semantic analysis module, which uses a pre-trained AI code model to perform semantic analysis on the target code content and generate high-dimensional semantic embedding vectors based on the extracted features; Semantic fingerprint generation module, used to perform hash calculation on high-dimensional semantic embedding vectors to obtain the binary semantic fingerprint of the target code; A structural feature extraction module is used to construct an abstract syntax tree of the target code using a static code analysis tool, and extract multiple structural features of the target code from the abstract syntax tree; The structural fingerprint generation module is used to arrange and combine the extracted multiple structural features and then perform hash calculations on them to obtain the structural fingerprint of the target code; A core package building module is used to receive the binary semantic fingerprint, structural fingerprint, unique identifier, developer digital identity and key submission metadata of the target code, and construct a code DNA core package based on the received data; The private key signature module is used to encrypt the hash value of the code DNA core package using the developer's private key and use the encrypted hash value as the developer's signature; The trusted timestamp acquisition module is used to hash the code DNA core package and the developer's signature as a whole to obtain the timestamp authentication hash value; and send the timestamp authentication hash value to the trusted timestamp service platform and receive the trusted timestamp certificate returned by it; The ownership authentication module is used to use the code DNA core package, developer signature and trusted timestamp certificate as the ownership authentication credentials of the target code.

[0006] According to some embodiments of the present application, optionally, the AI ​​semantic analysis module is specifically used to use the AI ​​code large model to identify multiple elements such as variable names, function names, and operators in the target code, and perform feature extraction on the multiple elements; and output a high-dimensional semantic embedding vector by performing weighted summation or nonlinear transformation operations on the features of the multiple elements, wherein the high-dimensional semantic embedding vector is a multidimensional array.

[0007] According to some embodiments of the present application, optionally, the multiple structural features of the target code include cyclomatic complexity, a topological summary of a function call graph, a feature vector of a module dependency graph, and a structural pattern occurrence frequency of a specific design pattern.

[0008] According to some embodiments of the present application, optionally, the structural fingerprint generation module is specifically used to arrange multiple structural features according to a preset arrangement order; then normalize the arranged multiple structural features to map the characteristic values ​​of multiple structural features in different ranges to a unified range; perform weighted splicing combination processing on the normalized multiple structural features or combine them through a small hash network to obtain a new feature vector through combination; perform hash calculation on the new feature vector to obtain the structural fingerprint of the target code.

[0009] In a second aspect, an embodiment of the present application provides a verification system for a software code ownership authentication credential, wherein the software code ownership authentication credential includes a forged software code ownership authentication credential or a software code ownership authentication credential generated by the system for generating a software code ownership authentication credential as provided in the first aspect. The verification system for the software code ownership authentication credential includes: A rights authentication certificate acquisition module is used to obtain a first rights authentication certificate for the software code to be verified, the first rights authentication certificate including a first code DNA core package, a first developer signature, and a first trusted timestamp certificate; A query module is configured to query a database for a second, confirmed authentication certificate corresponding to the first authentication certificate; if so, to call a subsequent module to further verify the first authentication certificate; if not, to output a verification failure; the second authentication certificate includes a second code DNA core package, a second developer signature, and a second trusted timestamp certificate; A code DNA core package verification module, configured to verify the consistency of the signature chain, time, and hash value in the first code DNA core package based on the second code DNA core package and its associated public key; A developer signature verification module, configured to decrypt the first developer signature using the developer's public key and compare the decrypted result with the hash value of the second code DNA core package; The trusted timestamp certificate verification module is used to send the first trusted timestamp certificate to the trusted timestamp service platform and receive the verification result of the first trusted timestamp certificate returned by the trusted timestamp service platform.

[0010] According to some embodiments of the present application, optionally, the code DNA core package verification module is specifically used to start from the starting point of the signature chain, and use the corresponding public key to verify each signature in sequence until the entire signature chain is verified; compare the time points of different operations recorded by the first code DNA core package with the time points of different operations recorded by the second code DNA core package one by one; compare multiple hash values ​​in the first code DNA core package with multiple hash values ​​in the second code DNA core package one by one, wherein the multiple hash values ​​include at least binary semantic fingerprints and structural fingerprints.

[0011] According to some embodiments of the present application, optionally, the verification system for software code ownership authentication credentials also includes: a fingerprint similarity analysis module, which is used to calculate a first similarity score between the binary semantic fingerprint in the first code DNA core package and the binary semantic fingerprint in the second code DNA core package, calculate a second similarity score between the structural fingerprint in the first code DNA core package and the structural fingerprint in the second code DNA core package, and perform a weighted sum of the first similarity score and the second similarity score according to preset weights to obtain a target similarity score, which is used for version evolution tracing or copyright ownership determination.

[0012] According to some embodiments of the present application, optionally, the verification system for software code ownership authentication credentials also includes: an evolution tracing module, which is used to construct the evolution relationship of the software codes of the first code DNA core package and the second code DNA core package based on the code version information of the first code DNA core package and the second code DNA core package, and the timestamp information of the first trusted timestamp certificate and the second trusted timestamp certificate when there are differences between the first code DNA core package and the second code DNA core package and the target similarity score is greater than a preset score, and input the evolution relationship into a visualization tool.

[0013] In a third aspect, an embodiment of the present application provides a method for generating a software code ownership authentication credential, characterized in that the method is implemented based on the system for generating a software code ownership authentication credential provided in the first aspect, and includes: Use the pre-trained AI code model to perform semantic analysis on the target code content and generate a high-dimensional semantic embedding vector based on the extracted features; Hash the high-dimensional semantic embedding vector to obtain the binary semantic fingerprint of the target code; Using static code analysis tools to construct an abstract syntax tree of the target code and extract multiple structural features of the target code from the abstract syntax tree; The extracted multiple structural features are arranged and combined, and then hashed to obtain the structural fingerprint of the target code; Receive the binary semantic fingerprint, structural fingerprint, unique identifier, developer digital identity and key submission metadata of the target code, and construct the code DNA core package based on the received data; Use the developer's private key to encrypt the hash value of the CodeDNA core package, and use the encrypted hash value as the developer's signature; Hash the CodeDNA core package and the developer's signature as a whole to obtain a timestamp authentication hash value; then send the timestamp authentication hash value to the trusted timestamp service platform and receive the trusted timestamp certificate returned by it; The code DNA core package, developer signature and trusted timestamp certificate are used together as the target code’s authentication credentials.

[0014] In a fourth aspect, an embodiment of the present application provides a method for verifying a software code ownership authentication credential. The method is implemented based on the verification system for software code ownership authentication credential provided in the second aspect, and includes: Obtain the first authentication certificate for the software code to be verified, which includes the first code DNA core package, the first developer's signature, and the first trusted timestamp certificate; Query the database to see if there is a second confirmed authentication certificate corresponding to the first authentication certificate; if so, call the subsequent module to further verify the first authentication certificate; if not, output verification failure; the second authentication certificate includes the second code DNA core package, the second developer signature, and the second trusted timestamp certificate; Based on the second code DNA core package and its associated public key, verify the consistency of the signature chain, time and hash value in the first code DNA core package; Decrypt the first developer's signature using the developer's public key and compare the decrypted result with the hash value of the second code DNA core package; The first trusted timestamp certificate is sent to the trusted timestamp service platform, and a verification result of the first trusted timestamp certificate returned by the platform is received.

[0015] The system and method for generating software code ownership authentication credentials, as well as the system and method for verifying software code ownership authentication credentials, provided in the embodiments of this application, utilize "semantic + structural" dual fingerprinting technology to analyze code from multiple dimensions. The semantic fingerprint provides a deep understanding of the code's functional semantics, while the structural fingerprint captures the code's structural characteristics. This multi-dimensional analysis approach comprehensively and meticulously characterizes the essential characteristics of the code. Compared to traditional methods, it can more accurately identify various forms of plagiarism, especially advanced "skin-changing" plagiarism. For example, in practical applications, traditional hashing methods may misidentify some plagiarized code with complex disguises as original. However, dual fingerprinting technology can accurately identify the plagiarized nature through in-depth analysis of semantics and structure, significantly improving the accuracy and reliability of plagiarism detection. Furthermore, the "Code DNA Core Package," which contains the dual code fingerprints, developer signatures, and key metadata, is timestamped using the Trusted Timestamp Service Platform (TSA). The timestamp provided by the TSA is authoritative and immutable, accurately recording the code's existence at a specific point in time. Once a copyright dispute occurs, this timestamp can serve as strong evidence to ensure the non-repudiation of the originality of the code and provide developers with reliable copyright protection. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings in the embodiments of the present application.

[0017] Figure 1 A structural block diagram of a system for generating software code ownership authentication credentials provided in an embodiment of the present application.

[0018] Figure 2 A structural block diagram of a verification system for software code ownership authentication credentials provided in an embodiment of the present application.

[0019] Figure 3 Another structural block diagram of a verification system for software code ownership authentication credentials provided in an embodiment of the present application.

[0020] Figure 4 A flowchart of a method for generating software code authentication credentials provided in an embodiment of the present application.

[0021] Figure 5 A flowchart of a method for verifying software code ownership authentication credentials provided in an embodiment of the present application.

[0022] Figure 6 A schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0023] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.

[0024] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.

[0025] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0026] It will be apparent to those skilled in the art that various modifications and variations can be made to this application without departing from the spirit or scope of this application. Therefore, this application is intended to cover modifications and variations of this application that fall within the scope of the corresponding claims (technical solutions claimed for protection) and their equivalents. It should be noted that the embodiments provided in the examples of this application may be combined with each other unless there is any inconsistency.

[0027] Before describing the technical solutions provided by the embodiments of the present application, in order to facilitate understanding of the embodiments of the present application, the present application first specifically describes the problems existing in the related art: Traditional code plagiarism detection methods, such as those based on text matching or simple hashing, primarily focus on the surface form of the code. Advanced "skin-changing" plagiarism, however, involves replacing variable and function names, fine-tuning the code's logical structure, and rewriting textual expressions. This creates surface differences from the original code, while preserving the core content and structure. This makes it difficult for traditional methods to accurately identify this type of plagiarism, making it unable to meet the growing demand for code copyright protection.

[0028] Furthermore, during the software development process, it's often difficult to definitively prove the creation time and originality of code. Developers may face the possibility of their code being plagiarized, but lacking effective time stamps and unalterable evidence, they are unable to definitively prove their code's originality and creation time. This puts developers at a disadvantage in copyright disputes, hindering the motivation and rights protection of code creators.

[0029] In view of the above research findings of the inventors, the embodiments of the present application provide a system and method for generating software code ownership authentication certificates, and a verification system and method for software code ownership authentication certificates, which can solve the technical problems existing in the related art that advanced "skin-changing" plagiarism of software codes is difficult to identify and code ownership certificates are easily tampered with.

[0030] Figure 1 This is a structural block diagram of a system for generating software code authentication credentials provided in an embodiment of the present application. Figure 1 As shown, the system 10 for generating software code ownership authentication credentials may include an AI semantic analysis module 101, a semantic fingerprint generation module 102, a structural feature extraction module 103, a structural fingerprint generation module 104, a core package construction module 105, a private key signature module 106, a trusted timestamp acquisition module 107 and a ownership authentication module 108.

[0031] The AI ​​semantic analysis module 101 can be used to perform semantic analysis on the content of the target code using a pre-trained AI code model, and generate a high-dimensional semantic embedding vector based on the extracted features.

[0032] The target code is the code that needs to be authenticated. It can be any code or any code fragment, and this application does not limit this. The AI ​​code model can be a code language model (CLM), such as a BERT-type CLM or a language-optimized CLM. The AI ​​code model utilizes a large amount of code data during pre-training to learn the syntax, semantics, and structure of the code, and generates a high-dimensional semantic embedding vector based on the extracted features.

[0033] For example, in some embodiments, the AI ​​semantic analysis module 101 can be specifically used to use the AI ​​code large model to identify multiple elements such as variable names, function names, operators, etc. in the target code, and perform feature extraction on the multiple elements; and output a high-dimensional semantic embedding vector by performing weighted summation or nonlinear transformation operations on the features of the multiple elements. Among them, the high-dimensional semantic embedding vector can be a multidimensional array. The high-dimensional semantic embedding vector is usually an array composed of real numbers, and its dimensions can be dozens or even thousands of dimensions. For example, in a common code semantic analysis scenario, the generated high-dimensional semantic embedding vector may be an array of length 512, and each element is a floating point number, such as [0.123, -0.456, 0.789, ..., 0.234]. The above is only an example and does not constitute a limitation of this application.

[0034] When generating a high-dimensional semantic embedding vector, the AI ​​semantic analysis module 101 can input the target code into a pre-trained AI code large model. The AI ​​code large model can perform an in-depth analysis of the syntax, semantics and other information of the target code, identify multiple elements such as variable names, function names, operators in the target code, and perform feature extraction on the multiple elements. For multiple elements such as variable names, function names, operators in the target code, the AI ​​code large model will map these elements to a high-dimensional space based on the patterns it has learned from a large amount of code data. Then, by performing a comprehensive calculation of the features of these elements in the target code, such as weighted summation, nonlinear transformation and other operations, a high-dimensional semantic embedding vector representing the semantics of this section of target code is finally output.

[0035] The semantic fingerprint generation module 102 can be used to perform a hash calculation on the high-dimensional semantic embedding vector to obtain a binary semantic fingerprint of the target code. For example, in some embodiments, the semantic fingerprint generation module 102 can convert the high-dimensional semantic embedding vector into a fixed-length binary semantic fingerprint SF using a locality-sensitive hashing algorithm (such as the SimHash algorithm). The core value of the binary semantic fingerprint SF lies in its ability to identify code snippets that are similar in functionality or logical intent but may differ in their specific implementation (e.g., variable names, comments, and code structure).

[0036] The structural feature extraction module 103 may be configured to construct an abstract syntax tree of the target code using a static code analysis tool, and extract multiple structural features of the target code from the abstract syntax tree.

[0037] The Abstract Syntax Tree (AST) is a tree structure that graphically displays the grammatical structure of the target code. The tree's nodes represent various syntactic elements in the code, such as function definitions, variable declarations, and expressions, while the edges represent the hierarchical relationships between these elements. For example, for a simple C language code like "int main() { int a = 5; return a;}," the root node of its abstract syntax tree might be the "Function Definition" node. It has two child nodes: a "Variable Declaration" node (corresponding to "int a = 5;") and a "Return Statement" node (corresponding to "return a;"). The "Variable Declaration" node, in turn, has child nodes representing the variable type "int," the variable name "a," and the assignment expression.

[0038] Specifically, the structural feature extraction module 103 can be used as a static code analysis tool to construct an abstract syntax tree of the target code. First, the lexical analyzer will decompose the target code into lexical units. Then, the syntax analyzer will combine these lexical units into a grammatical structure according to the grammatical rules of the programming language and construct an abstract syntax tree. During the construction process, each syntax element will be gradually added to the tree as a node according to the grammatical hierarchy of the code, and a parent-child relationship will be established between them. Next, multiple structural features of the target code can be extracted from the abstract syntax tree. Exemplarily, the multiple structural features include but are not limited to cyclomatic complexity, the topological summary of the function call graph, the feature vector of the module dependency graph, the frequency of occurrence of structural patterns of specific design patterns, etc.

[0039] Cyclomatic complexity represents the number of independent paths in the code, reflecting the complexity of the code's logical structure. The higher the cyclomatic complexity, the more complex the code is, and the more difficult it is to understand, test, and maintain.

[0040] The topological summary of the function call graph can be used to represent the calling relationships between functions in a program. Through the topological summary of the function call graph, you can quickly understand the calling patterns and overall architecture of the functions in the code.

[0041] A module dependency graph describes the dependencies between modules in the target code. Nodes represent modules, and edges indicate the direction of dependencies between modules. The eigenvector of a module dependency graph is a data structure used to quantify and represent the characteristics of the graph. It maps various structural information and attributes of the graph into a vector, where each element corresponds to a specific feature of the graph. By analyzing the eigenvector of the module dependency graph, we can identify key modules that significantly impact other modules.

[0042] The structural pattern frequency of a specific design pattern indicates how many times the structural pattern of the specific design pattern is used in the target code. By counting the structural pattern frequency, we can understand the use of design patterns in the target code and assess the target code's reliance on specific design patterns.

[0043] The structural fingerprint generation module 104 can be used to permutate and combine the multiple extracted structural features and then perform a hash calculation to obtain the structural fingerprint (STF) of the target code. The structural fingerprint (STF) primarily characterizes the organizational structure and morphological complexity of the target code. To ensure the reproducibility of the structural fingerprint, the permutations and combinations follow preset rules. For example, the extracted structural features are arranged in lexicographic order by their feature names, followed by normalization and weighted concatenation to ensure unique output for the same code.

[0044] The core package construction module 105 can be used to receive the binary semantic fingerprint SF of the target code, the structural fingerprint STF of the target code, the unique identifier of the target code, the developer digital identity of the target code and key submission metadata, and construct the code DNA core package based on the received data.

[0045] The unique identifier of the target code is the unique identifier of the target code in the code repository, such as the Git Commit ID (H1). The Git Commit ID (H1) represents the unique identifier corresponding to the most recently committed target code in the code repository. It is a string generated by a hash algorithm that uniquely identifies the content and status of a code commit. Developer digital identities include but are not limited to PKI (Public Key Infrastructure)-based certificates (Distinguished Names) and / or Decentralized Identifiers (DIDs). Key commit metadata includes but is not limited to the commit summary and project ID.

[0046] The core package construction module 105 can be used to construct the binary semantic fingerprint SF, structural fingerprint STF, unique identifier, developer digital identity and key submission metadata of the target code into a standardized data structure (such as JSON format or XML format), which is called the code DNA core package.

[0047] The private key signature module 106 can be used to encrypt the hash value of the code DNA core package using the developer's private key (such as a private key based on the SM2 algorithm, a private key based on the RSA algorithm, or a private key based on the ECDSA algorithm) and use the encrypted hash value as the developer signature Developer_Signature. The Developer_Signature ensures the integrity of the core package and the non-repudiation of the developer's identity.

[0048] The trusted timestamp acquisition module 107 can be used to perform hash calculation on the code DNA core package and the developer signature Developer_Signature as a whole to obtain the timestamp authentication hash value Hash_For_TSA.

[0049] For example, in some embodiments, when calculating Hash_For_TSA, the code DNA core package and the developer signature can be input as a whole into a hash algorithm (such as the SHA-256 hash algorithm). The hash algorithm will perform a series of bit operations and transformations on it, ultimately outputting a fixed-length timestamp authentication hash value Hash_For_TSA. This timestamp authentication hash value Hash_For_TSA can be used for subsequent trusted timestamp authentication operations to ensure data integrity and non-tampering.

[0050] The trusted timestamp acquisition module 107 may also be configured to send the timestamp authentication hash value Hash_For_TSA to a trusted timestamp service platform (TSA), and receive a trusted timestamp certificate of the target code returned by the trusted timestamp service platform.

[0051] The title authentication module 108 can be used to use the code DNA core package, developer signature and trusted timestamp certificate together as the title authentication certificate of the target code. That is, the title authentication module 108 can use the complete "code DNA core package + Developer_Signature + trusted timestamp certificate" as a "code DNA title authentication certificate" for the target code. Among them, the trusted timestamp certificate can be a certification certificate in PDF format, embedded as a PDF attachment, or displayed in the form of a string. In this application, the title authentication certificate is a "trinity" complete evidence chain certificate, that is, the title authentication certificate is a logically indivisible whole composed of three elements: code DNA core package (content), developer signature (identity) and trusted timestamp (time). Without any part, a complete and reliable chain of evidence of rights cannot be formed.

[0052] The system for generating software code ownership authentication credentials provided by the embodiments of the present application analyzes the code from multiple dimensions through the "semantic + structure" dual fingerprint technology. The semantic fingerprint deeply understands the functional semantics of the code, and the structural fingerprint grasps the structural characteristics of the code. This multi-dimensional analysis method can comprehensively and meticulously characterize the essential characteristics of the code. Compared with traditional methods, it can more accurately identify various forms of plagiarism, especially advanced "skin-changing" plagiarism. For example, in actual applications, for some plagiarized codes that have undergone complex disguises, traditional hash methods may be misjudged as original, but dual fingerprint technology can accurately judge the nature of plagiarism through in-depth analysis of semantics and structure, greatly improving the accuracy and reliability of plagiarism detection.

[0053] Furthermore, the Trusted Timestamp Service (TSA) platform is used to timestamp the "Code DNA Core Package," which contains dual code fingerprints, developer signatures, and key metadata. The timestamp provided by the TSA is authoritative and tamper-proof, accurately recording the state of the code at a specific point in time. In the event of a copyright dispute, this timestamp serves as strong evidence, ensuring the undeniable originality of the code and providing developers with reliable copyright protection.

[0054] In order to enable the structural fingerprint STF of the target code to better reflect the organizational structure and morphological complexity of the target code, according to some embodiments of the present application, optionally, the structural fingerprint generation module 104 can be specifically used to arrange multiple structural features according to a preset arrangement order; then normalize the arranged multiple structural features to map the characteristic values ​​of multiple structural features in different ranges to a unified range; perform weighted splicing and combination processing on the normalized multiple structural features or combine them through a small hash network to obtain a new feature vector through combination; perform hash calculation on the new feature vector to obtain the structural fingerprint STF of the target code.

[0055] Specifically, different structural features may be disordered in their original state, and disordered features are not conducive to subsequent unified processing and analysis. First, the structural fingerprint generation module 104 can be used to arrange the multiple structural features extracted from the abstract syntax tree according to a preset arrangement order. The arrangement order can be flexibly adjusted according to actual conditions, and this application does not limit this. Then, the structural fingerprint generation module 104 can be used to normalize the multiple structural features after arrangement to map the feature values ​​of multiple structural features in different ranges to a unified range. In this way, each feature can have the same weight basis in the subsequent combination process, avoiding the excessive impact of certain features on the results due to the large value range, so that the final generated structural fingerprint can more accurately reflect the structural features of the target code without causing deviations due to the value range problem of a certain feature.

[0056] Then, the structural fingerprint generation module 104 can be used to perform weighted concatenation and combination processing on the normalized multiple structural features or perform combination processing through a small hash network to obtain a new feature vector through the combination.

[0057] Specifically, for multiple normalized structural feature vectors, a corresponding weight can be assigned to each structural feature vector. For example, assume there are three structural feature vectors, namely structural feature vector A, structural feature vector B, and structural feature vector C. For example, the weight assigned to structural feature vector A is 0.6, the weight assigned to structural feature vector B is 0.3, and the weight assigned to structural feature vector C is 0.1. The weighted concatenation process is to first multiply each element in structural feature vector A by its weight, each element in structural feature vector B by its weight, and each element in structural feature vector C by its weight, and then concatenate the results in sequence to obtain a new feature vector.

[0058] Different structural features may have different importance for describing the structure of the code. Through weighting, different weights can be assigned to the features according to their importance, so that important features can be more prominently reflected in the combination results, and the generated new feature vector can more accurately reflect the essential characteristics of the code structure, thereby improving the quality and representativeness of the structural fingerprint.

[0059] Small hash networks, on the other hand, can map multiple features into a new feature space. This allows for compression and transformation of features, extracting more representative feature combinations and generating new feature vectors. Small hash networks are relatively simple and computationally inefficient, reducing feature dimensionality while preserving key information and improving computational efficiency.

[0060] Based on the same technical concept as the system 10 for generating software code ownership authentication credentials provided in the above embodiment, the present application also provides a verification system for software code ownership authentication credentials. The verification system can be used to verify the authenticity of software code ownership authentication credentials, such as verifying forged software code ownership authentication credentials, and can also verify software code ownership authentication credentials generated by the system 10 for generating software code ownership authentication credentials.

[0061] Figure 2 This is a structural diagram of a verification system for software code authentication credentials provided in an embodiment of the present application. Figure 2 As shown, the verification system 20 for software code authentication credentials may include an authentication credential acquisition module 201 , a query module 202 , a code DNA core package verification module 203 , a developer signature verification module 204 and a trusted timestamp certificate verification module 205 .

[0062] The authentication certificate acquisition module 201 can be used to obtain the first authentication certificate of the software code to be verified. Specifically, for the sake of distinction, the authentication certificate of the software code to be verified is referred to as the first authentication certificate. The first authentication certificate can be submitted by the user or uploaded or retrieved by other means, which is not limited in this application. Among them, the first authentication certificate can include the first code DNA core package, the first developer signature and the first trusted timestamp certificate.

[0063] Query module 202 can be used to query the database for a confirmed second authentication credential corresponding to the first authentication credential. If so, it calls the Code DNA Core Package Verification Module 203, the Developer Signature Verification Module 204, and the Trusted Timestamp Certificate Verification Module 205 to further verify the first authentication credential. If not, it outputs a verification failure result.

[0064] Similarly, the second authentication certificate may include a second code DNA core package, a second developer signature, and a second trusted timestamp certificate.

[0065] The code DNA core package verification module 203 can be used to verify the consistency of the signature chain, time and hash value in the first code DNA core package based on the second code DNA core package and its associated public key.

[0066] For example, in some embodiments, the code DNA core package verification module 203 can be specifically used to start from the starting point of the signature chain in the first code DNA core package, and use the corresponding public key to verify each signature in turn until the entire signature chain is verified.

[0067] Specifically, a signature chain can include multiple signatures, which can be added sequentially. Starting from the beginning of the signature chain, each signature is verified using its corresponding public key. For example, if the first signature was signed by the developer using their private key, then this signature is verified using the developer's public key. If verification succeeds, the signature is valid, meaning that the data was not tampered with when it was signed.

[0068] Next, the next public key is used to verify the next signature in the signature chain until the entire signature chain is verified. If any signature fails to verify, the entire signature chain of the FirstCode DNA core package fails to verify, and the verification failure result is output.

[0069] For example, in some embodiments, the code DNA core package verification module 203 can also be used to compare the time points of different operations recorded in the first code DNA core package with the time points of different operations recorded in the second code DNA core package one by one.

[0070] Specifically, the first DNA core package can record the time points of multiple different operations. For example, the signature time should be after the data is generated, and the subsequent operation time should be after the signature time. If the time points of different operations recorded by the first DNA core package are out of order, or differ from the time points of different operations recorded by the second DNA core package, verification fails and a verification failure result is output.

[0071] For example, in some embodiments, the code DNA core package verification module 203 may also be used to compare multiple hash values ​​in the first code DNA core package with multiple hash values ​​in the second code DNA core package one by one, wherein the multiple hash values ​​include at least a binary semantic fingerprint SF and a structural fingerprint STF.

[0072] Specifically, the hash value carried by or recalculated in the first code DNA core package can be compared one by one with multiple hash values ​​in the second code DNA core package, such as comparing the binary semantic fingerprint SF of the first code DNA core package with the binary semantic fingerprint SF of the second code DNA core package, and comparing the structural fingerprint STF of the first code DNA core package with the structural fingerprint STF of the second code DNA core package. For example, if any one or more hash values ​​are inconsistent, the verification fails and a verification failure result is output.

[0073] The developer signature verification module 204 can be used to decrypt the first developer signature using the developer's public key and compare the decrypted result with the hash value of the second code DNA core package. If the two are consistent, it means that the first developer signature is valid, otherwise the verification fails.

[0074] The trusted timestamp certificate verification module 205 may be configured to send the first trusted timestamp certificate to a trusted timestamp service platform (TSA) and receive a verification result of the first trusted timestamp certificate returned by the trusted timestamp service platform. The verification result of the first trusted timestamp certificate may include verification success or verification failure.

[0075] The embodiment of the present application provides a verification system for software code authentication credentials. The code DNA core package verification module verifies the consistency of the signature chain, time and hash value. The signature chain verification ensures the authenticity and integrity of the signature during the code circulation process. The time comparison ensures the rationality of the code operation sequence and time. The hash value comparison quickly determines whether the core content of the code is consistent, thereby improving the accuracy of the verification. The developer signature verification module further confirms the reliability of the code source by decrypting the signature and comparing it with the code hash value, preventing the code from being published in the name of the developer. The trusted timestamp certificate verification module uses an authoritative trusted timestamp service platform for verification, ensuring the authenticity and validity of the timestamp and enhancing the reliability of the entire verification system.

[0076] Figure 3 Another structural block diagram of the verification system for software code authentication credentials provided in the embodiment of the present application. Figure 3 As shown, according to some embodiments of the present application, optionally, the verification system 20 for software code ownership authentication credentials may further include a fingerprint similarity analysis module 301 .

[0077] The fingerprint similarity analysis module 301 can be used to calculate a first similarity score between the binary semantic fingerprint in the first code DNA core package and the binary semantic fingerprint in the second code DNA core package, calculate a second similarity score between the structural fingerprint in the first code DNA core package and the structural fingerprint in the second code DNA core package, and perform a weighted sum of the first similarity score and the second similarity score according to preset weights to obtain a target similarity score. The target similarity score is used for version evolution tracing or copyright attribution determination.

[0078] Specifically, the similarity score can be calculated from multiple dimensions, taking into account the semantic and structural similarity of the code. For example, for semantic similarity, the fingerprint similarity analysis module 301 can be used to calculate the first similarity between the binary semantic fingerprint SF in the first code DNA core package and the binary semantic fingerprint SF in the second code DNA core package, such as the Hamming distance or vector cosine similarity of the two binary semantic fingerprints SF. Then, based on the first similarity, a corresponding score is assigned, which is called the first similarity score or the semantic similarity score.

[0079] For example, for structural similarity, a second similarity between the structural fingerprint STF in the first code DNA core package and the structural fingerprint STF in the second code DNA core package can be calculated, such as the graph edit distance or feature vector distance between the two structural fingerprint STFs. Then, based on the second similarity, a corresponding score is assigned, which is called a second similarity score or a structural similarity score.

[0080] Then the first similarity score (i.e. semantic similarity score) and the second similarity score (i.e. structural similarity score) are weighted and summed according to their respective preset weights to obtain the final target similarity score. For example, if the weight of the first similarity score is 0.6 and the weight of the second similarity score is 0.4, then the target similarity score = the first similarity score 0.6 + Second Similarity Rating 0.4.

[0081] This target similarity score is used for version evolution tracing or copyright determination. For example, if the target similarity score is high, plagiarism is likely present, excluding reasonable coincidences. This helps determine code copyright ownership and protect the rights of original developers. Furthermore, by comparing the similarity between the binary semantic fingerprint (SF) and the structural fingerprint (STF), the code evolution path and trends can be understood. This helps development teams understand the code's evolutionary history, analyze which parts have undergone significant changes and which have remained relatively stable, and thus better facilitate code maintenance and subsequent development.

[0082] Accordingly, if Figure 3 As shown, in some embodiments, the verification system 20 for software code ownership authentication credentials may further include an evolution tracing module 302. The evolution tracing module 302 may be used to construct an evolutionary relationship of the software codes of the first code DNA core package and the second code DNA core package based on the code version information of the first code DNA core package and the second code DNA core package, and the timestamp information of the first trusted timestamp certificate and the second trusted timestamp certificate when there are differences between the first code DNA core package and the second code DNA core package and the target similarity score is greater than a preset score, and input the evolutionary relationship into a visualization tool. The visualized evolutionary relationship and the target similarity score can be used as an objective quantitative basis or indicator with technical credibility to determine the degree of code plagiarism or borrowing, and provide technical support for litigation, judicial appraisal and arbitration.

[0083] Specifically, when there are differences between the first code DNA core package and the second code DNA core package, it means that the software code corresponding to the first code DNA core package and the software code corresponding to the second code DNA core package are not exactly the same. However, when the target similarity score is greater than the preset score, it means that the similarity between the software code corresponding to the first code DNA core package and the software code corresponding to the second code DNA core package is high. At this time, the code version information of the first code DNA core package and the second code DNA core package can be extracted, and based on the timestamp information of the first trusted timestamp certificate and the second trusted timestamp certificate, the software codes corresponding to the first code DNA core package and the software codes corresponding to the second code DNA core package of different versions can be arranged in chronological order. For example, determine which version is the earliest and which is the subsequent updated version based on the timestamp.

[0084] Next, the binary semantic fingerprint SF and the structural fingerprint STF can be used to establish an association between the two versions of the software code. For example, if two different versions of the software code have similar structural fingerprints, it means that there may be an evolutionary relationship between them. The second similarity between the structural fingerprints can be used to determine the closeness of this relationship and construct the evolutionary relationship of the software code of the first code DNA core package and the second code DNA core package. For example, from a temporal perspective, the software code corresponding to the second code DNA core package is in front, and the software code corresponding to the first code DNA core package is in the back. From the code structure point of view, the first code DNA core package of version A has modified and replaced the relevant algorithm code in the second code DNA core package of version B, for example, it has also adjusted some module structures related to data processing and storage. These changes are reflected in the hash value and structural fingerprint of the code, so that the first code DNA core package of version A and the second code DNA core package of version B differ in these aspects.

[0085] Next, the constructed evolutionary relationship of the software code can be input into a visualization tool. For example, version information, software code identifiers, and the relationship between the two can be provided to the visualization tool in a pre-defined data format (such as JSON). Through the interactive features of the visualization tool, users can more intuitively understand the trusted evolution process of the software code between different versions.

[0086] Based on the same technical concept as the system 10 for generating software code ownership authentication credentials provided in the above embodiment, the present application also provides a method for generating software code ownership authentication credentials. The method can be implemented based on the system 10 for generating software code ownership authentication credentials provided in the above embodiment.

[0087] Figure 4A flow chart of a method for generating software code authentication credentials provided in an embodiment of the present application. Figure 4 As shown, the method for generating a software code authentication certificate may include the following steps: S401: Use the pre-trained AI code model to perform semantic analysis on the content of the target code and generate a high-dimensional semantic embedding vector based on the extracted features; S402: Perform hash calculation on the high-dimensional semantic embedding vector to obtain a binary semantic fingerprint of the target code; S403: Using a static code analysis tool to construct an abstract syntax tree of the target code, and extracting multiple structural features of the target code from the abstract syntax tree; S404: Arrange and combine the extracted multiple structural features, and then perform hash calculations on them to obtain a structural fingerprint of the target code; S405: Receive the binary semantic fingerprint, structural fingerprint, unique identifier, developer digital identity and key submission metadata of the target code, and construct a code DNA core package based on the received data; S406: Encrypt the hash value of the code DNA core package using the developer's private key, and use the encrypted hash value as the developer's signature; S407: Hash the code DNA core package and the developer signature as a whole to obtain a timestamp authentication hash value; and send the timestamp authentication hash value to the trusted timestamp service platform, and receive the trusted timestamp certificate returned by the platform; S408: The code DNA core package, developer signature and trusted timestamp certificate are used together as the target code’s authentication credentials.

[0088] The specific process of the above steps has been described in detail above and will not be repeated here.

[0089] Figure 4 Each step in the method shown has the functions of implementing the various modules / units of the system 10 for generating software code ownership authentication credentials provided in the above embodiment, and can achieve its corresponding technical effects. For the sake of brevity, it will not be repeated here.

[0090] Based on the same technical concept as the verification system 20 for software code ownership authentication credentials provided in the above embodiment, this application also provides a verification method for software code ownership authentication credentials. This method can be implemented based on the verification system 20 for software code ownership authentication credentials provided in the above embodiment.

[0091] Figure 5 A flow chart of a method for verifying software code rights authentication credentials provided in an embodiment of the present application. Figure 5As shown, the verification method for software code authentication credentials may include the following steps: S501: Obtain a first authentication certificate for the software code to be verified, where the first authentication certificate includes a first code DNA core package, a first developer signature, and a first trusted timestamp certificate; S502: Query the database to see whether there is a second confirmed authentication certificate corresponding to the first authentication certificate; if so, call the subsequent module to further verify the first authentication certificate; if not, output verification failure; the second authentication certificate includes the second code DNA core package, the second developer signature, and the second trusted timestamp certificate; S503: Based on the second code DNA core package and its associated public key, perform consistency verification on the signature chain, time and hash value in the first code DNA core package; S504: Decrypt the first developer's signature using the developer's public key, and compare the decrypted result with the hash value of the second code DNA core package; S505: Send the first trusted timestamp certificate to the trusted timestamp service platform, and receive the verification result of the first trusted timestamp certificate returned by the platform.

[0092] The specific process of the above steps has been described in detail above and will not be repeated here.

[0093] Figure 5 Each step in the method shown has the functions of implementing the various modules / units of the verification system 20 for software code authentication credentials provided in the above embodiment, and can achieve its corresponding technical effects. For the sake of brevity, they are not repeated here.

[0094] Based on the method for generating a software code ownership authentication credential or the verification method for a software code ownership authentication credential provided in the above embodiments, the present application also provides an electronic device.

[0095] The electronic device in the embodiment of the present application can be a user terminal device, a server, other computing devices, or a cloud server. Figure 6 This is a hardware structure diagram of an electronic device according to an embodiment of the present application. The electronic device may include a processor 601 and a memory 602 storing computer program instructions. When the processor 601 executes the computer program instructions, the process or function of any of the above-mentioned embodiments is implemented.

[0096] Specifically, processor 601 may include a central processing unit (CPU) or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application. Memory 602 may include a large-capacity memory for data or instructions. For example, memory 602 may be at least one of the following: a hard disk drive (HDD), read-only memory (ROM), random access memory (RAM), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, a magnetic tape, a universal serial bus (USB) drive, or other physical / tangible memory storage device. For another example, memory 602 may include removable or non-removable (or fixed) media. For another example, memory 602 may be internal or external to the integrated gateway disaster recovery device. Memory 602 may be non-volatile solid-state memory. In other words, memory 602 typically includes a tangible (non-transitory) computer-readable storage medium (such as a memory device) encoded with computer-executable instructions, and when the software is executed (e.g., by one or more processors), the operations described in the method of the embodiments of the present application can be performed. The processor 601 implements the process or function of any method in the above embodiments by reading and executing computer program instructions stored in the memory 602.

[0097] In one example, Figure 6 The electronic device shown may also include a communication interface 603 and a bus 610. The processor 601, memory 602, and communication interface 603 are connected via bus 610 and communicate with each other. The communication interface 603 is primarily used to enable communication between the various modules, devices, units, and / or equipment in the embodiments of the present application. Bus 610, which may comprise hardware, software, or both, couples the components of the online data traffic metering device. For example, the bus may include at least one of the following: an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industrial Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses. Bus 610 may include one or more buses. Although the embodiments of the present application describe or illustrate a specific bus, the embodiments of the present application may consider any suitable bus or interconnection method.

[0098] In combination with the method in the above embodiments, an embodiment of the present application also provides a computer-readable storage medium, which stores computer program instructions. When the computer program instructions are executed by a processor, they implement the process or function of any method in the above embodiments.

[0099] In addition, an embodiment of the present application further provides a computer program product, which stores computer program instructions. When the computer program instructions are executed by a processor, the process or function of any one of the methods in the above embodiments is implemented.

[0100] The flowcharts and / or block diagrams of the methods, devices, systems and computer program products of the embodiments of the present application are described above by way of example, and various aspects thereof are described. It should be understood that each box in the flowchart and / or block diagram or a combination thereof may be implemented by computer program instructions, or may be implemented by dedicated hardware that performs a specified function or action, or may be implemented by a combination of dedicated hardware and computer instructions. For example, these computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to form a machine that enables these instructions executed by such a processor to enable the implementation of the functions / actions specified in each box in the flowchart and / or block diagram or a combination thereof. Such a processor may be a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit.

[0101] The functional blocks shown in the structural block diagrams of the embodiments of the present application can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc.; when implemented in software, they are programs or code segments used to perform the required tasks. The programs or code segments can be stored in a memory or transmitted over a transmission medium or communication link via a data signal carried in a carrier wave. The code segments can be downloaded via a computer network such as the Internet or an intranet.

[0102] It should be noted that the present application is not limited to the specific configurations and processes described above or shown in the figures. The above is only a specific implementation method of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the described system, device, module or unit can refer to the corresponding process in the method embodiment without further description. It should be understood that the scope of protection of the present application is not limited to this. Any technician familiar with the technical field can think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the scope of protection of the present application.

Claims

1. A system for generating software code authentication credentials, characterized in that: include: AI semantic analysis module, which uses a pre-trained AI code model to perform semantic analysis on the target code content and generate high-dimensional semantic embedding vectors based on the extracted features; Semantic fingerprint generation module, used to perform hash calculation on high-dimensional semantic embedding vectors to obtain the binary semantic fingerprint of the target code; A structural feature extraction module is used to construct an abstract syntax tree of the target code using a static code analysis tool, and extract multiple structural features of the target code from the abstract syntax tree; The structural fingerprint generation module is used to arrange and combine the extracted multiple structural features and then perform hash calculations on them to obtain the structural fingerprint of the target code; A core package building module is used to receive the binary semantic fingerprint, structural fingerprint, unique identifier, developer digital identity and key submission metadata of the target code, and construct a code DNA core package based on the received data; The private key signature module is used to encrypt the hash value of the code DNA core package using the developer's private key and use the encrypted hash value as the developer's signature; The trusted timestamp acquisition module is used to hash the code DNA core package and the developer's signature as a whole to obtain the timestamp authentication hash value; The timestamp authentication hash value is sent to the trusted timestamp service platform, and the trusted timestamp certificate is returned. The ownership authentication module is used to use the code DNA core package, developer signature and trusted timestamp certificate as the ownership authentication credentials of the target code.

2. The system according to claim 1, wherein: The AI ​​semantic analysis module is specifically used to use the AI ​​code large model to identify multiple elements such as variable names, function names, and operators in the target code, and to extract features of multiple elements; and to output a high-dimensional semantic embedding vector by performing weighted summation or nonlinear transformation operations on the features of multiple elements, where the high-dimensional semantic embedding vector is a multidimensional array.

3. The system according to claim 1, wherein: Multiple structural features of the target code include cyclomatic complexity, topological summary of the function call graph, feature vectors of the module dependency graph, and structural pattern occurrence frequency of specific design patterns.

4. The system according to claim 1, wherein: The structural fingerprint generation module is specifically used to arrange multiple structural features according to a preset arrangement order; then normalize the arranged multiple structural features to map the characteristic values ​​of multiple structural features in different ranges to a unified range; perform weighted splicing and combination processing on the normalized multiple structural features or combine them through a small hash network to obtain a new feature vector through the combination; perform hash calculation on the new feature vector to obtain the structural fingerprint of the target code.

5. A verification system for software code authentication credentials, characterized in that: The software code ownership authentication credential includes a forged software code ownership authentication credential or a software code ownership authentication credential generated by the system for generating a software code ownership authentication credential according to any one of claims 1 to 4, wherein the verification system includes: A rights authentication certificate acquisition module is used to obtain a first rights authentication certificate for the software code to be verified, the first rights authentication certificate including a first code DNA core package, a first developer signature, and a first trusted timestamp certificate; A query module is configured to query a database for a second, confirmed authentication certificate corresponding to the first authentication certificate; if so, to call a subsequent module to further verify the first authentication certificate; if not, to output a verification failure; the second authentication certificate includes a second code DNA core package, a second developer signature, and a second trusted timestamp certificate; A code DNA core package verification module, configured to verify the consistency of the signature chain, time, and hash value in the first code DNA core package based on the second code DNA core package and its associated public key; A developer signature verification module, configured to decrypt the first developer signature using the developer's public key and compare the decrypted result with the hash value of the second code DNA core package; The trusted timestamp certificate verification module is used to send the first trusted timestamp certificate to the trusted timestamp service platform and receive the verification result of the first trusted timestamp certificate returned by the trusted timestamp service platform.

6. The verification system according to claim 5, characterized in that: The Code DNA core package verification module is specifically used to verify each signature in sequence using the corresponding public key starting from the starting point of the signature chain until the entire signature chain is verified; the time points of different operations recorded in the first Code DNA core package are compared with the time points of different operations recorded in the second Code DNA core package one by one; Multiple hash values ​​in the first code DNA core package are compared one by one with multiple hash values ​​in the second code DNA core package, wherein the multiple hash values ​​at least include binary semantic fingerprints and structural fingerprints.

7. The verification system according to claim 5, characterized in that: The verification system further comprises: The fingerprint similarity analysis module is used to calculate a first similarity score between the binary semantic fingerprint in the first code DNA core package and the binary semantic fingerprint in the second code DNA core package, calculate a second similarity score between the structural fingerprint in the first code DNA core package and the structural fingerprint in the second code DNA core package, and perform weighted summation of the first similarity score and the second similarity score according to preset weights to obtain a target similarity score, which is used for version evolution tracing or copyright ownership determination.

8. The verification system according to claim 7, characterized in that: The verification system further comprises: The evolution tracing module is used to construct the evolution relationship of the software codes of the first code DNA core package and the second code DNA core package based on the code version information of the first code DNA core package and the second code DNA core package and the timestamp information of the first trusted timestamp certificate and the second trusted timestamp certificate when there are differences between the first code DNA core package and the second code DNA core package and the target similarity score is greater than a preset score, and input the evolution relationship into the visualization tool.

9. A method for generating a software code authentication certificate, characterized in that: The method is implemented based on the system for generating software code ownership authentication credentials according to any one of claims 1 to 4, and includes: Use the pre-trained AI code model to perform semantic analysis on the target code content and generate a high-dimensional semantic embedding vector based on the extracted features; Hash the high-dimensional semantic embedding vector to obtain the binary semantic fingerprint of the target code; Using static code analysis tools to construct an abstract syntax tree of the target code and extract multiple structural features of the target code from the abstract syntax tree; The extracted multiple structural features are arranged and combined, and then hashed to obtain the structural fingerprint of the target code; Receive the binary semantic fingerprint, structural fingerprint, unique identifier, developer digital identity and key submission metadata of the target code, and construct the code DNA core package based on the received data; Use the developer's private key to encrypt the hash value of the CodeDNA core package, and use the encrypted hash value as the developer's signature; Hash the CodeDNA core package and the developer's signature as a whole to obtain a timestamp authentication hash value; then send the timestamp authentication hash value to the trusted timestamp service platform and receive the trusted timestamp certificate returned by it; The code DNA core package, developer signature and trusted timestamp certificate are used together as the target code’s authentication credentials.

10. A method for verifying software code ownership authentication credentials, characterized in that: The method is implemented based on the verification system for software code ownership authentication credentials according to any one of claims 5 to 8, and includes: Obtain the first authentication certificate for the software code to be verified, which includes the first code DNA core package, the first developer's signature, and the first trusted timestamp certificate; Query the database to see if there is a second confirmed authentication certificate corresponding to the first authentication certificate; if so, call the subsequent module to further verify the first authentication certificate; if not, output verification failure; the second authentication certificate includes the second code DNA core package, the second developer signature, and the second trusted timestamp certificate; Based on the second code DNA core package and its associated public key, verify the consistency of the signature chain, time and hash value in the first code DNA core package; Decrypt the first developer's signature using the developer's public key and compare the decrypted result with the hash value of the second code DNA core package; The first trusted timestamp certificate is sent to the trusted timestamp service platform, and a verification result of the first trusted timestamp certificate returned by the platform is received.

Citation Information

Patent Citations

  • Source code multi-tag graph neural network-based program code copying type detection method and system

    CN108446540A

  • Data right confirmation method and system based on block chain technology

    CN112651052A

  • Software copyright management system based on smart contract

    CN119903491A

  • Model weight confirmation method and device based on block chain and model fingerprint

    CN120372704A

  • KR20210082885A