A decentralized open-source software trusted tokenization protocol

By analyzing the structural and semantic information of open-source software in a trusted execution environment, non-fungible tokens are generated, which solves the problem of the lack of incentive mechanisms in open-source projects and realizes objective evaluation of the impact of commit records and developer incentives.

CN116956236BActive Publication Date: 2026-07-17THE BLOCKHOUSE TECHNOLOGY LIMITED +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
THE BLOCKHOUSE TECHNOLOGY LIMITED
Filing Date
2022-04-13
Publication Date
2026-07-17

Smart Images

  • Figure CN116956236B_ABST
    Figure CN116956236B_ABST
Patent Text Reader

Abstract

This invention relates to a decentralized open-source software protocol for global open-source software and developers. The protocol is based on non-fungible tokens (NFTs), which are linked to commit records of collaboratively developed source code, such as open-source software code. Through this protocol, contributors to open-source software code are incentivized to undertake clearly defined tasks and then claim their contributions in the form of NFTs. This invention proposes a Transparent Trust Center (TC) as its core technology to achieve trusted software analysis, thereby extracting essential information and value from given commit records in a software repository.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the generation of nonfungible tokens associated with commit records of software source code, particularly but not limited to the generation of nonfungible tokens relating to source code developed collaboratively, such as open-source software code. Background Technology

[0002] In the field of software development, collaboration among developers on joint software projects is a common practice, often using the open-source model. As those skilled in the art will understand, open-source software, as the name suggests, is software whose source code can be freely modified and redistributed. In other words, it's a model that allows for decentralized software development, using a developer community to collaborate on further developing the software.

[0003] Typically, developers undertake these types of projects for a variety of reasons, with a few examples including a general passion for the project's goals, academic purposes, or to enhance their own skills as developers. However, applicants recognize that mechanisms to incentivize developers to contribute to open-source projects are few (if any), and most projects fail to attract and reach developers globally.

[0004] according to According to the 2020 Digital Insight Report, 99.95% of developers are inactive, and 71.21% of open-source projects are supported by fewer than 10 developers. For a time, the open-source project "OpenSSL" was maintained by only a single active developer.

[0005] Developers who contribute to these open-source projects may provide code that has a significant impact on the project. For example, changes to the source code supplied by a specific developer (called a "commit") may add particularly important new features to the project, or may fix specific bugs or critical security vulnerabilities.

[0006] The applicant recognizes that it would be beneficial to provide a mechanism for determining the value associated with specific code and properly attributing it to one or more relevant developers. Summary of the Invention

[0007] From a first perspective, embodiments of the present invention provide a method for operating a trusted execution environment to generate non-fungible tokens associated with commit records of source code, the method comprising:

[0008] The trusted execution environment is supplied with a first source code and a second source code, wherein the second source code is the result of the first source code undergoing the commit record;

[0009] Parse the first source code and the second source code to generate a first structure representation and a second structure representation accordingly;

[0010] Compile the first source code and the second source code to generate a first executable code and a second executable code accordingly;

[0011] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0012] Analyze the first structural representation and the second structural representation, as well as the first control flow graph and the second control flow graph, to determine the structural information associated with the submission record;

[0013] Symbolic execution is performed on the first control flow graph and the second control flow graph to generate corresponding first semantic representations and second semantic representations;

[0014] Analyze the first semantic representation and the second semantic representation to determine the semantic information associated with the submission record;

[0015] Generate non-fungible tokens; and

[0016] The non-fungible token is associated with the structural information and semantic information associated with the submission record.

[0017] This first aspect of the invention extends to a trusted execution environment configured to generate non-fungible tokens associated with commit records to source code, the trusted execution environment being configured to:

[0018] Receive a first source code and a second source code, wherein the second source code is the result of the first source code undergoing the commit record;

[0019] Parse the first source code and the second source code to generate a first structure representation and a second structure representation accordingly;

[0020] Compile the first source code and the second source code to generate a first executable code and a second executable code accordingly;

[0021] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0022] Analyze the first structural representation and the second structural representation, as well as the first control flow graph and the second control flow graph, to determine the structural information associated with the submission record;

[0023] Symbolic execution is performed on the first control flow graph and the second control flow graph to generate corresponding first semantic representations and second semantic representations;

[0024] Analyze the first semantic representation and the second semantic representation to determine the semantic information associated with the submission record;

[0025] Generate non-fungible tokens; and

[0026] The non-fungible token is associated with the structural information and semantic information associated with the submission record.

[0027] The first aspect of the invention also extends to a non-transitory computer-readable medium containing instructions that, when executed by a processor, cause the processor to perform a method of operating a trusted execution environment to generate a non-fungible token associated with a commit record to source code, the method comprising:

[0028] The trusted execution environment is supplied with a first source code and a second source code, wherein the second source code is the result of the first source code undergoing the commit record;

[0029] Parse the first source code and the second source code to generate a first structure representation and a second structure representation accordingly;

[0030] Compile the first source code and the second source code to generate a first executable code and a second executable code accordingly;

[0031] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0032] Analyze the first structural representation and the second structural representation, as well as the first control flow graph and the second control flow graph, to determine the structural information associated with the submission record;

[0033] Symbolic execution is performed on the first control flow graph and the second control flow graph to generate corresponding first semantic representations and second semantic representations;

[0034] Analyze the first semantic representation and the second semantic representation to determine the semantic information associated with the submission record;

[0035] Generate non-fungible tokens; and

[0036] The non-fungible token is associated with the structural information and semantic information associated with the submission record.

[0037] The first aspect of the invention also extends to a computer software product comprising instructions that, when executed by a processor, cause the processor to perform a method of operating a trusted execution environment to generate a non-fungible token associated with a commit record to source code, the method comprising:

[0038] The trusted execution environment is supplied with a first source code and a second source code, wherein the second source code is the result of the first source code undergoing the commit record;

[0039] Parse the first source code and the second source code to generate a first structure representation and a second structure representation accordingly;

[0040] Compile the first source code and the second source code to generate a first executable code and a second executable code accordingly;

[0041] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0042] Analyze the first structural representation and the second structural representation, as well as the first control flow graph and the second control flow graph, to determine the structural information associated with the submission record;

[0043] Symbolic execution is performed on the first control flow graph and the second control flow graph to generate corresponding first semantic representations and second semantic representations;

[0044] Analyze the first semantic representation and the second semantic representation to determine the semantic information associated with the submission record;

[0045] Generate non-fungible tokens; and

[0046] The non-fungible token is associated with the structural information and semantic information associated with the submission record.

[0047] Therefore, it should be understood that embodiments of the present invention provide an arrangement in which a "trusted execution environment" (TEE) is used to analyze the impact of a commit record on software code by examining the semantic and structural content associated with a given commit record of software code, and subsequently generating a non-fungible token (NFT) associated with relevant information related to the semantic and structural content of the commit record. Thus, the NFT provides a "fingerprint" of the commit record c, thereby capturing the impact of the commit record c on the structural and / or semantic content of the relevant source code.

[0048] In other words, for a given code commit c: take the source code s before commit c and the source code s' after commit c as input (s, s', c) to the "transparency centre" (TC) service that provides a trusted computing environment based on TEE. This TC service can be set up on any cloud platform if needed.

[0049] Those skilled in the art will understand that a TEE is a secure processing arrangement that can form part of a larger processing arrangement or processor (such as a central processing unit (CPU)). Several different types of TEEs exist, as is known in the art itself, to which various aspects and embodiments of the present invention can be readily applied. Thus, a TEE is a computing environment capable of running general-purpose programs. It typically has its own memory, but this can be cautiously expanded by using external memory in a restricted and encrypted manner. It can securely authenticate itself to users, so that no user would reasonably believe they are using or viewing results from the TEE if it did not. This typically involves authentication and some form of key negotiation, as well as bootstrapping of a signature mechanism between the TEE and the user. Generally, a TEE will be configured to use only programs known and trusted by all parties. Furthermore, a TEE is typically capable of authenticating its configuration.

[0050] The use of a TEE-based transparent central service according to the present invention is highly advantageous because the TEE ensures the reliability of the process used to perform analysis. This is because, as outlined above, the code running on the TEE is securely stored and cannot be tampered with. Therefore, the structural and semantic analyses performed on commit records can be trusted to be objective. In particular, the use of a TEE-based transparent central service ensures the reliability of the process used to perform analysis because the code running on the TEE is securely stored and cannot be tampered with. Therefore, since code commit records undergo the same analysis as any other commit record input to this service, the analysis of commit records and thus the value provided by the developers of those commit records can be trusted to be objective.

[0051] The applicant recognizes that another benefit of using a TEE in the analysis process of this invention is that a TEE provides a high degree of flexibility to change the implementation of software analysis if needed. While cryptographic solutions can be used without a TEE, these solutions are typically implemented for specific analysis schemes and require complex implementations or redesigns if the software analysis is to be changed later. Furthermore, the use of a TEE supports confidential software analysis in which the confidentiality of the implementation scheme is desired is not disclosed. For example, if: t is the source code of the software analyzer; t' is an encrypted version of t; and x is an executable file of t, then the provider of t can send t' and x to a TEE-based transparent center to run the software analysis and prove that the results were generated by x associated with t (and t') without disclosing t to others.

[0052] Those skilled in the art will understand that the term "source code" is used to refer to low-level code, typically written in whole or in part by one or more humans, prior to any compilation process. This can be written in conventional programming languages ​​such as C; C++; C#; Java; Python; Visual Basic; JavaScript; R; SQL; or PHP, but it should be understood that this list is not exhaustive and that thousands of different programming languages ​​exist that are known in the art itself to to which the principles of the invention can be readily applied. The invention can also be used with proprietary programming languages ​​that are not widely used, provided that relevant analysis can be performed on the code.

[0053] The TC service performs a structural analysis process St to parse (s, s') into a pair of structural representations (t, t'). In some implementations, the pair of structural representations (t, t') may include a syntax tree (sometimes called an "abstract syntax tree"). Those skilled in the art will understand that the computation of these first and second structural representations (which, as described above, may include syntax trees) can be implemented via well-defined tree / graph algorithms or specific machine learning processes known in the art itself.

[0054] In at least some implementations, the step of analyzing the structural representation to determine the structural information associated with the submission record may include calculating a first structural code (M_structure_A). Therefore, this code M_structure_A can provide a proof of work for c based on (t, t').

[0055] As outlined above, the method includes compiling a first source code and a second source code (s, s') into corresponding executable code (b, b') as a low-level representation of (s, s'), and constructing a control flow graph (g, g') for the two executable codes. Therefore, in some potentially overlapping implementations, the step of analyzing the first and second control flow graphs to determine structural information associated with a commit record may include calculating a second structural code (M_structure_B). Thus, this encoded M_structure_B can provide a proof of work for c based on (g, g').

[0056] In some implementations, the process of generating structural information from an abstract syntax tree (or similar) and a control flow graph may involve, as appropriate, running a systematic traversal on a given graph structure (i.e., an abstract syntax tree, which is a type of graph) and / or a control flow graph. During this process, all nodes in the graph are visited according to a depth-first search (DFS) procedure. Using this procedure, the structural information is updated based on the content contained in that node each time it is visited. The process then continues to visit the next node based on DFS. This process continues until all nodes in the graph have been covered, at which point the process is complete.

[0057] Symbolic execution steps can work in a similar manner by systematically accessing the control flow graph (CFG), for example, using Depth-First Search (DFS). According to this implementation, before execution begins, the process assigns symbolic values ​​(e.g., X, Y, Z) to all variables used in the CFG, rather than concrete values. Then, all nodes and paths in the CFG can be symbolically executed one after another. Specifically, the execution process symbolically executes all instruction information within each basic block of control flow based on the virtual machine standard and symbolic context, updates the symbolic state of the software analysis process, and adds relevant semantic information such as the semantic graph structure. Exploring a complete path in the CFG yields a set of symbolic values ​​and their expressions.

[0058] In some implementations, a satisfiable modulo theory (SMT) solver can be used to check if an expression is solvable; that is, to find at least one set of concrete values ​​for all symbolic variables so that the expression evaluates to true. If so, the path is feasible because it can be triggered under specific conditions. Unsolvable paths are ignored, and the process continues until all paths have been explored. The SMT solver can be part of the same module that performs symbolic execution of the CFG, or it can be a separate module.

[0059] In a particular set of implementations, the structural information (M_structure) can therefore be a combination of the two structural codes M_structure_A and M_structure_B.

[0060] In some cases, commit records may be provided along with a commit log (i.e., an overview of what the commit record provides). This can be in the form of a written description and / or a set of options that indicate certain attributes of the commit record. For example, a commit log may indicate that a commit record introduced one or more new features and / or that a commit record fixed a specific bug in the software. In some implementations, the method further includes extracting data from the commit log associated with the commit record. This optional semantic analysis process may be referred to as "Se-a". In one set of such implementations, the step of extracting data from the commit log associated with the commit record may include calculating a first semantic code (M_semantic_A).

[0061] As summarized above, the method of the present invention includes performing a semantic analysis process (referred to as "Se-b") on a first control flow graph and a second control flow graph (g, g'). In some embodiments, the step of analyzing the first and second semantic representations includes computing a second semantic code M_semantic_B. This code M_semantic_B can be an encoding that provides proof-of-work for c via a specific form of formal verification technique, such as capturing and vectorizing semantic updates to the source code by symbolically traversing the two graphs. Generally, it should be understood that the term "semantic update" refers to logical modifications that do not take into account the structural form of the source code (e.g., including but not limited to conditional procedures, value dependencies, etc.).

[0062] In a particular set of implementations, semantic information (M_semantics) can therefore be a combination of the two structures encoding M_semantics_A and M_semantics_B.

[0063] Based on the presence of various encodings in a given set of implementations, a TEE can provide combinations of various encodings as output. In a specific set of implementations, the output of the TEE includes structural information M_structure and semantic information M_semantics.

[0064] In some implementations, the structural information M_structure and the semantic information M_semantics can each comprise corresponding numerical vectors. Each value in these numerical vectors can be "accountable" or "non-accountable." An accountable value means that the value is directly related to an attribute that can be independently interpreted and generated using vectorization techniques known in the art itself. Non-accountable values ​​are related to the vector as a whole, rather than to any individual value; these non-accountable values ​​can be generated by machine learning or deep learning models. Thus, the encoded or numerical vectors provide an objective abstraction of the software submission record.

[0065] An arrangement is envisioned in which a numerical score for a given software commit can be generated by applying a specific formula to the code to produce a numerical score, although the chosen formula will make the numerical score a subjective measure rather than an objective measure given by the code itself. Therefore, in general, the impact of a particular commit is confirmed through community consensus later in the software's development process or usage, rather than at the time of the commit (it is considered influential if most developers and / or users react positively to the commit). Therefore, embodiments of the invention use structural information M_structure and semantic information M_semantics from the commit to generate unbiased abstractions. As a result, the objective measure provided by M_structure and M_semantics remains unchanged regardless of how the community's perception of the commit changes over time (e.g., from positive to negative, or vice versa).

[0066] The term "source code" should also be understood to extend to "intermediate language" code, such as LLVM IR. Those skilled in the art will understand that intermediate code typically provides an intermediate representation between source code (written in high-level programming languages ​​such as those listed above) and machine code for execution. According to embodiments of various aspects of the invention, analysis and verification techniques applicable to source code can also be executed on such intermediate code. Therefore, vendors are able to supply "source code" in such intermediate representation forms.

[0067] Therefore, as outlined above, in some embodiments, source code includes software code. However, in addition to being used to generate NFTs associated with commit records of the software source code, the applicant recognizes that the principles of the invention can also be applied to hardware. Those skilled in the art will understand that hardware description languages ​​(HDLs) can be used to define electronic circuits, particularly complex digital circuits, where a synthesizer (similar to a compiler used in software development) can transform an HDL description of the desired circuit behavior into a “netlist,” that is, a list of physical electronic components (typically from a predefined library of components) and their associated connections, which, once constructed into a physical circuit, will have the properties defined in the HDL description. The term “source code” as used herein should also be understood to include code written in HDL. Two commonly used HDLs are Verilog and VHDL, but these are merely exemplary and the principles of the invention apply to any such HDL. Therefore, in some embodiments, source code includes HDL code.

[0068] Those skilled in the art will also understand that the term "executable code"—as used in association with certain embodiments of the present invention—is used to refer to code that can be executed by a processor to perform one or more associated functions. Typically, executable code is derived from source code via a compilation process, resulting in a "binary file" (also known as "machine code" or "machine-readable code"). Although this is often in a form that is difficult for humans to understand, the term "executable code" is also extended to "executable source code," where human-readable code is executable. The term "executable code" is further extended to encompass "bytecode" (sometimes called "portable code" or "p-code"), which those skilled in the art will understand is a set of instructions designed to be executed by a software interpreter or used for further compilation into machine code.

[0069] The code provided by the vendor may undergo some obfuscation process. For example, the source code (or some intermediate code) may be obfuscated, making the code incomprehensible to humans, but it can still be compiled into an executable file that provides the same functionality as an executable file compiled from the unobfuscated source code, or can be executed in its obscured source code form.

[0070] However, it should be understood that there are no strict requirements regarding the identifiability or understandability of source code or executable code to humans and / or machines. Generally, however, source code and executable code can be in a form where the source code is understandable for the purpose of analysis performed within the TEE, while the executable code may be incomprehensible, or may be understandable to a lesser extent than the source code. Although executable code is generally not understood by humans, it should be understood that source code does not necessarily need to be understood by humans either, as long as the analysis performed within the TEE can be executed on that source code.

[0071] The principles of this invention can be applied to any software project in which commit records are made to update the source code. While this could be, for example, a software project where a single developer works on it, the invention is particularly advantageous in arrangements where multiple different users contribute to the source code (e.g., in collaborative open-source software projects). Thus, in some implementations, the source code can be edited by multiple users. The ability to generate NFTs associated with a developer's contribution to the project can encourage developer participation.

[0072] As outlined above, the TEE operates using source code prior to a commit record (“first” source code) and source code subsequent to a commit record (“second” source code). In some implementations, the first and second source codes are directly supplied to the trusted execution environment. However, in a set of alternative implementations, the first source code and the commit record are directly supplied to the trusted execution environment, and the method further includes generating the second source code by subjecting the first source code to a commit record.

[0073] In some implementations, the method further includes extracting one or more intent tags from the submission record and adding the one or more intent tags to semantic information associated with the submission record.

[0074] In some implementations, non-fungible tokens (NFTs) include structural and semantic information associated with the commit record. In other words, the structural and semantic information—or “metadata”—can be stored in the NFT’s data fields. This allows for easy access to the information by simply examining the NFT itself; however, this can have drawbacks, as it increases the storage required for the NFT. This is a particularly important consideration because storing metadata in the NFT generally increases the cost of storing the NFT on a blockchain, since such blockchain systems typically require payment for every bit stored on the blockchain.

[0075] Therefore, in some preferred embodiments, the method further includes storing structural and semantic information associated with the submission record in a database, the structural and semantic information being stored in contrast to identifiers associated with the non-fungible token. In other words, metadata can be stored in external storage, where the NFT provides pointers (i.e., identifiers) indicating where the metadata can be found. It should be noted that even if the metadata is placed in external storage, an attacker cannot forge the metadata because the process of generating the metadata is verifiable. An observer with input from a trusted source can use links in the NFT to find the metadata and run the process locally to generate a copy of the metadata, and then check whether the stored metadata is valid.

[0076] In some implementations, the method further includes: extracting data from one or more information fields associated with a commit record; and associating the data with a non-fungible token. In one set of such implementations, the one or more information fields include one or more of the following: user identity information; software repository information; and / or a timestamp. Therefore, the TC service may, as appropriate, perform an extraction process E to collect basic information M_basic from one or more source codes, commit records, and / or commit logs, which includes one or more of the following: the creator of c, the timestamp of c, the project software repository associated with c, etc. This basic information M_basic can be provided as the output of a TEE, and in one set of implementations is the output of a TEE along with the M_structure and / or M_semantics as outlined above.

[0077] It should be understood that the TEE may include suitable components or modules configured to implement the features of the present invention. One or more (and potentially all) of the various functions may be implemented by the same components or modules, and / or one or more (and potentially all) of these functions may be implemented by respective independent components or modules.

[0078] In some implementations, the trusted execution environment includes a receiving module configured to receive a first source code and a second source code.

[0079] In some implementations, the trusted execution environment includes a parser configured to parse the first source code and the second source code to generate a first structural representation and a second structural representation therefrom accordingly.

[0080] In some implementations, the trusted execution environment includes a compiler configured to compile the first source code and the second source code to generate, respectively, first executable code and second executable code.

[0081] In some implementations, the trusted execution environment includes a control flow graph generator configured to generate corresponding first and second control flow graphs from the first and second executable code.

[0082] In some implementations, the trusted execution environment includes a tree analyzer configured to analyze the first and second structural representations.

[0083] In some implementations, the trusted execution environment includes a graph analyzer configured to analyze the first control flow graph and the second control flow graph.

[0084] In some implementations, the trusted execution environment includes a symbolic executor configured to perform symbolic execution of a first control flow graph and a second control flow graph to generate corresponding first and second semantic representations.

[0085] In some implementations, the trusted execution environment includes a graph analyzer configured to analyze the first semantic representation and the second semantic representation to determine semantic information associated with the submission record.

[0086] In some implementations, the trusted execution environment includes a non-fungible token generator configured to generate non-fungible tokens and associate the non-fungible tokens with structural and semantic information associated with a submission record.

[0087] The applicant recognizes that the use of structural and semantic analysis is highly beneficial because it provides an objective and comprehensive overview of the impact of software commit history. Using both to analyze the functionality of software code (especially commit history) is useful because it is possible to have two pieces of software code, a and b, with similar structures but exhibiting very different functionalities. Conversely, a and b may also have exactly the same functionality but very different structures, which often occurs when developers refactor software (e.g., from a to b) to make the code more readable and / or easier to maintain. However, the applicant recognizes that in some scenarios, only structural or semantic analysis is necessary.

[0088] Therefore, from a second aspect, embodiments of the present invention provide a method for operating a trusted execution environment to generate non-fungible tokens associated with commit records of source code, the method comprising:

[0089] The trusted execution environment is supplied with a first source code and a second source code, wherein the second source code is the result of the first source code undergoing the commit record;

[0090] Parse the first source code and the second source code to generate a first structure representation and a second structure representation accordingly;

[0091] Compile the first source code and the second source code to generate a first executable code and a second executable code accordingly;

[0092] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0093] Analyze the first structural representation and the second structural representation, as well as the first control flow graph and the second control flow graph, to determine the structural information associated with the submission record;

[0094] Generate non-fungible tokens; and

[0095] The nonfungible token is associated with the structural information linked to the submission record.

[0096] This second aspect of the invention extends to a trusted execution environment configured to generate non-fungible tokens associated with commit records to source code, the trusted execution environment being configured to:

[0097] Receive a first source code and a second source code, wherein the second source code is the result of the first source code undergoing the commit record;

[0098] Parse the first source code and the second source code to generate a first structure representation and a second structure representation accordingly;

[0099] Compile the first source code and the second source code to generate a first executable code and a second executable code accordingly;

[0100] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0101] Analyze the first structural representation and the second structural representation, as well as the first control flow graph and the second control flow graph, to determine the structural information associated with the submission record;

[0102] Generate non-fungible tokens; and

[0103] The nonfungible token is associated with the structural information linked to the submission record.

[0104] A second aspect of the invention extends to a non-transitory computer-readable medium containing instructions that, when executed by a processor, cause the processor to perform a method of operating a trusted execution environment to generate a non-fungible token associated with a commit record to source code, the method comprising:

[0105] The trusted execution environment is supplied with a first source code and a second source code, wherein the second source code is the result of the first source code undergoing the commit record;

[0106] Parse the first source code and the second source code to generate a first structure representation and a second structure representation accordingly;

[0107] Compile the first source code and the second source code to generate a first executable code and a second executable code accordingly;

[0108] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0109] Analyze the first structural representation and the second structural representation, as well as the first control flow graph and the second control flow graph, to determine the structural information associated with the submission record;

[0110] Generate non-fungible tokens; and

[0111] The nonfungible token is associated with the structural information linked to the submission record.

[0112] A second aspect of the invention extends to a computer software product comprising instructions that, when executed by a processor, cause the processor to perform a method of operating a trusted execution environment to generate a non-fungible token associated with a commit record to source code, the method comprising:

[0113] The trusted execution environment is supplied with a first source code and a second source code, wherein the second source code is the result of the first source code undergoing the commit record;

[0114] Parse the first source code and the second source code to generate a first structure representation and a second structure representation accordingly;

[0115] Compile the first source code and the second source code to generate a first executable code and a second executable code accordingly;

[0116] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0117] Analyze the first structural representation and the second structural representation, as well as the first control flow graph and the second control flow graph, to determine the structural information associated with the submission record;

[0118] Generate non-fungible tokens; and

[0119] The nonfungible token is associated with the structural information linked to the submission record.

[0120] Alternatively, from a third aspect, embodiments of the present invention provide a method for operating a trusted execution environment to generate non-fungible tokens associated with commit records to source code, the method comprising:

[0121] The trusted execution environment is supplied with a first source code and a second source code, wherein the second source code is the result of the first source code undergoing the commit record;

[0122] Compile the first source code and the second source code to generate a first executable code and a second executable code accordingly;

[0123] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0124] Symbolic execution is performed on the first control flow graph and the second control flow graph to generate corresponding first semantic representations and second semantic representations;

[0125] Analyze the first semantic representation and the second semantic representation to determine the semantic information associated with the submission record;

[0126] Generate non-fungible tokens; and

[0127] The nonfungible token is associated with the semantic information linked to the submission record.

[0128] This third aspect of the invention extends to a trusted execution environment configured to generate non-fungible tokens associated with commit records to source code, the trusted execution environment being configured to:

[0129] Receive a first source code and a second source code, wherein the second source code is the result of the first source code undergoing the commit record;

[0130] Compile the first source code and the second source code to generate a first executable code and a second executable code accordingly;

[0131] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0132] Symbolic execution is performed on the first control flow graph and the second control flow graph to generate corresponding first semantic representations and second semantic representations;

[0133] Analyze the first semantic representation and the second semantic representation to determine the semantic information associated with the submission record;

[0134] Generate non-fungible tokens; and

[0135] The nonfungible token is associated with the semantic information linked to the submission record.

[0136] A third aspect of the invention extends to a non-transitory computer-readable medium containing instructions that, when executed by a processor, cause the processor to perform a method of operating a trusted execution environment to generate a non-fungible token associated with a commit record to source code, the method comprising:

[0137] The trusted execution environment is supplied with a first source code and a second source code, wherein the second source code is the result of the first source code undergoing the commit record;

[0138] Compile the first source code and the second source code to generate a first executable code and a second executable code accordingly;

[0139] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0140] Symbolic execution is performed on the first control flow graph and the second control flow graph to generate corresponding first semantic representations and second semantic representations;

[0141] Analyze the first semantic representation and the second semantic representation to determine the semantic information associated with the submission record;

[0142] Generate non-fungible tokens; and

[0143] The nonfungible token is associated with the semantic information linked to the submission record.

[0144] A third aspect of the invention extends to a computer software product comprising instructions that, when executed by a processor, cause the processor to perform a method of operating a trusted execution environment to generate a non-fungible token associated with a commit record to source code, the method comprising:

[0145] The trusted execution environment is supplied with a first source code and a second source code, wherein the second source code is the result of the first source code undergoing the commit record;

[0146] Compile the first source code and the second source code to generate a first executable code and a second executable code accordingly;

[0147] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0148] Symbolic execution is performed on the first control flow graph and the second control flow graph to generate corresponding first semantic representations and second semantic representations;

[0149] Analyze the first semantic representation and the second semantic representation to determine the semantic information associated with the submission record;

[0150] Generate non-fungible tokens; and

[0151] The nonfungible token is associated with the semantic information linked to the submission record.

[0152] The applicant also recognizes that when only semantic analysis is performed, the code can be supplied to the TEE only in executable (i.e., binary code) form. Therefore, from a fourth aspect, embodiments of the present invention provide a method for operating a trusted execution environment to generate non-fungible tokens associated with commit records to source code, the method comprising:

[0153] The trusted execution environment is supplied with first executable code and second executable code, wherein the first executable code and second executable code are compiled versions of the first source code and the second source code, and the second source code is the result of the first source code undergoing the commit record;

[0154] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0155] Symbolic execution is performed on the first control flow graph and the second control flow graph to generate corresponding first semantic representations and second semantic representations;

[0156] Analyze the first semantic representation and the second semantic representation to determine the semantic information associated with the submission record;

[0157] Generate non-fungible tokens; and

[0158] The nonfungible token is associated with the semantic information linked to the submission record.

[0159] This fourth aspect of the invention extends to a trusted execution environment configured to generate non-fungible tokens associated with commit records to source code, the trusted execution environment being configured to:

[0160] Receive first executable code and second executable code, wherein the first executable code and the second executable code are compiled versions of the first source code and the second source code, and the second source code is the result of the first source code undergoing the commit record;

[0161] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0162] Symbolic execution is performed on the first control flow graph and the second control flow graph to generate corresponding first semantic representations and second semantic representations;

[0163] Analyze the first semantic representation and the second semantic representation to determine the semantic information associated with the submission record;

[0164] Generate non-fungible tokens; and

[0165] The nonfungible token is associated with the semantic information linked to the submission record.

[0166] A fourth aspect of the invention extends to a non-transitory computer-readable medium containing instructions that, when executed by a processor, cause the processor to perform a method of operating a trusted execution environment to generate a non-fungible token associated with a commit record to source code, the method comprising:

[0167] The trusted execution environment is supplied with first executable code and second executable code, wherein the first executable code and second executable code are compiled versions of the first source code and the second source code, and the second source code is the result of the first source code undergoing the commit record;

[0168] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0169] Symbolic execution is performed on the first control flow graph and the second control flow graph to generate corresponding first semantic representations and second semantic representations;

[0170] Analyze the first semantic representation and the second semantic representation to determine the semantic information associated with the submission record;

[0171] Generate non-fungible tokens; and

[0172] The nonfungible token is associated with the semantic information linked to the submission record.

[0173] A fourth aspect of the invention extends to a computer software product comprising instructions that, when executed by a processor, cause the processor to perform a method of operating a trusted execution environment to generate a non-fungible token associated with a commit record to source code, the method comprising:

[0174] The trusted execution environment is supplied with first executable code and second executable code, wherein the first executable code and second executable code are compiled versions of the first source code and the second source code, and the second source code is the result of the first source code undergoing the commit record;

[0175] Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code;

[0176] Symbolic execution is performed on the first control flow graph and the second control flow graph to generate corresponding first semantic representations and second semantic representations;

[0177] Analyze the first semantic representation and the second semantic representation to determine the semantic information associated with the submission record;

[0178] Generate non-fungible tokens; and

[0179] The nonfungible token is associated with the semantic information linked to the submission record.

[0180] It should be understood that the optional features described above with respect to the first aspect of the invention are equally applicable to the second, third, and fourth aspects of the invention as appropriate. Attached Figure Description

[0181] Some embodiments of the invention will now be described with reference to the accompanying drawings, in which:

[0182] Figure 1 This is a diagram illustrating the commit history of software source code in a software repository;

[0183] Figure 2 This is a block diagram of a transparent center based on a Trusted Execution Environment (TEE) according to an embodiment of the present invention; and

[0184] Figure 3 This is a block diagram illustrating a scheme for extracting basic information, semantic information, and structural information associated with a submission record according to an embodiment of the present invention. Detailed Implementation

[0185] Figure 1 This is a diagram illustrating the commit history of software source code in a software repository. Figure 1 As can be seen, software repository 2 is used to store and record the source code of software projects, such as open source software projects, in which many users can contribute changes to the source code in the form of commit records (c)4.

[0186] Software repository 2 stores the current version of source code(s)6. When commit record 4 is submitted, software repository 2 updates the source code to a new version of source code(s')8, where s' differs from s in that the changes implemented in commit record c are implemented.

[0187] Figure 2 This is a block diagram of a Transparent Center (TC) based on a Trusted Execution Environment (TEE) according to an embodiment of the present invention. Figure 2 As can be seen, TEE 10 takes commit record 4, source code 6 before commit record, and source code 8 after commit record as input. As outlined in more detail below, TEE 10 provides TC service, which generates NFT 12 and associated metadata 14.

[0188] Figure 3This is a block diagram illustrating a scheme for extracting basic information, semantic information, and structural information associated with a submission record according to an embodiment of the present invention.

[0189] The TC service can perform an extraction process E to collect basic information M_basic from source code s, source code s', commit record c, and / or commit log. In this embodiment, the basic information M_basic includes the creator of c, the timestamp of c, the project software repository associated with c, etc. This basic information M_basic can be provided as the output of the TEE, and in one set of embodiments, it is the output of the TEE along with the M_ structure and / or M_ semantics as outlined above.

[0190] As outlined above, TEE 10 provides source code 6 before the commit and source code 8 after the commit. TEE 10 uses parser 16 to parse these source codes (s, s') 6 and 8 to generate a first abstract syntax tree t(s) and a second abstract syntax tree t'(s') accordingly, which are structural representations of the corresponding source codes 6 and 8.

[0191] The first abstract syntax tree t(s) and the second abstract syntax tree t'(s') are input into the tree analyzer 18. The tree analyzer 18 evaluates the tree t(s) and the tree t'(s'). The tree analyzer 18 uses well-defined tree / graph algorithms and / or specific machine learning procedures to analyze the contents of the syntax tree t(s) and the syntax tree t'(s').

[0192] It should be understood that such algorithms are generally known in the art. However, for ease of understanding, a brief overview of suitable algorithms is provided below. It should also be understood that any suitable algorithm can be used according to the principles of the invention.

[0193] In the case of graph algorithms, the graph is systematically traversed, and a predefined vector structure is refined on the fly. For example, a ternary structure with three attributes, each representing the number of graph nodes of a particular type, can be used. The vectorization process involves traversing the graph and counting the nodes for each type.

[0194] In the context of machine learning or deep learning, an encoding of a given graph, such as an AST or CFG, is generated. This encoding can be generated via well-designed algorithms known in the art, including but not limited to MDS, IsoMap, DeepWalk, and graph2vec. More specifically, such generation is a two-stage process: training and prediction. In the deep learning-based training phase, a set of graphs is fed into a training engine as a dataset. For each graph, a Weisfeiler-Lehman graph kernel is extracted to produce a set of subgraphs. All kernels are then used as input to a multi-layer neural network for training based on backpropagation and stochastic gradient descent techniques to assign numerical vectors to the graph and all its subgraphs. In the prediction phase, the deep learning model computes to generate numerical vectors for a given graph by partitioning the given graph into subgraphs and then mapping those subgraphs to subgraphs in the model. In an embodiment of the invention, given a pair of graphs, such as an AST or CFG, a pair of encodings from the model is generated as high-dimensional numerical vectors. The vector distance between the pair of vectors is then computed to indicate whether they are structurally and semantically close to each other. Therefore, a pair of triplets of the previous and subsequent graphs is obtained, that is, one is the "previous" code and the other is the "after" code, and their distance values ​​ranging from -1 to 1 are obtained (where -1 indicates far and 1 indicates near).

[0195] It should be understood that each abstract syntax tree t(s) and t'(s') provides a tree-like representation of the code, in which structural elements such as code sequences, conditional statements (e.g., "if" statements), and loops (e.g., "while" loops, "for" loops, etc.) are laid out. The tree analyzer 18 can examine the trees t(s) and t'(s') before and after commit record c to determine what changes (if any) the commit record c made to the syntactic structure of the source code.

[0196] TEE 10 then computes a structure code (M_structure_A) that provides a proof of work for c based on (t, t'). This structure code (M_structure_A) can be computed by the tree analyzer 18 itself, or it can be computed by another component of TEE 10 based on the output of the tree analyzer 18.

[0197] TEE 10 includes compiler 20, which compiles source code 6 and source code 8 to generate a first binary file b(s) and a second binary file b'(s') accordingly. TEE 10 also includes a control flow graph (CFG) builder 24, which generates a first CFG g(s) and a second CFG g'(s') from the binary files b(s) and b'(s') that generated the source code s before and after the commit. As described in more detail below, these first CFG g(s) and second CFG g'(s') are also used in the semantic analysis process performed by TEE 10.

[0198] Syntax tree t(s) and syntax tree t'(s') provide a structural representation of the code at a high level (i.e., in the form of source code, which is generally easier for humans to read), while CFG g(s) and CFG g'(s') provide a structural representation of the code at a low level (i.e., in the form of machine-executable code (or "binary")).

[0199] The graph analyzer 24 in TEE 10 takes CFG g(s) and CFG g'(s') as input and examines CFG g(s) and CFG g'(s') before and after commit record c to determine what changes (if any) the commit record c made to the structure of the executable version of the source code.

[0200] TEE 10 then calculates a structure code (M_structure_B), which provides a proof of work for c based on (g, g'). This structure code (M_structure_B) can be calculated by the graph analyzer 24 itself, or by another component of TEE 10 based on the output of the graph analyzer 24.

[0201] Subsequently, TEE 10 combines the two structure codes M_structure_A and M_structure_B to generate the structure information M_structure associated with the commit record c.

[0202] Submission record 4 is also input to intent extractor 25, which extracts one or more intent tags, encoded as M_semantic_A.

[0203] The first CFG g(s) and the second CFG g'(s') generated by the CFG builder 22 are also input to the symbol executor 26 of the TEE 10. The symbol executor 26 performs symbolic execution on the CFG g(s) and CFG g'(s') to generate the corresponding first semantic graph sg(s) and second semantic graph sg'(s'), which are then input to the further graph analyzer 28.

[0204] Symbolic execution of a given CFG, performed by symbolic executor 26, operates by systematically accessing all nodes in the CFG, for example, using Depth-First Search (DFS). Before execution begins, the process assigns symbolic values ​​(e.g., X, Y, Z) to all variables used in the CFG, rather than concrete values. Then, each node and path in the CFG is executed symbolically, one after another. Specifically, the execution process symbolically executes all instruction information within each control flow basic block based on the virtual machine standard and symbolic context, updates the symbolic state of the software analysis process, and adds relevant semantic information such as the semantic graph structure. After exploring a complete path in the CFG, a set of symbolic values ​​and their expressions is obtained. It should be understood that, in practice, graph analyzer 28 can be the same functional unit as graph analyzer 24 used in structural analysis, or it can be a separate functional unit.

[0205] Next, symbolic executor 26 uses a satisfiable module theory (SMT) solver to check if the expression is solvable, that is, whether at least one set of concrete values ​​of all symbolic variables used to generate the expression can be evaluated as true. If so, the path is considered feasible because it can be triggered under certain conditions. Unsolvable paths are ignored, and the process continues until all paths have been explored. In this embodiment, the SMT solver is a separate module integrated into symbolic executor 26, but it is understood that it could alternatively be integrated into another component (such as graph analyzer 28), or it could be a separate component. An exemplary SMT is derived from... Z3 is used, however this is not limiting, and other SMT solvers are available, and those skilled in the art can provide their own implementations.

[0206] The SMT solver provides the functionality to automatically solve decision problems using a set of logical formulas (e.g., expressions, in the context of this invention). The process involves first converting the set of logical formulas into a set of Boolean formulas (i.e., formulas with only Boolean variables that can take the value true or false). Then, some form of backtracking algorithm is run to determine the satisfiability of the set of formulas. Specifically, literals are systematically selected and assigned specific values. Based on these values, the expressions are simplified according to the underlying theory introduced in them, and then divided into subroutines. For each subroutine, its satisfiability is determined. This process is repeated recursively (i.e., given specific values, simplified, and divided) until a final determination is achieved for the entire set of formulas.

[0207] The output of graph analyzer 28 is encoded as M_semantic_B, which is combined with M_semantic_A from intent extractor 25 to generate semantic information M_semantic.

[0208] Review and Reference Figure 2 TEE 10 then generates NFT 12, which is recorded on a suitable blockchain, such as Ethereum. While NFT 12 can contain basic information M_basic, structural information M_structure, and semantic information M_semantic, recording data on a blockchain can be costly, especially given the large volume of data. To avoid this, TEE 10 generates NFT 12 with pointers to metadata 14, which contains basic information M_basic, structural information M_structure, and semantic information M_semantic, and this metadata 14 is stored outside the blockchain.

[0209] Therefore, those skilled in the art will understand that embodiments of the present invention provide an arrangement in which the structural and semantic content of software code can be determined, and corresponding NFTs (i.e., digital tokens) associated with that content can be generated. This allows the value of the developer's work to be captured and transformed into digital assets. The use of a transparent central service based on a TEE ensures that the process used to perform the analysis can be trustworthy, because the code running on the TEE is securely stored and cannot be tampered with. As a result, the analysis of code commit records and thus the value provided by the developer of those commit records can be trusted to be objective, because the commit records undergo the same analysis as any other commit record input to this service.

[0210] Although specific embodiments of the invention have been described in detail, those skilled in the art will understand that the detailed embodiments do not limit the scope of the claimed invention.

Claims

1. A method for operating a trusted execution environment to generate a non-fungible token associated with a record of commits to source code, the method comprising: The trusted execution environment is supplied with a first source code and a second source code, wherein the second source code is the result of the first source code undergoing the commit record; Parse the first source code and the second source code to generate a first structure representation and a second structure representation accordingly; Compile the first source code and the second source code to generate a first executable code and a second executable code accordingly; Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code; Analyze the first structural representation and the second structural representation, as well as the first control flow graph and the second control flow graph, to determine the structural information associated with the submission record; Symbolic execution is performed on the first control flow graph and the second control flow graph to generate corresponding first semantic representations and second semantic representations; Analyze the first semantic representation and the second semantic representation to determine the semantic information associated with the submission record; Generate non-fungible tokens; and The non-fungible token is associated with the structural information and semantic information associated with the submission record.

2. The method of claim 1, wherein the source code can be edited by multiple users.

3. The method according to claim 1 or 2, wherein the first source code and the second source code are directly supplied to the trusted execution environment.

4. The method of claim 1 or 2, wherein the first source code and the commit record are directly supplied to the trusted execution environment, and the method further comprises generating the second source code by subjecting the first source code to the commit record.

5. The method of claim 1, wherein the first structure representation and the second structure representation respectively include a first syntax tree and a second syntax tree.

6. The method of claim 1, wherein the step of analyzing the structural representation to determine structural information associated with the submission record includes calculating a first structural code, wherein the structural information includes the first structural code.

7. The method of claim 1, wherein the step of analyzing the first control flow graph and the second control flow graph to determine structural information associated with the submission record includes calculating a second structural code, wherein the structural information includes the second structural code.

8. The method of claim 1, further comprising extracting data from a commit log associated with the commit record.

9. The method of claim 8, wherein the step of extracting data from the commit log associated with the commit record includes calculating a first semantic code, wherein the semantic information includes the first semantic code.

10. The method of claim 1, wherein the step of analyzing the first semantic representation and the second semantic representation includes calculating a second semantic code, wherein the semantic information includes the second semantic code.

11. The method of claim 1, further comprising extracting one or more intent tags from the submission record, and adding the one or more intent tags to the semantic information associated with the submission record.

12. The method of claim 1, wherein the nonfungible token includes the structural information and the semantic information associated with the submission record.

13. The method of claim 1, further comprising storing the structural information and the semantic information associated with the submission record in a ledger, wherein the structural information and the semantic information are stored in contrast to an identifier associated with the nonfungible token.

14. The method of claim 13, wherein the ledger comprises a blockchain ledger.

15. The method according to claim 1, further comprising: Extract data from one or more information fields associated with the submitted record; as well as Associate the data with the non-fungible token.

16. The method of claim 15, wherein the one or more information fields include one or more of the following: user identity information; software repository information; and / or timestamp.

17. A trusted execution environment configured to generate non-fungible tokens associated with commit records to source code, said trusted execution environment being configured to: Receive a first source code and a second source code, wherein the second source code is the result of the first source code undergoing the commit record; Parse the first source code and the second source code to generate a first structure representation and a second structure representation accordingly; Compile the first source code and the second source code to generate a first executable code and a second executable code accordingly; Generate a corresponding first control flow graph and a second control flow graph from the first executable code and the second executable code; Analyze the first structural representation and the second structural representation, as well as the first control flow graph and the second control flow graph, to determine the structural information associated with the submission record; Symbolic execution is performed on the first control flow graph and the second control flow graph to generate corresponding first semantic representations and second semantic representations; Analyze the first semantic representation and the second semantic representation to determine the semantic information associated with the submission record; Generate non-fungible tokens; and The non-fungible token is associated with the structural information and semantic information associated with the submission record.

18. A non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to perform a method of an operational trusted execution environment according to any one of claims 1 to 16.