Software license-based code suggestion
By generating source code license home databases and filtering code suggestions based on these licenses, the problem of difficulty in considering licensing restrictions in development sessions in the prior art when providing code suggestions is solved, and code suggestions generation complying with license terms is achieved.
Patent Information
- Application Number
- CN202380069824.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2023-09-19
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art has difficulty considering licensing restrictions in development sessions when providing code suggestions, which may lead to developers violating the terms of the software license when integrating open source code.
By generating a source code license attribution database, determine the corresponding license for each source code file and filter and generate code suggestions based on these licenses to ensure that the suggested code complies with the licensing standards specified by the developer.
This implements the consideration of licensing restrictions when providing code suggestions, ensuring that developers comply with relevant license terms when using open source code, and avoid potential legal risks and technical issues.
Smart Images

Figure CN119968612A_ABST
Abstract
Description
Background Art
[0001] Open source software is considered a collaborative effort between different developers with different goals. Open source code is available for public inspection and review on public websites. Repositories host open source code for merging, forking, or pull requests from other developers who may want to integrate at least part of the code into their own projects. Typically, open source code is subject to a software license that specifies what a developer can do with the open source code written by another person.
[0002] Predictive code suggestions may provide suggested code snippets that may be based on existing open source code. Code suggestions may be based on existing source code files that are subject to their own license terms. If the source code originated from another open source project, code suggestions may not necessarily take into account restrictions that may apply to the current development session. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Figure 1 A data flow diagram of a system for a code suggestion service according to some embodiments is presented.
[0004] Figure 2 is a block diagram of a system for source code license attribution according to various embodiments.
[0005] Figure 3 A flow diagram of a system configured to attribute a license to source code files is presented in accordance with some embodiments.
[0006] Figure 4 is a logical block diagram illustrating a provider network that implements different services including code development services according to some embodiments.
[0007] Figure 5 is a logical block diagram illustrating code suggestion processing according to some embodiments.
[0008] Figure 6 is a logical block diagram illustrating code suggestions according to some embodiments.
[0009] Figure 7 is a logical block diagram illustrating an example interface of a development environment according to some embodiments.
[0010] Figure 8 is a flow chart of a method for providing code suggestions based on licensing criteria according to some embodiments.
[0011] Fig. 9 is a flow diagram of a method for attributing a license to a source code file according to some embodiments.
[0012] Fig.10is a flow chart of a method for parsing license format data to generate a source code attribution database according to some embodiments.
[0013] Fig.11 is a block diagram illustrating an example computing device for implementing the various techniques described herein, in accordance with some embodiments.
[0014] Although embodiments are described herein by way of example with respect to several embodiments and illustrative figures, those skilled in the art will recognize that embodiments are not limited to the described embodiments or figures. It should be understood that the drawings and detailed descriptions thereof are not intended to limit the embodiments to the specific forms disclosed, but on the contrary, are intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope defined by the appended claims. As used throughout this application, the word "may" is used in a permissive sense (i.e., meaning "possibly"), rather than a mandatory sense (i.e., meaning "must"). Similarly, the words "include, including, and includes" are meant to include but are not limited to.
[0015] "Based on." As used herein, this term is used to describe one or more factors that influence a determination. This term does not exclude additional factors that may influence the determination. That is, the determination may be based only on those factors or at least in part on those factors. Consider the phrase "A is determined based on B." While B may be a factor that influences the determination of A, this phrase does not exclude that the determination of A is also based on C. In other cases, A may be determined based only on B.
[0016] This specification includes references to "one embodiment" or "an embodiment." The appearances of the phrases "in one embodiment" or "in an embodiment" are not necessarily referring to the same embodiment. The particular features, structures, or characteristics may be combined in any suitable manner consistent with the present disclosure. DETAILED DESCRIPTION
[0017] Various systems and methods for generating code suggestions are described. Code suggestions can be provided to a developer during a development session. The development session can be conducted using an integrated development environment ("IDE"). The developer can enter code during a development session and request code suggestions from various external source code files. For example, the code suggestion can include a function or definition in a source code file to be merged into the development session.
[0018] Code suggestions may be generated based on a programming language suggestion model that is configured to apply a machine learning model to various source code files to determine portions of code to suggest. Source code files may be provided by one or more source code repositories hosted by other developers. Source code files may be subject to one or more software licenses that affect the rights and obligations of developers who integrate source code files into their own source code. For example, certain software licenses may require developers to make code open source, or restrict developers from obtaining certain intellectual property protections, such as patents. The code suggestion service may be configured to limit potential code suggestions by filtering out source code files that are subject to particular types of software licenses.
[0019] Source code files may be filtered based on one or more licensing criteria. For example, licensing criteria may indicate license terms or conditions that the developer intends to exclude from any source code files provided through the code suggestion service. As another example, licensing criteria may indicate that the developer does not want to give up its patent rights in any software developed using the imported source code files.
[0020] The source code license attribution service can generate a source code license attribution database. The source code license attribution database can associate source code files with their corresponding licenses. For example, a particular source code file can be indicated as being subject to a particular software license. The database can also include data indicating licensing criteria included in the license. For example, the database can indicate that a particular software license has various criteria that can be used to filter or restrict specific source code files from being included as code suggestion candidates for use by developers who choose to restrict based on licensing criteria.
[0021] In one aspect, a system is described. The system may include one or more computing devices that implement a code suggestion service. The code suggestion service may be configured to receive a request specifying one or more licensing criteria through an interface of the code suggestion service. The code suggestion service may be configured to determine a corresponding license for a corresponding source code file in a plurality of source code files based on a source code attribution database, the source code attribution database containing indications of corresponding licenses applicable to the plurality of source code files identified by parsing the plurality of source code files. The code suggestion service may be configured to generate a set of candidate code suggestions for a received code input based at least in part on the plurality of source code files. The code suggestion service may be configured to determine one or more code suggestions that meet the one or more licensing criteria based at least in part on the corresponding licenses for the corresponding source code files in the plurality of source code files based on the set of candidate code suggestions. The code suggestion service may be configured to provide one or more code suggestions that meet the one or more licensing criteria determined based on the set of candidate source code files.
[0022] In another aspect, a method is described. The method may include being performed by one or more computing devices implementing a code suggestion service. The method may include determining a corresponding license for a corresponding source code file in a plurality of source code files based on a source code attribution database, the source code attribution database containing an indication of the corresponding license for the corresponding source code file. The method may include generating a set of candidate code suggestions for a received code input based at least in part on the plurality of source code files. The method may include determining one or more code suggestions that satisfy one or more licensing criteria based on the set of candidate code suggestions. The method may include providing one or more code suggestions that satisfy one or more licensing criteria determined based on the set of candidate source code files.
[0023] In yet another aspect, one or more computer-readable storage media storing instructions are described. When executed on or across one or more processors, the instructions cause the one or more processors to perform operations. The operations may include determining a corresponding license for a corresponding source code file in a plurality of source code files based on a source code attribution database, the source code attribution database containing an indication of the corresponding license for the corresponding source code file. The operations may include generating a set of candidate code suggestions for a received code input based at least in part on the plurality of source code files. The operations may include determining one or more code suggestions that satisfy one or more licensing criteria based on the set of candidate code suggestions. The operations may include providing one or more code suggestions that satisfy one or more licensing criteria determined based on the set of candidate source code files.
[0024] Figure 1 A data flow diagram of a system for a code suggestion service 100 according to some embodiments is shown. According to various embodiments, the system 100 may include an integrated development environment ("IDE") 110, code suggestion generation 120, source code license attribution 130, and a license data repository 140. The system 100 may include various computing devices to implement a corresponding one of the following: IDE 110, code suggestion generation 120, source code license attribution 130, and license data repository 140. In some embodiments, the system 100 may be hosted as part of a provider network configured to provide computing-based services to clients.
[0025] According to various embodiments, the IDE 110 may be deployed at a computing device. For example, the IDE 110 may be implemented at a client computing device. As another example, the IDE 110 may be implemented as a service provided by a provider network, such as a web-based interface. According to some embodiments, the IDE 110 may receive input from a client, such as a software developer. For example, the input may include a code file input 102. As another example, the input may include text input, software source code, source code file, software library, content asset, etc. The code suggestion processing 112 may actively obtain and verify the code suggestion 116 before providing it to the display 104. In this way, a higher latency programming language suggestion model 122 implemented as part of the code suggestion generation 120 may be employed, the higher latency programming language suggestion model providing better and more usable code suggestions 116, even if its latency is longer than a smaller but less accurate model, because the active request can make the apparent latency of the code suggestion 0 or close to 0, while still ensuring by verification that the suggestion is still valid before display 104 (in view of the context of potential changes, such as other code file inputs 102).
[0026] According to some embodiments, IDE 110 can provide code suggestions to the client based on various criteria. For example, IDE 110 can provide code suggestions based on predicting upcoming text based on current text input from the client. In some embodiments, IDE 110 can send code suggestion request 114 to code suggestion generation 120. Request 114 can also include various criteria to limit or filter out potential candidate code suggestions. For example, a developer can request that the presented code suggestions exclude specified features related to their corresponding software license.
[0027] Source code files may be subject to a software license based on text provided by their respective developers. For example, a given source code file for a given application may indicate that the open source code is subject to a given open source license. For example, the open source license may include one or more of the Apache License, the BSD 3-clause License, the BSD 2-clause License, the GNU General Public License (GPL), the GNU Library or "Lesser" General Public License (LGPL), the MIT License, the Mozilla Public License, the Public Development and Distribution License, the Eclipse Public License, or any other applicable open source license. Different license types may impose different restrictions or obligations on any developer who adopts the open source code.
[0028] According to various embodiments, the developer of the code file input 102 can modify the request 114 to indicate that certain license types or features are to be included or excluded as part of the code suggestion generation 120. For example, the request 114 can include licensing criteria for excluding certain licenses from the code suggestions. Example licensing criteria can include intellectual property restrictions, open source requirements, monetization restrictions, copy restrictions, or other types of features related to open source licenses. In some cases, the developer may not want to apply a particular type of license to the code file input 102, so licensing criteria can be generated to exclude a particular type of license in the request 114 for code suggestions.
[0029] As another example, IDE 110 can provide code suggestions based on an external source code file from a source code currently being edited by a client. In some embodiments, code suggestion generation 120 can provide code suggestions to IDE 110. In some cases, the code suggestions can include the text of a function, variable, value, system call, or other text input that a client may enter. In some embodiments, the code suggestions can be based on code imported from one or more external source code repositories.
[0030] According to some embodiments, licensing criteria filtering 124 may filter candidate code suggestions by applying licensing criteria from request 114. According to some embodiments, candidate code suggestions may be determined based on source code and license information 136 provided by source code license attribution 130. Source code license attribution 130 may include code license attribution database builder 132. In some embodiments, code license attribution database builder 132 may generate code license attribution database 134. Code license attribution database 134 may include database records indicating corresponding licenses of corresponding source code files. For example, a given source code file may be indicated as having a given open source license, which will be applied to any application that employs the given source code file.
[0031] According to some embodiments, the code license attribution database builder 132 can generate the code license attribution database 132 based on the license information obtained from the license data repository 140. The license data repository 140 may include information related to various licenses that can be applied to various software. For example, the license data repository 140 may include a software license database 142, which is configured to include records indicating different software licenses. The software license database 142 may include license identifiers and license texts for various licenses. In some embodiments, the source code license attribution 130 can retrieve the license from the license data repository 140 and the software license database 142.
[0032] Code license affiliation database builder 132 may generate code license affiliation database 134 based in part on the licenses obtained from software license database 142. For example, code license affiliation database builder 132 may include information about the licenses, such as license terms and various criteria or characteristics about the licenses.
[0033] According to some embodiments, source code license attribution 130 may obtain source code files from a source code repository. Source code license attribution 130 may parse source code files to identify which licenses apply to the corresponding source code files. In some embodiments, text included in the source code files may indicate the licenses attributable to the source code files. For example, source code license attribution 130 may analyze the text in the source code files to determine the attributable licenses. Source code license attribution 130 may populate or generate a code license attribution database 134 based on the attributable licenses of the corresponding source code files.
[0034] According to some embodiments, source code license attribution 130 can provide source code and license information 136 to code suggestion generation 120. Code suggestion generation 120 can apply licensing criteria filtering 124 to generate a list of candidate code suggestions. In some embodiments, code suggestion generation 120 may include a programming language prediction model configured to determine a list of candidate code suggestions. For example, the programming language prediction model can apply one or more machine learning techniques to predict candidate code suggestions. In some embodiments, licensing criteria filtering 124 can be used to train the programming language prediction model to improve subsequent results of candidate code suggestions. The list of candidate code suggestions can exclude licenses that do not meet licensing criteria. Code Suggestions Code suggestions 116 can be provided to IDE 110. Code suggestion processing 112 can verify code suggestions 116 to display 104 code suggestions.
[0035] Figure 2 2 is a block diagram of a system 200 for source code license attribution according to various embodiments. The system 200 may include a license attribution service 202 and a code license attribution database builder 204. The system 200 may be implemented on or across one or more computing devices including one or more processors and memory storing instructions that cause the one or more processors to perform various operations.
[0036] According to some embodiments, license attribution service 202 may receive an indication of source code repository 210 to attribute a license to a source code file of source code repository 210. Source code parser 212 may parse the source code file to determine whether a portion of the source code file includes an indication of a license attributable to the source code file. In some embodiments, source code parser 212 may include comment extractor 214 to extract annotated code from the source code file for analysis. For example, an open source license may be added to the source code file as a text comment that does not directly affect the execution of the source code. According to some embodiments, the parsed text may be sent to text comparator 230 to be compared with the license text from license database 220.
[0037] According to some embodiments, the license attribution service 202 may also include a metadata parser 216. According to some embodiments, the metadata parser 216 may parse the metadata from the source code repository 210 to determine whether the metadata contains license-related information. For example, the metadata parser 216 may obtain the metadata to extract license-based information. According to some embodiments, the metadata parser 216 may send the metadata to the text comparator 230 to compare with the license text from the license database 220.
[0038] According to some embodiments, text comparator 230 may include string matching 232 configured to compare one or more strings of input text with strings contained in licenses identified in license database 220. For example, text extracted from a source code file may be compared to licenses to determine which licenses are included as part of the source code file. In some cases, the license text in the source code file may be different or different from known versions of the licenses in license database 220. For example, a developer who copied the license text from a secondary source may have slight changes in various terms or phrases in the license text.
[0039] According to some embodiments, string matching 232 can be based on similarity matching. For example, the input text can be provided to the text comparator 230 as N-grams. String matching 232 can generate a similarity score 234, which indicates the probability that the N-gram of the corresponding input text of the source code file may have a specific license. For example, a similarity score of 0.95 or higher can indicate a high probability that the input text represents a specific license. As another example, a similarity score of 0.90 to 0.95 can indicate a low probability that the input text represents a specific license.
[0040] According to some embodiments, based on the similarity score, license attribution service 202 can apply license attribution 236 to attribute a specific license to the corresponding source code file. License attribution 236 can associate a given source code file with a given license. License attribution 236 can register the given source code file with the license corresponding thereto in source code attribution database 242 stored in data storage device 240.
[0041] According to some embodiments, at least a portion of the source code attribution database 232 may be generated by the code license attribution database builder 204. For example, the code license attribution database builder 204 may generate a baseline version of the database 232 populated by the license attribution 236. According to some embodiments, the code license attribution database builder 204 may include a license format data collector 252. The license format data collector 252 may collect license format data 254 from the license data repository. For example, the license format data collector 252 may obtain the license format data 254 from the Software Package Data Exchange (SPDX) license list. The license format data 254 may include data structures of various software licenses that may be applied to open source software. The other source converter 250 may receive user input (such as input from a developer) to identify other licenses that are not necessarily provided from the license data repository. For example, a developer may decide to write his own license terms instead of adopting an existing license type. According to some embodiments, the database generator 256 may establish the source code attribution database 242 based on the license format data 254 and other license data provided by the other source converter 250.
[0042] Figure 3 A flow chart of a system 300 configured to attribute a license to a source code file according to some embodiments is shown. The system 300 may be implemented by one or more computing devices including one or more processors and memory. In some embodiments, the system 300 may include a processor and a processor. Figure 1 The source code license is attributable to 130 or Figure 2 The license attribution service 202 corresponds to the license attribution service.
[0043] Source code repository 302 may provide source code files to system 300. According to some embodiments, text from the source code files may be parsed into extracted text 304. Extracted text 304 may be processed in one of a variety of ways to determine the license attributed to the source code files.
[0044] At 310, the system 300 can determine a similarity score with a known license. In some embodiments, the similarity score can be expressed as a number from 0 to 1. At 312, based on the similarity score indicating a high probability of matching with a given license, the high probability match can be verified. In some embodiments, cluster licenses can be separated from non-cluster licenses. Cluster licenses can be verified using string-based matching to remove false positives. Based on a determination that a high probability match is verified, an attributed license can be assigned to the source code file 340. Based on a determination that a high probability match is not verified, verification can be retried as if the match is a low probability match.
[0045] At 314, based on the similarity score indicating a low probability of matching with a given license, a low probability match can be verified. In some embodiments, a unique string based match can be used to verify the cluster license. Based on the verification of the low probability match, the source code file can be assigned an attributed license 340.
[0046] At 320, the extracted text 304 may also be normalized. At 322, the normalized text may be parsed based on a keyword-based string match. Based on a string match with a given license, an attributable license 340 may be assigned to the source code file. At 330, the extracted text 304 may be hashed to generate a hashed N-gram. At 332, the hashed N-gram may be parsed based on a title text match. Based on a title text match with a given license, an attributable license 340 may be assigned to the source code file.
[0047] Figure 4 is a logical block diagram illustrating a provider network implementing different services, including code development services, according to some embodiments. Provider network 400 (which may be referred to as a "cloud provider network" or simply "cloud" in some embodiments) refers to a network-accessible pool of computing resources (such as computing, storage and networking resources, applications and services), which may be virtualized or bare metal. Provider network 400 can provide convenient, on-demand network access to a shared pool of configurable computing resources that can be programmatically provisioned and released in response to customer commands. These resources can be dynamically provisioned and reconfigured to accommodate variable loads.
[0048] The provider network 400 can be formed into multiple zones, where a zone is a separate geographic area in which a cloud provider clusters data centers. Each zone may include two or more availability zones connected to each other by a dedicated high-speed network (e.g., a fiber optic communication connection). An availability zone (also referred to as an availability domain, or simply "zone") refers to an isolated fault domain that includes one or more data center facilities, which have separate power, separate networking, and separate cooling separated from facilities in another availability zone. Preferably, the availability zones within a zone are far enough apart from each other so that the same natural disaster should not simultaneously take more than one availability zone offline. Customers can connect to the availability zones of the provider network 400 through a publicly accessible network (e.g., the Internet, a cellular communication network). The zone is connected to a global network that includes a dedicated networking infrastructure (e.g., a fiber optic connection controlled by a cloud provider) that connects each zone to at least one other zone. The provider network 400 can deliver content from a point of presence outside of these zones but networked with these zones by means of edge locations and regional edge cache servers. This division and geographic distribution of computing hardware enables the provider network 400 to provide customers with low-latency resource access with a high degree of fault tolerance and stability on a global scale.
[0049] As described above, the provider network 410 can implement various computing resources or services, such as code development services 410 and other services 430, which can be any other type of network-based services, including various other types of storage services (for example, database services or object storage services), computing services, data processing services, machine learning services, analysis services, communication services, event processing services, visualization services, and security services not shown.
[0050] In various embodiments, Figure 4 The components shown may be implemented directly in computer hardware, as instructions executable directly or indirectly by computer hardware (e.g., a microprocessor or computer system), or using a combination of these technologies. Figure 4 The components of can be implemented by a system including multiple computing nodes (or simply referred to as nodes), each of which can be similar to Fig.11 4. In various embodiments, the functionality of a given system or service component (e.g., a component of code development service 410) may be implemented by a specific node or may be distributed across several nodes. In some embodiments, a given node may implement the functionality of more than one service system component (e.g., more than one data storage component).
[0051] In some embodiments, code development service 410 may be implemented by provider network 400. Code development service 410 may implement various features to write code for different systems, applications, or devices, providing features for suggesting, identifying, reviewing, building, and deploying code. For example, code development service 410 may implement development environment 411. Code development environment 411 may provide various code entry tools (e.g., text-based, chart / graphic application development) to specify, call, or otherwise write (or cause to be written) code for different hardware or software applications.
[0052] Code development service 410 may implement code suggestion delivery 414, which may implement various computing resources to host and implement code suggestions 413 in a scalable manner to deliver on-demand code suggestions across a large number of clients using a highly driven machine learning model for high-quality code suggestion results. For example, code suggestion delivery 414 may implement workload balancing features and request management features to process and return code suggestions in a timely manner, thereby providing real-time code suggestions to code suggestion processing 420 (within or without provider network 400) with little or no apparent latency.
[0053] To avoid having the development environment wait to send multiple code suggestions in one communication, in some embodiments, code suggestion delivery 414 can implement a paging feature for code suggestions to allow multiple code suggestions to be delivered from the host or other computing resource implementing code suggestion 413 to the recipient development environments 419 and 411 over time through multiple communications. In this way, valid code suggestions can be made and presented, and then updated as more suggestions are received. Such techniques provide a simulated streaming experience without actually needing to support bidirectional streaming in the development environment. In this way, the benefits of fast delivery and updating of code suggestions can be provided without introducing additional requirements for the development environment, which may not necessarily be maintained by the provider network 400 operator.
[0054] To implement paging, code suggestions may be stored in the service 410 as they are generated and then returned in multiple exchanges by utilizing paging tokens accompanying requests for code suggestions to allow additional code suggestions to be retrieved from storage and sent back to the development environment 419 or 411.
[0055] In various embodiments, code suggestions 413 can generate code suggestions based on text input in development environment 411 or 419 (e.g., using a plug-in or other connection that can provide real-time analysis and suggestions of code as code is entered into development environment 411 or 419), as described below with respect to Figure 5Discussed in detail. Code suggestions 413 can use a generative model, a machine learning model, such as a generative pre-trained transformer (GPT), that is trained to generate code suggestions. Generative models are typically trained on large amounts of data for a specific task. In the case of generating code recommendations, this corpus (e.g., from a code suggestion code repository 415 or other code repository used to train the generative model) can include code repositories or snippets from various sources. Depending on the source code or owner, the code may be subject to certain licenses that need to be attributed during any use or reproduction. Since generative models can sometimes reproduce matches with the training data verbatim or nearly verbatim, it may also be necessary to provide metadata for attributing the original source as part of the suggestion. Code suggestion metadata (not shown) can provide the ability to provide metadata for code suggestions that can be provided.
[0056] Code development service 410 may implement (or access) code repository 415. Code repository 415 may store various code files, objects, or other code that may interact with various other features of code development service 410 (e.g., development environment 411 for writing, building, compiling, and / or testing code). In some embodiments, code repository 415 may implement various version and / or other access controls to track and / or maintain consistent versions of code collections for various development projects. In some embodiments, code repository may be stored or implemented external to provider network 400 (e.g., hosted in a private network or other location).
[0057] Code development service 410 may include code license attribution 417. Code license attribution 417 may determine the license attributable to the corresponding source code file, such as code repository 415. The attributable license may include an open source license that may be employed when the corresponding source code file is used in other projects.
[0058] Code development service 410 may implement interfaces to access and / or utilize various features of code development service 410. Such interfaces may include various types of interfaces, such as command line interfaces, graphical user interfaces, and / or programming interfaces (e.g., application programming interfaces (APIs)) to perform requested operations, including operations of development environment 411. An API refers to an interface and / or communication protocol between a client and a server such that if a client issues a request in a predefined format, the client should receive a response in a specific format or initiate a defined action. In the context of a cloud provider network, an API provides a gateway for clients to access cloud infrastructure by allowing them to obtain data from the cloud provider network or initiate actions within the cloud provider network, thereby enabling the development of applications that interact with resources and services hosted in the cloud provider network. An API may also enable different services of the cloud provider network to exchange data with each other.
[0059] In general, client 450 can cover any type of client, which can be configured to submit network-based requests to provider network 400 through network 460, including requests for services (e.g., requests for code searches or suggestions, etc.). For example, a given client 450 may include a suitable version of a web browser, or may include a plug-in module or other type of code module that can be executed as an extension of an execution environment provided by a web browser or executed within the execution environment. Alternatively, client 450 can cover applications (or their user interfaces), media applications, office applications, or any other applications that can utilize resources in provider network 400 to implement various applications. In some embodiments, such applications may include sufficient protocol support (e.g., for a suitable version of hypertext transfer protocol (HTTP)) for generating and processing network-based service requests without having to implement full browser support for all types of network-based data. That is, client 450 can be an application that interacts directly with provider network 400. In some embodiments, client 450 can generate network-based service requests based on a network-based service architecture of a representative state transfer (REST) style, a network-based service architecture based on documents or messages, or another suitable network-based service architecture.
[0060] In some embodiments, client 450 can provide access to provider network 400 to other applications in a transparent manner to these applications. For example, client 450 can be integrated with code development services 410. However, the operating system or file system can present different storage interfaces to the application program (such as the conventional file system hierarchy of files, directories and / or folders). In such embodiments, it may not be necessary to modify the application program to utilize the storage system service model. Instead, the details of docking with the data storage service can be coordinated by the client 450 and the operating system or file system on behalf of the application program executed in the operating system environment.
[0061] Client 450 can transmit network-based service request to provider network 400 through network 460, and receive response from the provider network. In various embodiments, network 460 can cover any suitable combination of networking hardware and protocol necessary for establishing network-based communication between client 450 and provider network 400. For example, network 460 can generally cover various telecommunication networks and service providers that implement the Internet together. Network 460 can also include private networks, such as local area network (LAN) or wide area network (WAN) and public or private wireless networks. For example, both given client 450 and provider network 400 can be provided in enterprises with their own internal networks respectively. In such embodiments, network 460 can include hardware (e.g., modem, router, switch, load balancer, proxy server, etc.) and software (e.g., protocol stack, accounting software, firewall / security software, etc.) necessary for establishing networking link between given client 450 and the Internet and between the Internet and provider network 400. It should be noted that in some embodiments, client 450 can communicate with provider network 400 using private network instead of public Internet.
[0062] In some embodiments, the provider network 400 may include the hardware (e.g., modems, routers, switches, load balancers, proxy servers, etc.) and software (e.g., protocol stacks, accounting software, firewall / security software, etc.) necessary to establish networking links between different components of the provider network 400, such as virtualization hosts, control plane components, and external networks 460 (e.g., the Internet). In some embodiments, the provider network 400 may employ Internet Protocol (IP) tunneling technology to provide an overlay network through which encapsulated packets may be tunneled through an internal network. The IP tunneling technology may provide a mapping and encapsulation system for creating an overlay network, and may provide separate namespaces for the overlay layer and the internal network layer. Packets in the overlay layer may be checked against a mapping directory to determine what their tunnel destination should be. The IP tunneling technology provides a virtual network topology; the interface presented to the client 450 may be attached to the overlay network so that when the client 450 provides the IP address to which it wants to send a packet, the IP address operates in the virtual space by communicating with a mapping service that knows where the IP overlay address is.
[0063] Perceived delays in code suggestions may reduce the utilization of code suggestions as a feature. For example, if a user must wait for a noticeable period of time after requesting code suggestions, the user workflow may be interrupted. To eliminate perceived delays, code suggestions may be pre-fetched. However, since the input context may have changed since the code suggestions were requested, techniques for validating proactively obtained code completion suggestions may be implemented, which ensure that the recommendations given are no longer consistent with the current state of the code. Figure 5 is a logical block diagram illustrating code suggestion processing according to some embodiments.
[0064] The code suggestion processing 420 can implement automatic suggestion event detection 510 that can evaluate keystrokes 542, elapsed time, special keys, or user-specific data to detect events. In various embodiments, this information can be maintained as part of a user-specific state that can be updated or reset when a code suggestion request is submitted. For example, keystrokes, elapsed time, or other measurements can be reset. Special keys can also be triggering events (and can also be evaluated in combination with other criteria (such as elapsed time)). For example, an event that triggers obtaining code completion suggestions can include entering "{""["" ("":" "ENTER key" or "TAB key") and an elapsed time threshold. In some embodiments, the automatic suggestion event detection 510 can use client-specific events, such as entering specific keys or characters in a client-specific mode (or configured / described by the client in the request to configure the suggestion processing 420).
[0065] Code suggestion request execution 520 may handle the formation, assembly, sending, and processing of responses from code suggestions 413, including sending requests 522 to obtain code completion suggestions and processing returned code suggestions 524. For example, code suggestion request execution 520 may obtain a context window of tokens (e.g., the N previous tokens before the cursor) from file state 540, as indicated at 544. In some embodiments, file and other context information as provided by file and other context extraction 550 may be sent.
[0066] File and other context extraction 550 can utilize different techniques to obtain file and other context information outside the context window (e.g., outside the N previous tokens). For example, file context can be obtained from the same file as the code suggestions being generated for inclusion therein. Information that can be obtained for file context can include boundaries in the current scope (e.g., code and comments limited by the current function to provide local context), class-level information (including class declarations, class constructors (e.g., __init__ functions)), and function-level information about all other public or protected methods defined by the class, function-level information, including all functions declared on both sides of the cursor in the current file. In some embodiments, signatures, docstrings, and return statements and / or variable-level information can be extracted, including all previous variable declarations that are visible to the current generation focus.
[0067] Other contexts that can be extracted at 550 can be intra-project contexts. In modern code development, classes and functions are typically defined in hierarchical files. Simple retrospective contexts do not include information outside the current file, which can create certain scenarios where machine learning models are unlikely to generate correct code. Because many code files use imported classes / functions / variables, adding this context can significantly improve code generation performance. Therefore, in some embodiments, an intra-project context can be added, where all classes, functions, and variables imported from the same project are used to obtain code completion suggestions.
[0068] Other context that can be extracted at 550 can be out-of-project context. Out-of-project context can refer to categories / functions / variables imported into the current file from other packages. This may affect the quality of suggestions when the imported packages are in a zero-shot setting (e.g., when the pre-trained model has no prior knowledge about the packages). Therefore, other context can be obtained by scanning out-of-project contexts for packages that are not included in the pre-trained data and including the corresponding categories / functions / variables as context in the request.
[0069] File and other context extraction 550 can perform regular expression based searches (e.g., searching for keywords such as "import") and extraction to obtain the various types of context discussed above. In some embodiments, parsing based extraction (e.g., by generating a symbol tree or other parsing graph of the code to obtain other context information) can be used.
[0070] The code suggestion request execution 520 can interact with the code suggestions provided in a paginated form. For example, a response 524 to a request for code suggestions can include a paginated token indicating that additional suggestions can be retrieved. The code suggestion request execution 520 can still continue to verify and provide suggestions through the code suggestion verification 530, while also submitting a subsequent request 522 with a paginated token to obtain additional code suggestion results, which can then be returned, verified, and provided. In this way, multiple suggestions can be made, allowing different execution times of the code suggestions to be generated, including potentially better code suggestions that may be provided when the user reviews the initial suggestion.
[0071] In some embodiments, file status 540 may provide information to various stages and may include both the code file and its associated metadata. File status 540 may also provide information for code suggestion verification 530, such as the current character before the cursor.
[0072] Code suggestion validation 530 can validate received code suggestions before providing them for display. For example, code suggestion validation can use one or more validation criteria to determine whether the added characters of the code suggestion are a match or a near match of the character before the cursor (and are added after the time the code suggestion request was made). As indicated at 552, valid code suggestions can be provided for display. In some embodiments, acceptance (or rejection) of the suggestions can be received as indicated at 554 and passed to or included in the file state 540 as indicated at 556.
[0073] Code suggestion validation 530 may identify and display 548 valid coding suggestions and handle acceptance or rejection of the suggestions. Code suggestions 413 may be implemented in different scenarios to provide various code suggestions. Figure 6 6 is a logical block diagram illustrating code suggestions according to some embodiments. A code suggestion request 601 may be received at a tokenizer 610. For example, various different tokenizers or tokenization techniques may be used. A token may be a single word in a sentence, including the leading space character "[space] word" before the word as a token. Punctuation, empty spaces, carriage returns, or various other characters may also be grouped into or individually considered tokens.
[0074] The programming language word prediction model 620 can use the provided code 622 as well as other context 624, such as file context outside the word window or other files that can be obtained using techniques such as regular expressions or parsing, as described above with respect to Figure 5 Programming language specific word prediction models 620 can be used (e.g., model A for language A, model B for language B, etc.). Each of them can be trained with code in the corresponding programming language and other contexts to generate recommendations. The programming language word prediction model 620 can also be trained on other context information (e.g., the same file, the same project, or other contexts), as described above with respect to Figure 5 discussed.
[0075] In some embodiments, code development service 410 may support a custom programming language model. For example, training data or code data from a user's specific code repository may be provided to train a custom programming language model so that it can be used for code suggestions.
[0076] The predictions may be provided to license filtering 630, which may filter or limit the predictions based on the licensing criteria provided by the client. After license filtering 630, the filtered predictions may be provided to selection 640, which may select a prediction to provide as a code suggestion based on the confidence score. In some embodiments, license filtering 630 may be used to train a programming language prediction model to improve subsequent results of candidate code suggestions. In some embodiments, multiple predictions may be provided in a paginated format or other multi-result format, as described above with respect to Figure 5 discussed.
[0077] One scenario where machine learning models that generate text recommendations can fail is when the input has partial words (like "Syst"). In these scenarios, machine learning models tend to provide poor predictions and, therefore, poor recommendations (e.g., generating nonsense data or generating illogical content). This happens because the model only considers word lemmas as input units. To overcome this scenario, backtrack to the last complete lemma and constrain the generation to match the prompt suffix (here "Syst"). As discussed below, constraining the generation helps improve the accuracy of the subword data metric without compromising the gain on the general evaluation set.
[0078] Given a string hint, inconsistencies from normal decoding may be caused by a suffix of the hint, which may appear with a subword that is not a complete word. Matching of the input string suffix can be performed with all available word-grams that start with the suffix or with which the suffix starts. In some embodiments, matching is performed efficiently using a character dictionary tree data structure (e.g., using a local Pytorch array as a fast cascaded node list). Based on a list of matching word-grams, other word-grams can be masked during the next word-gram prediction to ensure that the generation will match the suffix character by character. Further latency optimization can be achieved by maintaining a Boolean mask to cache very frequent suffixes (such as a single space). For each step of matching with the suffix, after each word-gram generation step, the matching word-gram (by character) is removed from the left side of the suffix and a constrained generation is performed until the suffix is an empty string. In some embodiments, partial tokens (e.g., suffixes) are determined by using the same pre-word splitting as in the pre-word strategy of tokenizer 610, which uses word boundaries to perform the splitting - this allows efficient backtracking for character matching because it is known with certainty that any partial token cannot cross a pre-word boundary.
[0079] Figure 7 700 is a logic block diagram showing an example interface of a development environment according to some embodiments. Figure 4 The implementation on the client of the code development service 410 as depicted, or as Figure 4 Depicted as being hosted as part of code development service 410 .
[0080] The integrated development environment interface 700 may include permission criteria 702, which may allow a user to select different criteria 704 to be included in or excluded from code suggestions. Figure 7 As shown, Criteria 1 704a and Criteria 3 704c are selected, while Criteria 2 704b is not selected. According to various embodiments, source code files having licenses matching Criteria 2 may be excluded from potential code suggestions.
[0081] The integrated development environment interface 700 may implement a code editor 710 (e.g., a text editor) that may allow a user to enter code in a programming language. The code suggestion feature 413 of the code development service 410 may analyze the entered characters to determine code suggestions 720, which may be displayed along with an indication of the code source 722 and license type 724. The code suggestion may be added, as indicated at 726. Although not shown, various other information regarding the code metadata may be displayed (e.g., a style guide).
[0082] Figure 8is a flow chart of a method 800 for providing code suggestions based on licensing standards according to some embodiments. The method 800 may be implemented by one or more computing devices. In some embodiments, the method 800 may be implemented as a code suggestion service. According to various embodiments, the code suggestion service may be used with Figure 1 The code suggests generating 120 or Figure 4 The code suggestion corresponds to handling 413.
[0083] Method 800 may include receiving, through an interface of a code suggestion service, a request specifying one or more permission criteria. The request may be provided by a client (eg, a developer) to restrict or filter code suggestions according to the permission criteria.
[0084] Method 800 may also include determining a corresponding license for a corresponding source code file in the plurality of source code files from a source code attribution database, the source code attribution database including indications of corresponding licenses applicable to the plurality of source code files identified by parsing the plurality of source code files. In some embodiments, determining the corresponding license may include parsing the source code file to identify the license included in the source code file. In some embodiments, the source code attribution database may include a record indicating the corresponding license.
[0085] The method 800 may further include generating a set of candidate code suggestions for the received code input based at least in part on the plurality of source code files. The method 800 may also include determining one or more code suggestions that satisfy one or more licensing criteria based at least in part on respective licenses of respective source code files in the plurality of source code files from the set of candidate code suggestions. The method 800 may end by providing one or more code suggestions that satisfy the one or more licensing criteria determined from the set of candidate source code files.
[0086] Fig. 9 is a flow chart of a method 900 for attributing a license to a source code file according to some embodiments. The method 900 may be implemented by one or more computing devices. According to some embodiments, the method 900 may be implemented by Figure 1 The source code license belongs to 130. Figure 2 License attribution service 202 or Figure 4 The code is implemented under license 417.
[0087] Method 900 may include identifying a plurality of source code files for license attribution. Method 900 may include parsing the plurality of source code files to extract one or more text blocks. Method 900 may further include comparing the one or more text blocks to a plurality of licenses to determine a similarity score between a corresponding one of the one or more text blocks and the plurality of licenses. Method 900 may include identifying one or more of the plurality of licenses applicable to an individual source code file of the plurality of source code files based on the similarity score. Method 900 may end by storing the indication of the corresponding license applicable to the corresponding one of the plurality of source code files in the source code attribution database.
[0088] Fig.10 is a flowchart of a method 1000 for parsing license format data to generate a source code attribution database according to some embodiments. The method 1000 may be implemented by one or more computing devices. According to some embodiments, the method 1000 may be implemented by Figure 1 The code license belongs to the database builder 132 or Figure 2 The database builder 204 is implemented.
[0089] The method 1000 may include, at 1002, obtaining license format data from a license data repository at a license database builder. The method 1000 may include determining a corresponding text type of a corresponding portion of the license format data at 1004. The method 1000 may further include annotating the corresponding portion with the determined corresponding text type at 1006. The method 1000 may end by generating at least a portion of a source code attribution database based on the marked license format data at 1008.
[0090] Fig.11 1 shows an example computing device for implementing the various techniques described herein according to some embodiments. For example, in one embodiment, the algorithm execution management system described above can be implemented by a computer device (e.g., Fig.11 The computer system 1100 is implemented as a computer device in a computer system including one or more processors that execute program instructions stored on a computer-readable storage medium coupled to the processor. In the illustrated embodiment, the computer system 1100 includes one or more processors 1110 coupled to a system memory 1120 via an input / output (I / O) interface 1130. The computer system 1100 further includes a network interface 1140 coupled to the I / O interface 1130. Although Fig.11Computer system 1100 is shown as a single computing device, but in various embodiments, computer system 1100 may include one computing device or any number of computing devices configured to work together as a single computer system 1100 .
[0091] In various embodiments, computer system 1100 may be a uniprocessor system including one processor 1110, or a multiprocessor system including several processors 1110 (e.g., two, four, eight, or another suitable number). Processor 1110 may be any suitable processor capable of executing instructions. For example, in various embodiments, processor 1110 may be a general-purpose or embedded processor that implements any of a variety of instruction set architectures (ISAs), such as an x86, PowerPC, SPARC, or MIPS ISA, or any other suitable ISA. In a multiprocessor system, each of processors 1110 may typically, but not necessarily, implement the same ISA.
[0092] The system memory 1120 may be one embodiment of a computer-accessible medium configured to store instructions and data accessible by the processor 1110. In various embodiments, the system memory 1120 may be implemented using any non-transitory storage medium or memory medium (e.g., magnetic or optical media) (e.g., a disk or DVD / CD coupled to the computer system 1100 via the I / O interface 1130). The non-transitory computer-accessible storage medium may also include any volatile or non-volatile media (e.g., RAM (e.g., SDRAM, DDR SDRAM, RDRAM, SRAM, etc.), ROM, etc.), which may be included in some embodiments of the computer system 1100 as the system memory 1120 or another type of memory. In addition, the computer-accessible medium may include a transmission medium or a signal (e.g., an electrical signal, an electromagnetic signal, or a digital signal) conveyed via a communication medium (e.g., a network and / or a wireless link), such as may be implemented via the network interface 1140. In the illustrated embodiment, program instructions (e.g., code) and data (e.g., as described above in the embodiments) that implement one or more desired functions are stored in the computer system 1100. Figure 1-8 The algorithm execution management system described in is shown as being stored within system memory 1130 as code 1125 and data 1126.
[0093] In one embodiment, the I / O interface 1130 may be configured to coordinate I / O traffic between the processor 1110, the system memory 1120, and any peripheral device in the device including the network interface 1140 or other peripheral interface. In some embodiments, the I / O interface 1130 may perform any necessary protocol, timing, or other data transformations to convert data signals from one component (e.g., the system memory 1120) into a format suitable for use by another component (e.g., the processor 1110). In some embodiments, for example, the I / O interface 1130 may include support for devices attached through various types of peripheral buses, such as variants of the peripheral component interconnect (PCI) bus standard or the universal serial bus (USB) standard. In some embodiments, for example, the functionality of the I / O interface 1130 may be split into two or more separate components, such as a north bridge and a south bridge. Moreover, in some embodiments, some or all of the functionality of the I / O interface 1130 (such as the interface of the system memory 1120) may be directly incorporated into the processor 1110.
[0094] The network interface 1140 may be configured to allow data to be exchanged between the computer system 1100 and other devices 1160 attached to one or more networks 1150. In various embodiments, the network interface 1140 may support communications over any suitable wired or wireless general data network (e.g., Ethernet type), for example. Additionally, the network interface 1140 may support communications over a telecommunications / telephone network (e.g., an analog voice network or a digital fiber optic communications network), over a storage area network (e.g., a Fiber Channel SAN), or over any other suitable type of network and / or protocol.
[0095] In some embodiments, the system memory 1120 may be configured to store Figure 1-8 An embodiment of a computer-accessible medium for program instructions and data described herein. In general, a computer-accessible medium may include a non-transitory storage medium or memory medium (such as a magnetic or optical medium), such as a disk or DVD / CD coupled to the computer system 1100 via an I / O interface 1130. A non-transitory computer-accessible storage medium may also include any volatile or non-volatile medium (such as RAM (e.g., SDRAM, DDR SDRAM, RDRAM, SRAM, etc.), ROM, etc.), which may be included in some embodiments of the computer system 1100 as system memory 1120 or another type of memory. In addition, a computer-accessible medium may include a transmission medium or a signal (such as an electrical signal, an electromagnetic signal, or a digital signal) conveyed via a communication medium (such as a network and / or a wireless link), such as may be implemented via a network interface 1140.
[0096] Embodiments of the present disclosure may be described in light of the following terms:
[0097] Clause 1. A system comprising:
[0098] One or more computing devices implementing a code suggestion service, the code suggestion service being configured to:
[0099] receiving, via an interface of the code suggestion service, a request to specify one or more licensing criteria;
[0100] determining a respective license for a respective source code file of a plurality of source code files from a source code attribution database, the source code attribution database containing indications of respective licenses applicable to the plurality of source code files identified by parsing the plurality of source code files;
[0101] generating a set of candidate code suggestions for the received code input based at least in part on the plurality of source code files;
[0102] determining, based at least in part on the corresponding license of the corresponding source code file in the plurality of source code files, one or more code suggestions that satisfy the one or more licensing criteria from the set of candidate code suggestions; and
[0103] The one or more code suggestions determined from the set of candidate source code files that satisfy the one or more licensing criteria are provided.
[0104] Clause 2. The system of clause 1, further comprising one or more additional computing devices, the one or more additional computing devices implementing a license attribution service, the license attribution service configured to:
[0105] identifying the plurality of source code files for license attribution;
[0106] parsing the plurality of source code files to extract one or more text blocks;
[0107] comparing the one or more blocks of text to a plurality of licenses to determine a similarity score between a corresponding block of text in the one or more blocks of text and the plurality of licenses;
[0108] Based on the similarity score, identifying one or more licenses from the plurality of licenses applicable to an individual source code file from the plurality of source code files; and
[0109] The indication of the respective license applicable to the respective one of the plurality of source code files is stored in the source code attribution database.
[0110] Clause 3. The system of clause 2, wherein the license attribution service is further configured to:
[0111] retrieving the plurality of licenses from a license repository; and
[0112] The plurality of licenses are stored in a data storage device.
[0113] Clause 4. A system according to any one of clauses 2 to 3,
[0114] wherein to compare the one or more text blocks with the plurality of licenses to determine the similarity score between the corresponding one of the one or more text blocks and the plurality of licenses, the license attribution service is configured to:
[0115] identifying character strings of the one or more blocks of text that match character strings of the plurality of licenses; and
[0116] The similarity score is determined based on a number of the one or more identified blocks of text that match the character strings of the plurality of licenses.
[0117] Clause 5. The system of any one of clauses 1 to 4, further comprising one or more computing devices configured to implement a database builder service, the database builder service configured to:
[0118] obtain license format data from a license data repository;
[0119] determining a corresponding text type of a corresponding portion of the license format data;
[0120] annotating the corresponding portion with the determined corresponding text type; and
[0121] At least a portion of the source code attribution database is generated based on the annotated license format data.
[0122] Clause 6. A method comprising:
[0123] Executed by one or more computing devices implementing the code suggestion service:
[0124] determining a corresponding license of a corresponding source code file among a plurality of source code files according to a source code attribution database, the source code attribution database containing an indication of the corresponding license of the corresponding source code file;
[0125] generating a set of candidate code suggestions for the received code input based at least in part on the plurality of source code files;
[0126] determining one or more code suggestions that satisfy one or more licensing criteria based on the set of candidate code suggestions; and
[0127] The one or more code suggestions determined from the set of candidate source code files that satisfy the one or more licensing criteria are provided.
[0128] Clause 7. The method according to clause 6, further comprising:
[0129] identifying the plurality of source code files for license attribution;
[0130] parsing the plurality of source code files to extract one or more text blocks;
[0131] comparing the one or more blocks of text to a plurality of licenses to determine a similarity score between a corresponding block of text in the one or more blocks of text and the plurality of licenses;
[0132] identifying, based on the similarity score, one or more licenses from the plurality of licenses applicable to an individual source code file from the plurality of source code files; and
[0133] The indication of the respective license applicable to the respective one of the plurality of source code files is stored in the source code attribution database.
[0134] Clause 8. The method of clause 7, wherein the one or more blocks of text contain code comments included as part of one or more source code files.
[0135] Clause 9. The method according to any one of clauses 7 to 8, further comprising:
[0136] retrieving the plurality of licenses from a license repository; and
[0137] The plurality of licenses are stored in a data storage device.
[0138] Clause 10. The method of any one of clauses 7 to 9, wherein comparing the one or more blocks of text to the plurality of licenses to determine the similarity score between the corresponding block of text in the one or more blocks of text and the plurality of licenses comprises:
[0139] identifying character strings of the one or more blocks of text that match character strings of the plurality of licenses; and
[0140] The similarity score is determined based on a number of the one or more identified blocks of text that match the character strings of the plurality of licenses.
[0141] Clause 11. The method according to any one of clauses 6 to 10, further comprising:
[0142] obtain license format data from a license data repository;
[0143] determining a corresponding text type of a corresponding portion of the license format data;
[0144] annotating the corresponding portion with the determined corresponding text type; and
[0145] At least a portion of the source code attribution database is generated based on the annotated license format data.
[0146] Clause 12. The method of any one of clauses 6 to 11, wherein determining the one or more code suggestions comprises:
[0147] The set of candidate code suggestions is filtered to exclude code suggestions that do not satisfy the one or more permissive criteria.
[0148] Clause 13. A method according to any one of clauses 6 to 12, wherein the licensing criteria include one or more of the following: intellectual property restrictions, open source requirements, monetization restrictions, and copying restrictions.
[0149] Clause 14. One or more computer-readable storage media storing instructions that, when executed on or across one or more processors, cause the one or more processors to:
[0150] determining a corresponding license of a corresponding source code file among a plurality of source code files according to a source code attribution database, the source code attribution database containing an indication of the corresponding license of the corresponding source code file;
[0151] generating a set of candidate code suggestions for the received code input based at least in part on the plurality of source code files;
[0152] determining one or more code suggestions that satisfy one or more licensing criteria based on the set of candidate code suggestions; and
[0153] The one or more code suggestions determined from the set of candidate source code files that satisfy the one or more licensing criteria are provided.
[0154] Clause 15. The one or more computer-readable storage media of Clause 14, further comprising instructions that, when executed on or across the one or more processors, cause the one or more processors to:
[0155] identifying the plurality of source code files for license attribution;
[0156] parsing the plurality of source code files to extract one or more text blocks;
[0157] comparing the one or more blocks of text to a plurality of licenses to determine a similarity score between a corresponding block of text in the one or more blocks of text and the plurality of licenses;
[0158] Based on the similarity score, identifying one or more licenses from the plurality of licenses applicable to an individual source code file from the plurality of source code files; and
[0159] The indication of the respective license applicable to the respective one of the plurality of source code files is stored in the source code attribution database.
[0160] Clause 16. One or more computer-readable storage media as recited in Clause 15, wherein the one or more blocks of text contain code comments included as part of one or more source code files.
[0161] Clause 17. One or more computer-readable storage media according to any one of clauses 15 to 16, further comprising instructions that, when executed on or across the one or more processors, cause the one or more processors to:
[0162] retrieving the plurality of licenses from a license repository; and
[0163] The plurality of licenses are stored in a data storage device.
[0164] Clause 18. One or more computer-readable storage media according to any one of clauses 15 to 17, wherein in order to compare the one or more text blocks with the plurality of licenses to determine the similarity score between the corresponding one of the one or more text blocks and the plurality of licenses, the one or more computer-readable storage media further comprises instructions that, when executed on or across the one or more processors, cause the one or more processors to:
[0165] identifying character strings of the one or more blocks of text that match character strings of the plurality of licenses; and determining the similarity score based on a number of the identified one or more blocks of text that match character strings of the plurality of licenses.
[0166] Clause 19. One or more computer-readable storage media according to any one of clauses 14 to 18, further comprising instructions that, when executed on or across the one or more processors, cause the one or more processors to:
[0167] obtain license format data from a license data repository;
[0168] determining a corresponding text type of a corresponding portion of the license format data;
[0169] annotating the corresponding portion with the determined corresponding text type; and
[0170] At least a portion of the source code attribution database is generated based on the annotated license format data.
[0171] Clause 20. One or more computer-readable storage media according to any one of clauses 14 to 19, further comprising instructions that, when executed on or across the one or more processors, cause the one or more processors to:
[0172] The set of candidate code suggestions is filtered to exclude code suggestions that do not satisfy the one or more permissive criteria.
[0173] Various embodiments may further include receiving, sending or storing instructions and / or data implemented according to the foregoing description on a computer-accessible medium. In general, a computer-accessible medium may include a storage medium or memory medium (such as a magnetic or optical medium, for example, a disk or a DVD / CD-ROM), a volatile or non-volatile medium (such as a RAM (e.g., SDRAM, DDR, RDRAM, SRAM, etc.), a ROM, etc.), and a transmission medium or signal (such as an electrical signal, an electromagnetic signal, or a digital signal) conveyed by a communication medium (such as a network and / or a wireless link).
[0174] The various systems and methods as shown in the accompanying drawings and described herein represent example embodiments of the methods. The systems and methods can be implemented manually, in software, in hardware, or a combination thereof. The order of any method can be changed, and various elements can be added, reordered, combined, omitted, modified, etc.
[0175] It will be apparent to those skilled in the art having the benefit of this disclosure that various modifications and changes may be made. The embodiments are intended to encompass all such modifications and changes, and, therefore, the above description is to be regarded as illustrative rather than limiting.
Claims
1. A system comprising: One or more computing devices implementing a code suggestion service, the code suggestion service being configured to: determining a corresponding license of a corresponding source code file of a plurality of source code files according to a source code attribution database, the source code attribution database containing indications of the corresponding licenses of the plurality of source code files; generating a set of candidate code suggestions for the received code input based at least in part on the plurality of source code files; determining one or more code suggestions that satisfy one or more licensing criteria based on the set of candidate code suggestions; and The one or more code suggestions determined from the set of candidate source code files that satisfy the one or more licensing criteria are provided.
2. The system of claim 1 , further comprising one or more additional computing devices, the one or more additional computing devices implementing a license attribution service, the license attribution service being configured to: identifying the plurality of source code files for license attribution; parsing the plurality of source code files to extract one or more text blocks; comparing the one or more blocks of text to a plurality of licenses to determine a similarity score between a corresponding block of text in the one or more blocks of text and the plurality of licenses; identifying, based on the similarity score, one or more licenses from the plurality of licenses applicable to an individual source code file from the plurality of source code files; and The indication of the respective license applicable to the respective one of the plurality of source code files is stored in the source code attribution database.
3. The system of claim 2, wherein the license attribution service is further configured to: retrieving the plurality of licenses from a license repository; and The plurality of licenses are stored in a data storage device.
4. A system according to any one of claims 2 to 3, wherein to compare the one or more text blocks with the plurality of licenses to determine the similarity score between the corresponding one of the one or more text blocks and the plurality of licenses, the license attribution service is configured to: identifying character strings of the one or more blocks of text that match character strings of the plurality of licenses; and The similarity score is determined based on a number of the one or more identified blocks of text that match the character strings of the plurality of licenses.
5. The system according to any one of claims 1 to 4, further comprising one or more computing devices, wherein the one or more computing devices are configured to implement a database builder service, wherein the database builder service is configured to: obtain license format data from a license data repository; determining a corresponding text type of a corresponding portion of the license format data; annotating the corresponding portion with the determined corresponding text type; and At least a portion of the source code attribution database is generated based on the annotated license format data.
6. A method comprising: Executed by one or more computing devices implementing the code suggestion service: determining a corresponding license of a corresponding source code file among a plurality of source code files according to a source code attribution database, the source code attribution database containing an indication of the corresponding license of the corresponding source code file; generating a set of candidate code suggestions for the received code input based at least in part on the plurality of source code files; determining one or more code suggestions that satisfy one or more licensing criteria based on the set of candidate code suggestions; and The one or more code suggestions determined from the set of candidate source code files that satisfy the one or more licensing criteria are provided.
7. The method according to claim 6, further comprising: identifying the plurality of source code files for license attribution; parsing the plurality of source code files to extract one or more text blocks; comparing the one or more blocks of text to a plurality of licenses to determine a similarity score between a corresponding block of text in the one or more blocks of text and the plurality of licenses; identifying, based on the similarity score, one or more licenses from the plurality of licenses applicable to an individual source code file from the plurality of source code files; as well as The indication of the respective license applicable to the respective one of the plurality of source code files is stored in the source code attribution database.
8. The method of claim 7, wherein the one or more blocks of text contain code comments included as part of one or more source code files.
9. The method according to any one of claims 7 to 8, further comprising: retrieving the plurality of licenses from a license repository; and The plurality of licenses are stored in a data storage device.
10. The method of any one of claims 7 to 8, wherein comparing the one or more text blocks to the plurality of licenses to determine the similarity score between the corresponding one of the one or more text blocks and the plurality of licenses comprises: identifying character strings of the one or more blocks of text that match character strings of the plurality of licenses; and The similarity score is determined based on a number of the one or more identified blocks of text that match the character strings of the plurality of licenses.
11. The method according to any one of claims 6 to 10, further comprising: obtain license format data from a license data repository; determining a corresponding text type of a corresponding portion of the license format data; annotating the corresponding portion with the determined corresponding text type; and At least a portion of the source code attribution database is generated based on the annotated license format data.
12. The method of any one of claims 6 to 11, wherein determining the one or more code suggestions comprises: The set of candidate code suggestions is filtered to exclude code suggestions that do not satisfy the one or more permissive criteria.
13. The method of any one of claims 6 to 12, wherein the licensing criteria include one or more of the following: intellectual property restrictions, open source requirements, monetization restrictions, copying restrictions.
14. One or more computer-readable storage media storing instructions that, when executed on or across one or more processors, cause the one or more processors to: determining a corresponding license of a corresponding source code file among a plurality of source code files according to a source code attribution database, the source code attribution database containing an indication of the corresponding license of the corresponding source code file; generating a set of candidate code suggestions for the received code input based at least in part on the plurality of source code files; determining one or more code suggestions that satisfy one or more licensing criteria based on the set of candidate code suggestions; and The one or more code suggestions determined from the set of candidate source code files that satisfy the one or more licensing criteria are provided.
15. The one or more computer-readable storage media of claim 14, further comprising instructions that, when executed on or across the one or more processors, cause the one or more processors to: identifying the plurality of source code files for license attribution; parsing the plurality of source code files to extract one or more text blocks; comparing the one or more blocks of text to a plurality of licenses to determine a similarity score between a corresponding block of text in the one or more blocks of text and the plurality of licenses; identifying, based on the similarity score, one or more licenses from the plurality of licenses applicable to an individual source code file from the plurality of source code files; and The indication of the respective license applicable to the respective one of the plurality of source code files is stored in the source code attribution database.