Distributed Version Control Method Based on Object Storage and Fine-Grained Access Control
By adopting distributed object storage and fine-grained permission access control in distributed version control systems, the problems of storage space expansion, branch access efficiency and permission management are solved, and high security and efficient data storage and access are achieved.
Patent Information
- Application Number
- CN202310333285.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-31
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-03-31
AI Technical Summary
The existing distributed version control system has shortcomings in storage space scalability, unstructured data and small file storage, and the permission management is not granular enough, resulting in low data storage security and access efficiency.
Use distributed object storage as the storage medium of the version library to quickly expand through a flat storage architecture; monitor the branch information of the version library, create, update or delete the working directory in the bucket, and realize online branch access; adopt fine-grained permission access control method, set permissions through scope, resources, policies and authorized object templates to ensure that access permissions are controlled to the minimum level.
It improves the security of data storage and the scalability of storage space, realizes efficient branch online access and fine-grained permission management, and improves the security and efficiency of data access.
Smart Images

Figure CN116467280B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and further relates to a distributed version control method based on object storage and fine-grained access control in the field of computer data processing technology. The present invention can be used for distributed version control of file data that needs to be version-controlled and shared for team collaboration. Background Art
[0002] Distributed version control is a design pattern of a version control system. Each developer has a local complete code repository and can perform operations such as code modification, version control, commit, and rollback locally without directly interacting with the central server. Developers share and synchronize modifications through the central server in the form of collaborative pushing and pulling of code. With the advent of the big data era, the order of magnitude of file data has increased exponentially. Especially for file data such as data sets that have numerous sub-files, occupy a large amount of storage space, and need to be version-controlled, management has become extremely cumbersome, and file version management is chaotic. Currently, most distributed file version control systems use a file storage system as the storage medium and directly store file data in a specific directory structure of the file storage system. However, the file storage system is restricted by the hardware of the storage server, resulting in limited storage space capacity and poor scalability of the storage space, unable to meet the needs of storing huge amounts of data. In addition, as the number of file data to be managed increases, the creator of a specific Git repository will add new members to jointly manage the Git repository and share the data in the Git repository. This requires management of member users and reasonable division of access rights of member users to the file repository according to their roles.
[0003] Beijing Jingdong Shangke Information Technology Co., Ltd. published a data processing method for a distributed version control system in the patent document "A Data Processing Method, Device and System for a Distributed Version Control System" (Patent Application No.: 201310726376.6, Publication No.: CN 103647850A). This method is applied to a Git service device, which establishes a first link with a distributed database service device and interacts with the distributed database device through the first link to process the storage location of the Git repository; it also establishes a second link with a distributed file storage service device and interacts with the distributed file storage system through the second link to process the storage of the Git repository. This method uses HDFS as the distributed file system. Although it solves the problem of expanding the storage space for storing the Git repository to a certain extent, there are still deficiencies. HDFS distributed file storage stores the metadata of data in the memory space of the NameNode node. Due to the limitation of the NameNode node's memory space, the scalability of the storage space of this distributed version control system is restricted by the memory size of the NameNode. Secondly, HDFS organizes data in blocks, so it cannot well meet the storage requirements of unstructured data and a large number of small files.
[0004] 58.com, Inc. published a branch access method and device based on a distributed version control system in the patent document "Branch Access Method and Device Based on a Distributed Version Control System" (Patent Application No.: 201810960642.4, Publication No.: CN 109271194 A). This method determines several branches of the complete file data in the distributed version control system, creates a corresponding first branch directory for the specified branch, and stores the file data of the specified branch under the corresponding first branch directory; according to the target branch to be accessed, it searches for the corresponding first branch directory and accesses the file data in the found first branch directory. Although this method realizes the convenient access to several branches in the distributed version control system, there are still deficiencies. It is necessary to manually determine the integrity of the branch file data and update the file data in the corresponding branch directory, resulting in the inability to update the branch file data in real time, low update efficiency of the branch file data, and when accessing the branch file, it may still access the old version file. Summary of the Invention
[0005] The object of the present invention is to address the deficiencies of the above-mentioned existing technologies and propose a distributed version control method based on object storage and fine-grained access control, which is used to solve the problems of poor scalability of the storage space of the distributed version control system server, inability to meet large-scale data storage, online convenient access to branch files in the repository, and access control of the repository.
[0006] The specific idea for achieving the object of the present invention is as follows: Since the present invention uses distributed object storage as the storage medium of the repository of the distributed version control system, relying on its flat storage architecture, by adding storage nodes, the rapid expansion of the storage space of the repository of the distributed version control system can be quickly achieved, so as to meet the needs of storing large-scale repository data; in addition, through the data recovery, data multi-copy and data error correction characteristics of distributed object storage, the security of data storage is ensured, and the problem of data loss caused by a single point of failure of the storage module is avoided. When creating a repository, create a corresponding bucket, mount the bucket in the object storage service to the file system directory where the distributed version control system service is located, create and initialize the repository in this directory, and realize the one-to-one mapping between the bucket of the object storage service and the server-side repository, which simplifies the data management process of the server-side repository and easily realizes data migration and backup with the help of the object storage service. Secondly, the present invention monitors the branch information of the server-side repository, creates, updates or deletes the corresponding working directory in the bucket, and realizes convenient online access to branches by creating the corresponding working directory in the bucket for several branches existing in the repository in the distributed version control system. Finally, the present invention adopts a fine-grained access control method to control access to the repository, sets the permission set, basic information, access rules and granted permissions of the repository through scope, resource, policy and authorization object template respectively, and controls the access permission to the smallest level, realizing fine-grained access control of the repository.
[0007] The steps of the present invention are as follows:
[0008] Step 1, create a bucket:
[0009] Use distributed object storage as the storage medium of the repository of the distributed version control system. In the distributed object storage service, create a bucket for storing the server-side repository, and use distributed object storage to save the server-side repository;
[0010] Step 2, mount the bucket:
[0011] Mount the bucket of the storage server-side repository and map it to the path agreed by the file system of the server where the distributed version control server is located, and use the distributed version control server to manage and call the repository;
[0012] Step 3, initialize the repository:
[0013] The distributed version control server executes the repository initialization command to initialize the repository under the bucket mapping path;
[0014] Step 4, create a corresponding working directory for the main branch:
[0015] Create a corresponding working directory for the main branch that exists by default in the repository, monitor the creation process of the distributed version control server repository, and create a working directory for storing the branch content in the bucket where the repository is stored;
[0016] Step 5, create a collection of scope objects:
[0017] Create a collection of scope objects corresponding to the permission set executed by the user on the server repository, set the scope object name, and obtain a one-to-one mapping relationship from repository permissions to scope objects;
[0018] Step 6, create a resource object describing the repository:
[0019] Set the name field of the resource object indicating the repository name, the type field indicating whether the type of the repository is a public type, the creator field indicating the creator information of the repository, and the permission set field indicating the set of operation permissions open for the repository respectively, and the permission set is a subset of the permission set mapped by the scope object;
[0020] Step 7, create a policy object describing the repository access rules:
[0021] Set the name field of the policy object indicating the access policy name, the role conditions constituting the access repository rules, the description field indicating the access policy description information, and the logical condition field indicating whether the access rule logic is positive or negative respectively;
[0022] Step 8, create an authorization object describing the permission granting information:
[0023] Adopt a fine-grained permission access control method, and set the resource name field of the authorization object indicating the repository information, the policy field indicating the adopted policy information, and the permission set attribute field indicating the set of access permissions to the repository granted after meeting the access policy respectively, and the access permission set is a subset of the access permission set in the resource object in Step 6, bind the repository and the access policy, specify the access permission set, and grant the access permission of the repository to the user or a class of users with a specific role;
[0024] Step 9, process requests and implement data interaction:
[0025] Step 9.1, after the distributed version control server receives the client's request for obtaining repository reference information, parse the repository from the mounted bucket, obtain the repository reference information according to the role-based access rules of the repository, and return the reference information;
[0026] Step 9.2, determine whether the current request is a data transfer request for the download data process. If so, execute Step 9.3; otherwise, execute Step 9.4;
[0027] Step 9.3, the distributed version control server receives and processes the data transfer request during the download data process, parses the repository from the mounted storage bucket, packs the repository data and returns the data, and then executes Step 10;
[0028] Step 9.4, the distributed version control server receives and processes the data transfer request during the upload data process, parses the repository from the mounted storage bucket, receives the data and updates the repository branch information, monitors the server repository branch information, and creates, updates or deletes the corresponding working directory in the storage bucket;
[0029] Step 10, complete the data interaction.
[0030] The present invention has the following advantages compared with the prior art:
[0031] First, since the present invention uses distributed object storage as the storage medium of the repository of the distributed version control system, it avoids the problem of data loss caused by a single point of failure of the storage module in the prior art, making the present invention have higher data storage security and high scalability of the storage space.
[0032] Second, the present invention overcomes the problem of frequent branch switching or untimely update of file data under different branches of the online access repository in the prior art by monitoring the server repository branch information and creating, updating or deleting the corresponding working directory in the storage bucket, making the present invention have higher efficiency of online access to branches.
[0033] Third, the present invention uses a fine-grained permission access control method to control access to the repository, overcomes the problems of unreasonable or rough permission division in the prior art, enables the present invention to refine the repository authorization, grants different permissions to users with different roles by setting member roles, prevents unauthorized access and data leakage, protects sensitive information, and thus improves the security of data access. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is the overall flowchart of the present invention;
[0035] Figure 2 is the flowchart of the present invention for receiving Git requests and implementing data interaction. DETAILED DESCRIPTION OF THE INVENTION
[0036] The following will further describe the present invention with reference to the drawings and embodiments.
[0037] Refer to Figure 1 , taking Git version management as an example, the implementation steps of the embodiments of the present invention will be further described.
[0038] In the embodiments of the present invention, the Git server establishes links with the authentication service and the distributed object storage service respectively. The Git server completes data interaction by calling the reserved interfaces of the authentication service and the distributed object storage service, and respectively realizes creating an original Git repository and managing the Git repository in the distributed object storage service, binding an access control policy to the created Git version in the authentication service, and judging the access permission when receiving a Git request according to the access control policy, thereby realizing fine-grained access permission control.
[0039] After the embodiment receives a request to create a Git repository, the following steps are executed to create a Git repository in the distributed object storage service and bind an access rule to the Git repository to realize fine-grained access control of the Git repository.
[0040] Step 1, create a bucket.
[0041] According to the Git repository name, use the API interface provided by the distributed object storage service to create a bucket with the same name in the object storage service. The naming rule of the bucket must meet the following conditions: it can only be composed of numbers, lowercase letters, and hyphens, must start with a number and a lowercase letter, and the name must be between 3 and 63 characters. By using the distributed object storage to save the Git repository, its features of data recovery, data multi-copy, and data error correction ensure the security and high availability of data storage; through its flat storage structure, rapid expansion of the storage space of the Git repository is realized.
[0042] Step 2, mount the bucket.
[0043] The Git repository stored in the distributed object storage service and the Git server are located on different servers. Use the S3FS mounting tool to mount the bucket storing the Git repository and map it to the path agreed upon by the file system of the server where the Git server is located, which is convenient for the Git server to manage and call the Git repository.
[0044] Step 3, initialize the repository.
[0045] The Git server executes the Git repository initialization command and sets it as a bare repository. Under the mapped path of the bucket, execute the instruction "git init --bare ${name}.git" to create and initialize the Git repository, where ${name} is the name of the Git repository.
[0046] Step 4, create a working directory with the same name for the master branch.
[0047] Create a working directory with the same name as the master branch that exists by default in the Git repository. By monitoring the creation process of the Git repository, after the Git repository is created, enter the path of the created Git repository and execute the "git worktree add.. / worktree / master master" command to create a working tree with the same name as the master branch of the bare repository in the storage bucket, enabling online access to the file contents under different branches of the Git repository through the API provided by the distributed object storage service without switching branches.
[0048] Step 5, create a collection of range objects.
[0049] Create a corresponding collection of range objects according to the operation permissions that the user can perform on the server-side Git repository. Git operation requests can be roughly divided into two types of operations. One is the download operation including clone, pull, and fetch operations, and the other is the upload operation centered around the push operation. The upload operation includes creating, updating, deleting branches, and creating, deleting tags. Therefore, create a collection of range objects according to the operation permissions, and set the name field of the range object to establish a one-to-one mapping relationship between the operation permissions and the range objects.
[0050] Step 6, create a resource object that describes the Git repository.
[0051] Create a resource object for this Git repository to describe its information. Create a resource object for the Git repository according to the basic information of the Git repository, and clarify the public operation permissions of this Git repository through the resource object. With the help of the resource object, mark the repository name through the name field, mark whether the type of this repository is a public type through the type field, mark the creator of the repository through the owner field, mark the operation permissions opened by this repository through the resource_scopes field and resource_scopes is a subset of the collection of range objects, and mark some additional information of this repository through the attributes field.
[0052] Step 7, create a policy object that describes the access rules of the Git repository.
[0053] Set the access rules for this Git repository by creating role - based policy objects. Create different policy objects according to the types of member roles that the Git repository has, and set role - based access rules. With the help of policy objects, the role conditions for this access rule are indicated by the "roles.id" field, whether this role condition must be met is indicated by the "roles.required" field, the name of this access rule is indicated by the "name" field, the description of this access rule is in the "description" field, and whether this access rule is positive or negative is indicated by the "logic" field.
[0054] Step 8, create an authorization object that describes the permission - granting information.
[0055] Bind this repository to specific access rules and clarify the access permissions by creating an authorization object. Create different authorization objects according to the need to grant different sets of access permissions to different types of roles for the Git repository, and clarify that certain access permissions to the Git repository are granted to users under a specific role condition. With the help of the authorization object, the protected resource, i.e., the Git repository, is indicated by the "resources" field, the access rules adopted are indicated by the "policies" field, and which access permissions to the Git repository are granted after meeting the access rules are indicated by the "scopes" field. The set of range objects in this field is a subset of the set of range objects represented by the "resource_scopes" field in the resource object in Step 6.
[0056] Step 9, process the request and implement data interaction.
[0057] Refer to Figure 2 and the embodiments to further illustrate the steps of processing requests to implement data interaction in the present invention. The data interaction process can be divided into an upload data process and a download data process. Both the upload data process and the download data process involve two HTTP requests. The difference between the two data interaction processes lies in the data transfer requests;
[0058] Step 9.1, the Git server receives and processes the request for obtaining the reference information of the Git repository in the data interaction process, parses the repository from the storage bucket where the Git repository is mounted, obtains the reference information of the repository according to the role - based access rules of the Git repository, and returns the reference information;
[0059] Step 9.2, determine whether the Git request is for a download data process request or an upload data process request. If it is for downloading data, execute Step 9.3; otherwise, execute Step 9.4;
[0060] Step 9.3, The Git server receives and processes the data transfer request during the download data process, parses the repository from the mounted bucket, packages the corresponding data of the Git repository according to the data required by the client indicated in the request parameters, and returns the data, then executes Step 10;
[0061] Step 9.4, The Git server receives and processes the data transfer request during the upload data process, parses the repository from the mounted bucket, receives the data and updates the Git repository;
[0062] Step 9.5, After the version library branch information is updated through the Git server hook post-receive, according to the creation, data update or deletion of the branch, then the creation, data update and deletion of the corresponding branch working directory are carried out in sequence for each case.
[0063] Step 10, Complete the data interaction of the Git repository.
Claims
1. A distributed version control method based on object storage and fine-grained access control, characterized in that, Using a distributed object storage as the storage medium for the repository of a distributed version control system, monitoring the branch information of the server-side repository, creating, updating, or deleting corresponding working directories in the storage bucket, and using a fine-grained access control method to control access to the repository; the steps of this control method are as follows: Step 1, create a storage bucket: Using a distributed object storage as the storage medium for the repository of a distributed version control system, in the distributed object storage service, create a storage bucket for storing the server-side repository, and use the distributed object storage to save the server-side repository; Step 2, mount the storage bucket: Mount the storage bucket of the server-side repository for storing the repository, and map it to the path specified by the file system of the server where the distributed version control server is located, and use the distributed version control server to manage and call the repository; Step 3, initialize the repository: The distributed version control server executes the repository initialization command to initialize the repository under the mapped path of the storage bucket; Step 4, create a corresponding working directory for the main branch: Create a corresponding working directory for the main branch that exists by default in the repository, monitor the creation process of the server-side repository of the distributed version control system, and create a working directory for storing branch content in the storage bucket where the repository is stored; Step 5, create a set of scope objects: Create a set of scope objects corresponding to the permission set executed by the user on the server-side repository, set the scope object name, and obtain a one-to-one mapping relationship from the repository permissions to the scope objects; Step 6, create a resource object describing the repository: Respectively set the name field of the resource object indicating the repository name, the type field indicating whether the type of the repository is a public type, the creator field indicating the creator information of the repository, and the permission set field indicating the set of operation permissions open for the repository, and the permission set is a subset of the permission set mapped by the scope object; Step 7, create a policy object describing the repository access rules: Respectively set the name field of the policy object indicating the access policy name, the role condition indicating the composition of the rules for accessing the repository, the description field indicating the description information of the access policy, and the logical condition field indicating whether the access rule logic is positive or negative; Step 8, create an authorization object describing the permission granting information: Using a fine-grained access control method, respectively set the resource name field of the authorization object indicating the repository information, the policy field indicating the adopted policy information, and the permission set attribute field indicating the set of access permissions to the repository granted after meeting the access policy, and the access permission set is a subset of the access permission set in the resource object in Step 6, bind the repository and the access policy, specify the access permission set, and grant the access permission of the repository to a user or a class of users with a specific role; Step 9, process requests and implement data interaction: Step 9.1, after the distributed version control server receives a request from the client to obtain repository reference information, parse the repository from the mounted storage bucket, obtain the repository reference information according to the role-based access rules of the repository, and return the reference information; Step 9.2: Determine whether the current request is a data transfer request for downloading process data. If so, proceed to Step 9.3; otherwise, proceed to Step 9.
4. Step 9.3: The distributed version control server receives and processes the data transfer request during the download data process, parses the repository from the mounted storage bucket, packs the repository data, returns the data, and then proceeds to Step 10. Step 9.4: The distributed version control server receives and processes the data transfer request during the upload data process, parses the repository from the mounted storage bucket, receives the data and updates the repository branch information, monitors the server repository branch information, and creates, updates, or deletes the corresponding working directory in the storage bucket. Step 10: Complete the data interaction.
2. The distributed version control method based on object storage and fine-grained access control according to claim 1, characterized in that, The set of scope objects described in Step 5 includes scope objects for accessing the server repository, downloading the repository, updating the main branch of the repository, creating, updating, and deleting non-main branches of the repository, merging repository branches, releasing and deleting repository versions, and mapping deletion permissions of the repository.
3. The distributed version control method based on object storage and fine-grained access control according to claim 1, wherein The update of the repository branch information described in Step 9.4 includes three cases, namely the creation of a branch, the data update of a branch, and the deletion of a branch. Subsequently, the creation, data update, and deletion of the corresponding branch working directory are performed for each case in sequence.
Citation Information
Patent Citations
Data processing method, device and system of distributed version control system
CN103647850A
A data processing method, device and system for a distributed version control system
CN103647850B
Branch access method and device based on distributed version control system
CN109271194A
Information processing method and device, equipment, storage medium and computer program product
CN114047943A
System method and article of manufacture for building, managing, and supporting various components of a system
US6957186B1