Gradle-based software rpm package construction method and system
By constructing a two-layer cache and optimizing the resource competitor algorithm, the Gradle RPM package building process is optimized, solving the problems of low efficiency in dependency resource management and insufficient parallelism, and achieving efficient RPM package building and stable software delivery.
Patent Information
- Application Number
- CN202511249142.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-09-03
AI Technical Summary
Existing Gradle-based RPM package building methods suffer from inefficient dependency resource management, insufficient parallelism in the build process, and non-standard service configuration management, resulting in low build efficiency and high operational complexity.
A two-layer buffer is constructed using a time-sequenced replacement algorithm. The parallelism parameter is calculated using a resource competitor algorithm. A depth-first search algorithm is applied to extract the configuration item hierarchy structure and generate a standardized service control instruction set, thereby realizing intelligent parallelization of the construction task and systematic management of configuration items.
It significantly shortens build time, improves resource utilization, ensures consistency and reliability of RPM packages in different environments, reduces deployment failure rate, and enhances the stability of software delivery.
Smart Images

Figure CN120762733B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of software packaging technology, and in particular to a method and system for building software RPM packages based on Gradle. Background Technology
[0002] In the field of enterprise software development, packaging and distribution of Java Web projects is a crucial step in the software delivery process. RPM (Red Hat Package Manager), a widely used package management system on Linux systems, provides a standardized approach to software installation, upgrades, and uninstallation. With the increasing prevalence of cloud computing and microservice architectures, the demand for deployment on different processor architecture platforms is growing. Heterogeneous environments with multiple architectures such as x86 and ARM, for example, make package building more complex.
[0003] Gradle, as an advanced build automation tool, has gradually become dominant in Java project building due to its flexible configuration capabilities and efficient incremental build features. Traditional RPM package building typically uses spec files to define build rules, while combining Gradle with RPM package building can fully leverage Gradle's dependency management and build automation advantages, simplifying the RPM packaging process for Java Web projects.
[0004] However, existing Gradle-based RPM package building methods have shortcomings, including inefficient dependency resource management. During large project builds, frequent downloads of the same dependencies lead to wasted network resources, and existing caching mechanisms often use simple local storage, failing to intelligently manage dependencies based on usage frequency, resulting in excessively long build environment initialization times. Furthermore, the build process suffers from insufficient parallelism. Existing methods typically use fixed parallelism parameters or simple system resource assessment methods, failing to dynamically adjust parallelism based on actual resource contention and task characteristics, leading to low resource utilization and limited build efficiency in complex projects. Finally, service configuration management is often inconsistent. In traditional RPM builds, service control scripts are often manually written or generated using simple templates, lacking systematic analysis and standardized processing of configuration items. This results in compatibility issues with service startup and shutdown operations in different environments, increasing operational complexity. Summary of the Invention
[0005] This invention provides a method and system for building software RPM packages based on Gradle, which can solve the problems in the prior art.
[0006] A first aspect of this invention provides a method for building software RPM packages based on Gradle, comprising:
[0007] Receive the package path and RPM package build parameters of the Java Web project, and parse out the target processor architecture type;
[0008] A two-layer cache is constructed based on the time-series eviction algorithm. The frequency of use of construction dependencies is calculated to generate a feature sequence. A data channel is built between the local cache and the shared cache to form a set of construction dependency resources.
[0009] Initialize the Gradle environment using the build dependency resource set and generate a list of build task configurations;
[0010] Based on the build task configuration list, the parallelism parameter is calculated using the resource competitor algorithm. The program package is then split into build task units according to the parallelism parameter, parallel build operations are executed, and the RPM base package of the target architecture is output.
[0011] Perform a configuration scan on the RPM base package, apply a depth-first search algorithm to extract the configuration item hierarchy, generate a service parameter table, and build a standardized service control instruction set based on the service parameter table;
[0012] Integrate the standardized service control instruction set into the RPM base package to build the RPM target package;
[0013] Perform the RPM target package installation test in an isolated environment, and output the final RPM package after successful verification.
[0014] In one optional embodiment, constructing a two-level buffer based on a time-series eviction algorithm includes:
[0015] A local cache area is established using off-heap memory, a shared cache area is established using a distributed file system, and a data channel is constructed between the local cache area and the shared cache area;
[0016] The frequency of use of the build dependency is statistically analyzed, and the time decay value of the build dependency is calculated based on the usage time. The time decay value is multiplied by the usage frequency to obtain the usage weight of the build dependency.
[0017] Slide the sampling window within a preset time period to record the usage time interval of the build dependencies and generate a feature sequence of the build dependencies.
[0018] The resource activity score of the build dependency is calculated based on the weight and the feature sequence, and a first cache threshold and a second cache threshold are set based on the resource activity score.
[0019] When the resource activity score of a build dependency is lower than the first cache threshold, the build dependency is moved out of the local cache; when the resource activity score is lower than the second cache threshold, the build dependency is moved to the shared cache.
[0020] In an optional embodiment, calculating the resource activity score of the build dependency based on the usage weight and the feature sequence, and setting a first cache threshold and a second cache threshold based on the resource activity score includes:
[0021] A base score is set based on the weight used, and a fluctuation range is calculated based on the feature sequence. The product of the base score and the fluctuation range is used as the resource activity score of the corresponding construction dependency.
[0022] Collect a sample set of resource activity scores for dependencies built within a preset time window, and normalize the data to obtain a standard distribution sequence.
[0023] Calculate the expected value and variance of the resource activity score corresponding to the dependency item based on the standard distribution sequence; obtain the total capacity and current occupancy rate of the cached resources, and calculate the capacity adjustment coefficient based on the current occupancy rate;
[0024] The threshold base is determined by weighting the expected value, the variance value, and the capacity adjustment coefficient.
[0025] A first cache threshold is set based on the product of the threshold base and a preset high-order decay factor to trigger local cache eviction; a second cache threshold is set based on the product of the threshold base and a preset low-order decay factor to trigger migration to a shared cache.
[0026] The system monitors the current occupancy rate of cached resources in real time. When the current occupancy rate exceeds a preset alarm value, the capacity adjustment coefficient is dynamically adjusted according to a preset increment.
[0027] The resource activity score is updated based on the update cycle to determine the latest capacity adjustment coefficient, and the first cache threshold and the second cache threshold are calibrated based on the latest capacity adjustment coefficient.
[0028] In one alternative embodiment, constructing the data channel includes:
[0029] Receive build dependency migration requests from the local cache;
[0030] Obtain the version identifier of the build dependency to be migrated, and check the data integrity of the build dependency;
[0031] Build dependencies are grouped according to data block size to generate a data transmission queue. A mapping table is established between the data transmission queue and the shared cache. Build dependencies in the data transmission queue are written to the corresponding storage location in the shared cache according to the mapping table.
[0032] Record the storage status of the build dependencies in the shared cache and generate a resource location index;
[0033] The build dependencies in the local cache and the shared cache are integrated to form a build dependency resource set.
[0034] In one optional embodiment, based on the build task configuration list, a parallelism parameter is calculated using a resource competitor algorithm. The package is then split into build task units according to the parallelism parameter, parallel build operations are performed, and the output RPM base package of the target architecture includes:
[0035] Parse the build task configuration list to obtain package dependency information, generate a dependency relationship topology graph based on the package dependency information, traverse the dependency relationship topology graph to calculate the build complexity weight of each package, and output the initial build task sequence according to the build complexity weight;
[0036] Collect computational resource usage data of the construction environment, construct a competition matrix based on resource usage overlap rate and perform eigenvalue decomposition to determine resource competition factors, and optimize the initial construction task sequence using graph coloring algorithm to obtain parallelism parameters;
[0037] Based on the parallelism parameter, the file dependencies in the package are scanned to obtain the dependency set. The dependency set is analyzed to identify the smallest independent compilation unit and the smallest independent compilation unit is combined to form a build task unit.
[0038] The execution priority sequence is obtained by prioritizing the build task units, and the build task units are allocated to the build queue to perform parallel build operations according to the execution priority sequence;
[0039] The execution status of the construction task unit is collected in real time. When the execution status shows an abnormality, the task is retried and the task that has exceeded the number of retries is migrated to a new computing node.
[0040] The execution progress of the construction task unit is periodically obtained. When the execution progress meets the preset performance decay threshold, the parallelism parameter is recalculated and the task is reallocated.
[0041] The integrity and consistency of the build artifacts of the build task unit are checked to obtain the check results. Based on the check results, the build artifacts are organized to generate the RPM base package of the target architecture.
[0042] In one optional embodiment, computational resource usage data of the construction environment is collected, a competition matrix is constructed based on the resource usage overlap rate, and eigenvalue decomposition is performed to determine the resource competition factor. Parallelism parameters are obtained by optimizing the initial construction task sequence using a graph coloring algorithm, including:
[0043] Collect CPU utilization data, memory usage data, disk I / O bandwidth utilization data, and disk space remaining data in the construction environment, and perform normalization processing to obtain computing resource usage data;
[0044] The computational resource usage data is converted into a competition matrix. The resource usage overlap rate of the construction task in each data dimension is calculated and weighted summation is performed to obtain the competition coefficient. The competition coefficient is used to fill the corresponding element positions of the competition matrix. Based on the construction complexity weight, the competition matrix is decomposed into eigenvalues, and the eigenvector corresponding to the largest eigenvalue is taken to obtain the resource competition factor.
[0045] The construction tasks are treated as vertices in a graph. When the resource contention factor between task pairs is greater than a preset contention threshold, an edge is established between the corresponding vertices to construct a conflict graph. The vertices are sorted in descending order according to the size of the resource contention factor of the construction tasks. An integer set of available color numbers is established, starting from zero and incrementing. Starting from the vertex with the highest degree, the smallest color number that has not been used by adjacent vertices is assigned to each vertex. The total number of colors used is recorded as a baseline value. The parallelism parameter is obtained by dynamically adjusting the baseline value in combination with the current load status.
[0046] In one optional embodiment, a configuration scan is performed on the RPM base package, a depth-first search algorithm is applied to extract the configuration item hierarchy, a service parameter table is generated, and a standardized service control instruction set is constructed according to the service parameter table, including:
[0047] The configuration files in the RPM base package are preprocessed to convert the configuration information into key-value pairs to obtain the configuration content.
[0048] The configuration content is traversed using a depth-first search algorithm. A tree structure of configuration items is constructed with the configuration group as the root node. The parent-child relationship and dependency relationship between configuration items are identified. The association relationship of sibling configuration items is extracted. The hierarchical depth information of configuration items is recorded. The namespace mapping of configuration items is established to obtain the hierarchical structure of configuration items.
[0049] A service parameter table is generated based on the configuration item hierarchy to record the attribute information of the service parameters;
[0050] Based on the service parameter table, define the basic operation instructions, parameter configuration instructions, and status monitoring instructions for the service, establish instruction combination rules and execution order constraints, and generate a standardized service control instruction set.
[0051] A second aspect of this invention provides a Gradle-based software RPM package building system, comprising:
[0052] The first unit is used to receive the package path and RPM package build parameters of the Java Web project and parse out the target processor architecture type.
[0053] The second unit is used to construct a two-layer cache based on the time-series eviction algorithm, calculate the frequency of use of construction dependencies to generate a feature sequence, build a data channel between the local cache and the shared cache, and form a set of construction dependency resources.
[0054] The third unit is used to initialize the Gradle environment using the build dependency resource set and generate a build task configuration list.
[0055] The fourth unit is used to calculate the parallelism parameter based on the build task configuration list using the resource competitor algorithm, split the program package into build task units according to the parallelism parameter, execute parallel build operations, and output the RPM base package of the target architecture.
[0056] The fifth unit is used to perform configuration scanning on the RPM base package, apply a depth-first search algorithm to extract the configuration item hierarchy, generate a service parameter table, and build a standardized service control instruction set according to the service parameter table;
[0057] The sixth unit is used to integrate the standardized service control instruction set into the RPM base package to build the RPM target package;
[0058] Unit 7 is used to perform RPM target package installation tests in an isolated environment, and outputs the final RPM package after successful verification.
[0059] A third aspect of the present invention provides an electronic device, comprising:
[0060] processor;
[0061] Memory used to store processor-executable instructions;
[0062] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0063] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0064] In this embodiment of the invention, the Gradle-based RPM package building method achieves efficient management and optimization of the build process. The dual-layer caching mechanism built using the time-sequence elimination algorithm effectively improves the speed of acquiring dependent resources, reduces repeated downloads, and significantly shortens the build time. The innovative resource competitor algorithm realizes intelligent parallelization of the build task, making full use of system resources while avoiding resource contention, thus maximizing the build efficiency. In particular, the optimized utilization of multi-core processors improves the build speed compared to traditional methods. The application of the depth-first search algorithm makes configuration item extraction more accurate, and the generated service parameter table and standardized service control instruction set ensure the consistency and reliability of the RPM package in different environments. The isolated environment testing mechanism further guarantees the quality of the generated package, significantly reduces the deployment failure rate, and enhances the stability of software delivery. Attached Figure Description
[0065] Figure 1 This is a flowchart illustrating the Gradle-based software RPM package construction method according to an embodiment of the present invention.
[0066] Figure 2 Build a logic flowchart for the RPM base package. Detailed Implementation
[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0068] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0069] Figure 1 This is a flowchart illustrating the Gradle-based software RPM package building method according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0070] Receive the package path and RPM package build parameters of the Java Web project, and parse out the target processor architecture type;
[0071] A two-layer cache is constructed based on the time-series eviction algorithm. The frequency of use of construction dependencies is calculated to generate a feature sequence. A data channel is built between the local cache and the shared cache to form a set of construction dependency resources.
[0072] Initialize the Gradle environment using the build dependency resource set and generate a list of build task configurations;
[0073] Based on the build task configuration list, the parallelism parameter is calculated using the resource competitor algorithm. The program package is then split into build task units according to the parallelism parameter, parallel build operations are executed, and the RPM base package of the target architecture is output.
[0074] Perform a configuration scan on the RPM base package, apply a depth-first search algorithm to extract the configuration item hierarchy, generate a service parameter table, and build a standardized service control instruction set based on the service parameter table;
[0075] Integrate the standardized service control instruction set into the RPM base package to build the RPM target package;
[0076] Perform the RPM target package installation test in an isolated environment, and output the final RPM package after successful verification.
[0077] In one optional implementation, constructing a two-level buffer based on a time-series eviction algorithm includes:
[0078] A local cache area is established using off-heap memory, a shared cache area is established using a distributed file system, and a data channel is constructed between the local cache area and the shared cache area;
[0079] The frequency of use of the build dependency is statistically analyzed, and the time decay value of the build dependency is calculated based on the usage time. The time decay value is multiplied by the usage frequency to obtain the usage weight of the build dependency.
[0080] Slide the sampling window within a preset time period to record the usage time interval of the build dependencies and generate a feature sequence of the build dependencies.
[0081] The resource activity score of the build dependency is calculated based on the weight and the feature sequence, and a first cache threshold and a second cache threshold are set based on the resource activity score.
[0082] When the resource activity score of a build dependency is lower than the first cache threshold, the build dependency is moved out of the local cache; when the resource activity score is lower than the second cache threshold, the build dependency is moved to the shared cache.
[0083] In one specific implementation, the two-tier cache comprises a local cache and a shared cache. The local cache is implemented using off-heap memory, while the shared cache is constructed using a distributed file system. The two caches are interconnected via a data channel to ensure data flow between the two layers. Off-heap memory refers to memory space not managed by the JVM, which can reduce the impact of JVM garbage collection and improve system performance; the distributed file system provides shared storage capabilities across multiple servers.
[0084] To construct the local cache, 1GB of off-heap memory is allocated and managed using the DirectByteBuffer class. This approach does not consume JVM heap memory, avoiding frequent garbage collection. Off-heap memory is allocated using the ByteBuffer.allocateDirect(1024×1024×1024) method, and an LRU eviction policy is configured as the basic cache eviction mechanism. The shared cache is implemented through a distributed file system deployed on a 3-node cluster, with each node providing 5TB of storage space, forming a total 15TB shared storage pool. This system supports high-concurrency read and write operations and is configured with a data redundancy mechanism to ensure data reliability.
[0085] This embodiment establishes a bidirectional data channel between two cache layers. Data transfer from the local cache to the shared cache is implemented through asynchronous writes. When a build dependency needs to be transferred from the local cache to the shared cache, a dedicated transfer task is created, and data is transferred in 64MB chunks. Data transfer from the shared cache to the local cache adopts an on-demand loading strategy. When an application requests access to a build dependency in the shared cache, it checks whether the item is already cached locally. If not, the data is read from the shared cache and loaded into the local cache.
[0086] When tracking the frequency of build dependencies, a counter mapping table is maintained, where the key is the unique identifier of the build dependency and the value is the number of times it is accessed. Each time a build dependency is accessed, the corresponding counter is incremented. For example, for the build dependency "libA.so", the initial count is 0, the count becomes 1 after the first access, and then increases to 6 after 5 consecutive accesses.
[0087] The time decay value is calculated based on the difference between the last access time of a dependency and the current time. A time decay function is used so that the weight of the dependency gradually decreases over time. Specifically, the current time T_current and the last access time T_last of the dependency are recorded, and the time difference T_diff (in hours) is calculated. The decay factor is set to 0.9, and the time decay value is calculated as 0.9 raised to the power of T_diff. For example, if a dependency was accessed 2 hours ago, its time decay value is 0.9², approximately 0.81.
[0088] The weighting calculation multiplies the usage frequency by the time decay value. For example, for a build dependency with an access frequency of 6, if its time decay value is 0.81, then its usage weight is 6 × 0.81 = 4.86.
[0089] Feature sequence generation is achieved through a sliding sampling window, with a window size of 24 hours and a sliding step of 1 hour. Within this window, the timestamp of each access to the build dependency is recorded, and the time interval between consecutive accesses is calculated. For example, for the build dependency "libB.so", if the access times within the 24-hour window are [9:00, 10:30, 15:45, 16:20], then the time interval sequence is [90 minutes, 315 minutes, 35 minutes]. These time intervals form the feature sequence of this dependency.
[0090] Resource activity scores combine usage weights with a feature sequence, calculating the mean and standard deviation of time intervals within the feature sequence. Shorter intervals and smaller fluctuations indicate more stable usage. For example, if a dependency has an average time interval of 120 minutes, a standard deviation of 50 minutes, and a usage weight of 4.86, its resource activity score could be 4.86 × (1 - 120 / 1440) × (1 - 50 / 120) = 2.27. Scores range from 0 to 10, with higher scores indicating greater activity.
[0091] Based on resource activity scores, two cache thresholds are dynamically set. The local cache threshold (first cache threshold) is set to 3.0, and the shared cache threshold (second cache threshold) is set to 1.0. These thresholds can be dynamically adjusted according to resource conditions. When the local cache utilization rate exceeds 85%, the first cache threshold may be increased to 3.5; when the shared cache utilization rate exceeds 90%, the second cache threshold may be increased to 1.5.
[0092] The cache eviction policy is executed based on resource activity scores. When the resource activity score of the build dependency "libC.so" is found to be 2.5, which is lower than the local cache threshold of 3.0 but higher than the shared cache threshold of 1.0, this dependency is removed from the local cache and transferred to the shared cache via the data channel. If another dependency "libD.so" has a score of 0.8, which is lower than the shared cache threshold of 1.0, it is deleted from the shared cache, freeing up storage space.
[0093] The two-tiered cache management process runs as a scheduled task, performing cache evaluation and eviction operations every 15 minutes. In each evaluation, a resource activity score is calculated for all cached items, and corresponding cache migration or deletion operations are performed based on two thresholds. This approach automatically adjusts the caching strategy according to the usage patterns of build dependencies, improving cache utilization efficiency and build performance.
[0094] In one optional implementation, calculating the resource activity score of the build dependency based on the usage weight and the feature sequence, and setting a first cache threshold and a second cache threshold based on the resource activity score includes:
[0095] A base score is set based on the weight used, and a fluctuation range is calculated based on the feature sequence. The product of the base score and the fluctuation range is used as the resource activity score of the corresponding construction dependency.
[0096] Collect a sample set of resource activity scores for dependencies built within a preset time window, and normalize the data to obtain a standard distribution sequence.
[0097] Calculate the expected value and variance of the resource activity score corresponding to the dependency item based on the standard distribution sequence; obtain the total capacity and current occupancy rate of the cached resources, and calculate the capacity adjustment coefficient based on the current occupancy rate;
[0098] The threshold base is determined by weighting the expected value, the variance value, and the capacity adjustment coefficient.
[0099] A first cache threshold is set based on the product of the threshold base and a preset high-order decay factor to trigger local cache eviction; a second cache threshold is set based on the product of the threshold base and a preset low-order decay factor to trigger migration to a shared cache.
[0100] The system monitors the current occupancy rate of cached resources in real time. When the current occupancy rate exceeds a preset alarm value, the capacity adjustment coefficient is dynamically adjusted according to a preset increment.
[0101] The resource activity score is updated based on the update cycle to determine the latest capacity adjustment coefficient, and the first cache threshold and the second cache threshold are calibrated based on the latest capacity adjustment coefficient.
[0102] In one specific implementation, when calculating the resource activity score of build dependencies, a base score is set based on usage weights. Different usage weights are assigned to different types of build dependencies, such as code repositories, binary packages, and build tools. For example, the weight of a code repository is 0.4, a binary package is 0.3, a build tool is 0.2, and other dependencies are 0.1. These weights can be adjusted according to actual conditions to reflect the importance of different dependencies in the build process. The base score can be set by multiplying the weight by a fixed coefficient. For example, the base score for a code repository is 0.4 × 10 = 4, the base score for a binary package is 0.3 × 10 = 3, the base score for a build tool is 0.2 × 10 = 2, and the base score for other dependencies is 0.1 × 10 = 1.
[0103] The fluctuation range is calculated based on the feature sequence, which includes dependency size, access frequency, and last access time. The fluctuation range calculation method comprehensively considers the dependency size factor, access frequency factor, and time decay factor. For example, for a 50MB codebase accessed 10 times per hour with a last access time of 5 minutes ago, the size factor can be set to 0.5, the frequency factor to 0.8, and the time decay factor to 0.9, resulting in a fluctuation range of 0.5 × 0.8 × 0.9 = 0.36. The product of the base score and the fluctuation range is used as the resource activity score for this dependency, i.e., 4 × 0.36 = 1.44.
[0104] Within a preset time window, such as 24 hours, a sample set of resource activity scores for dependency items is collected. Assume the sample set of scores for a certain dependency item during this period is {1.44, 1.52, 1.38, 1.47, 1.55, 1.42, 1.49, 1.53, 1.45, 1.51}. The sample set is normalized to distribute it between 0 and 1, resulting in the standard distribution sequence {0.12, 0.28, 0, 0.18, 0.35, 0.08, 0.22, 0.31, 0.14, 0.27}. The expected value and variance of the resource activity scores are calculated based on the standard distribution sequence. The expected value is the average of all standardized scores, 0.195, and the variance is the dispersion of the standardized scores, 0.012.
[0105] Obtain the total capacity and current occupancy of the cached resources. Assume the total capacity is 1000MB and the current occupancy is 60%. Calculate the capacity adjustment factor based on the current occupancy using a piecewise function: when the occupancy is below 30%, the adjustment factor is 0.8; when the occupancy is between 30% and 70%, the adjustment factor is 1.0; and when the occupancy is above 70%, the adjustment factor is 1.2. In this example, the occupancy is 60%, so the capacity adjustment factor is 1.0.
[0106] The threshold base is determined by weighting the expected value, variance, and capacity adjustment coefficient. The calculation formula can be set as: Threshold base = Expected value × 0.7 + Variance × 3 × Capacity adjustment coefficient. Substituting the values: Threshold base = 0.195 × 0.7 + 0.012 × 3 × 1.0 = 0.1725. The first cache threshold is set based on the product of the threshold base and a preset high-level decay factor, used to trigger local cache eviction. The high-level decay factor can be set to 0.85, then the first cache threshold = 0.1725 × 0.85 = 0.1466. The second cache threshold is set based on the product of the threshold base and a preset low-level decay factor, used to trigger migration to the shared cache. The low-level decay factor can be set to 0.65, then the second cache threshold = 0.1725 × 0.65 = 0.1121.
[0107] The system monitors the current occupancy rate of cached resources in real time. When the current occupancy rate exceeds a preset alarm value, the capacity adjustment coefficient is dynamically adjusted according to a preset increment. The preset alarm value can be set to 75%, and the preset increment can be set to 0.05. For example, if the monitored occupancy rate rises to 78%, exceeding the preset alarm value, the dynamic capacity adjustment coefficient will be 1.0 + 0.05 = 1.05.
[0108] The resource activity score is updated according to the set update cycle, such as every 4 hours. Assuming that after 4 hours, the score of a certain dependency changes from 1.44 to 1.58, and the cache utilization rate becomes 82%, the capacity adjustment coefficient is 1.0 + 0.05 × 2 = 1.1 (an increment is added for each time the preset alarm value is exceeded). The threshold base and cache threshold are recalculated to calibrate the first and second cache thresholds.
[0109] When the build process needs to access a dependency, the resource activity score of that dependency is calculated and compared with a first cache threshold and a second cache threshold. If the score is lower than the second cache threshold, the dependency is migrated from the local cache to the shared cache; if the score is lower than the first cache threshold but higher than the second cache threshold, the dependency is kept in the local cache but marked as evictionable; if the score is higher than the first cache threshold, the dependency is kept in the local cache and marked as high priority.
[0110] In this way, the cache management system can dynamically adjust its caching strategy based on the actual usage of dependencies, improving cache utilization efficiency. For example, for a large software project containing hundreds of build dependencies, this method can keep frequently used dependencies such as core libraries in the local cache, while migrating less frequently used dependencies such as testing tools to the shared cache, thereby reducing local cache usage and speeding up the build process.
[0111] Traditional Gradle build cache management schemes typically employ a simple FIFO (First-In, First-Out) strategy for cache eviction, failing to adequately consider the usage characteristics of dependencies and system resource conditions. These methods lack the ability to differentiate between build dependencies of varying importance, resulting in important dependencies being evicted prematurely while less important dependencies occupy cache space for extended periods.
[0112] This embodiment introduces a resource activity scoring mechanism, comprehensively considering the usage weight of dependencies, access characteristics, and system resource status, to achieve more refined cache management. This addresses the inefficiency of traditional caching strategies in software development, particularly the decline in build efficiency caused by unreasonable cache resource allocation in large-scale distributed team collaboration environments. By dynamically adjusting cache thresholds, it can adapt to different resource pressures, raising the eviction criteria when resources are scarce and relaxing retention conditions when resources are plentiful.
[0113] In one alternative implementation, constructing the data channel includes:
[0114] Receive build dependency migration requests from the local cache;
[0115] Obtain the version identifier of the build dependency to be migrated, and check the data integrity of the build dependency;
[0116] Build dependencies are grouped according to data block size to generate a data transmission queue. A mapping table is established between the data transmission queue and the shared cache. Build dependencies in the data transmission queue are written to the corresponding storage location in the shared cache according to the mapping table.
[0117] Record the storage status of the build dependencies in the shared cache and generate a resource location index;
[0118] The build dependencies in the local cache and the shared cache are integrated to form a build dependency resource set.
[0119] In one specific implementation, during the build dependency migration process, when the resource activity score of the local cache falls below the second cache threshold, the local cache sends a build dependency migration request to the shared cache. This request contains basic information about the build dependency, such as the dependency name, version number, and size. Upon receiving the migration request, the shared cache immediately parses the request and extracts the build dependency information to be migrated. Specifically, the migration request is encapsulated in JSON format, containing fields such as requestId, dependencyName, version, size, and priority. For example, a migration request for a core library dependency can be represented as: {"requestId":"req-20250812-001","dependencyName":"core-utils","version":"2.5.3","size":"15MB","priority":"medium"}.
[0120] After obtaining the version identifier of the build dependency to be migrated from the shared cache, the data integrity of the build dependency needs to be checked. The version identifier is usually a combination of a semantic version number (e.g., 2.5.3) and a build timestamp (e.g., 20250812143022). Data integrity checking uses a checksum mechanism, which is implemented by calculating the SHA-256 hash value of the dependency file and comparing it with the original hash value. If the hash values do not match, the data is determined to be incomplete, the migration process is stopped, and an error message is returned to the local cache; if the hash values match, subsequent migration steps continue. For example, for the core-utils-2.5.3.jar file, its original SHA-256 hash value is "7f83b1657ff1fc53b92dc18148a1d65dfc2d4b1fa3d677284addd200126d9069". The calculated hash value should match this to ensure data integrity.
[0121] After data integrity verification passes, the shared cache will group dependencies according to a preset data block size, generating a data transmission queue. The choice of data block size significantly impacts transmission efficiency; excessively large blocks may cause network congestion, while excessively small blocks increase transmission overhead. In practice, a suitable data block size is typically 1MB-4MB, which can be dynamically adjusted based on network conditions. For large dependencies, such as a 15MB core-utils-2.5.3.jar file, it can be divided into four data blocks: [0-3.99MB], [4-7.99MB], [8-11.99MB], and [12-15MB]. The data transmission queue is implemented using a priority queue, with priority determined by the importance of the dependency. The queue structure includes fields such as blockId, dataRange, priority, and status.
[0122] After generating the data transfer queue, the shared cache establishes a mapping table from the data transfer queue to the shared cache storage location. The mapping table uses key-value pairs, where the key is the data block identifier and the value is the target storage path. Storage paths are typically generated according to dependency types, versions, and other information for efficient retrieval. For example, for the first data block of core-utils-2.5.3.jar, the mapping table entry could be: {"blockId":"core-utils-2.5.3-block-1","storagePath":" / shared / cache / libs / core-utils / 2.5.3 / block-1"}. After the mapping table is generated, the shared cache writes the build dependencies from the data transfer queue to the corresponding storage locations based on the mapping table. The write process uses asynchronous I / O operations to improve transmission efficiency. For each data block, a temporary file is first created on the target path, and then renamed to the final file after writing is complete, ensuring atomicity of the write operation. If an error occurs during the write process, such as insufficient storage space, the written data block is rolled back, and an error message is returned to the local cache.
[0123] After data writing is complete, the shared cache records the storage status of the build dependencies, generating a resource location index. The storage status includes basic information about the dependency, storage path, write time, access count, and last access time. For example, for a successfully migrated core-utils-2.5.3.jar, its storage status can be represented as: {"dependencyId":"core-utils-2.5.3","storagePath":" / shared / cache / libs / core-utils / 2.5.3","storeTime":"20250812144530","accessCount":0,"lastAccessTime":"20250812144530","status":"active"}. The resource location index uses a multi-level index structure, organized according to dependency type, name, version, and other dimensions, supporting fast retrieval. Index entries contain references to the actual storage location, as well as metadata information such as size and hash value. To improve index performance, a memory cache can be used to store the index information of popular dependencies, employing an LRU strategy to manage the cache.
[0124] After generating the resource location index, the shared cache integrates build dependencies from both the local and shared caches to form a build dependency resource set. During integration, for dependencies that exist in both the local and shared caches, the version in the local cache is used first to reduce network transmission overhead. The build dependency resource set uses a unified access interface, hiding the underlying storage details from upper-layer applications. A proxy pattern can be used for this implementation. When an application requests access to a dependency, the proxy object first checks the local cache; if it doesn't exist, it automatically retrieves it from the shared cache. For example, when the build process needs to access the core-utils library, the resource set provides the interface: getDependency("core-utils", "2.5.3"). This method returns the actual access path of the dependency, and the application doesn't need to know its physical storage location.
[0125] In real-world applications, such as a Gradle-based RPM package build process, hundreds of dependencies may be involved. The method described in this example migrates dependencies with low resource activity to a shared cache, while still ensuring fast access. For instance, when building an RPM package for an enterprise application, the core framework library, due to frequent use, is kept in the local cache, while the quarterly updated reporting component library is migrated to the shared cache. When the build requires the reporting component, the build system locates the resource through the build dependency resource set and automatically retrieves the component from the shared cache; the entire process is transparent to the developers.
[0126] The shared cache also implements dependency version management. For different versions of the same dependency, such as versions 2.5.3 and 2.5.4 of core-utils, the shared cache stores them simultaneously and distinguishes them using version identifiers. When the build process specifies the use of a particular version, it can accurately locate and provide the required version. Furthermore, for dependency versions that are used very infrequently and have not been accessed for a long time, the shared cache periodically performs cleanup to free up storage space. The cleanup strategy is based on access time and frequency to ensure the availability of important dependencies.
[0127] Through the above implementation methods, the system can efficiently manage local and shared cache resources, improve dependency access efficiency, reduce duplicate downloads, and accelerate the software RPM package building process. Especially in large projects involving multi-developer collaboration, the shared caching mechanism effectively reduces the storage pressure on each developer's environment while ensuring the stability and efficiency of the build process.
[0128] In one optional implementation, based on the build task configuration list, the parallelism parameter is calculated using a resource competitor algorithm. The package is then split into build task units according to the parallelism parameter, parallel build operations are performed, and the output RPM base package of the target architecture includes:
[0129] Parse the build task configuration list to obtain package dependency information, generate a dependency relationship topology graph based on the package dependency information, traverse the dependency relationship topology graph to calculate the build complexity weight of each package, and output the initial build task sequence according to the build complexity weight;
[0130] Collect computational resource usage data of the construction environment, construct a competition matrix based on resource usage overlap rate and perform eigenvalue decomposition to determine resource competition factors, and optimize the initial construction task sequence using graph coloring algorithm to obtain parallelism parameters;
[0131] Based on the parallelism parameter, the file dependencies in the package are scanned to obtain the dependency set. The dependency set is analyzed to identify the smallest independent compilation unit and the smallest independent compilation unit is combined to form a build task unit.
[0132] The execution priority sequence is obtained by prioritizing the build task units, and the build task units are allocated to the build queue to perform parallel build operations according to the execution priority sequence;
[0133] The execution status of the construction task unit is collected in real time. When the execution status shows an abnormality, the task is retried and the task that has exceeded the number of retries is migrated to a new computing node.
[0134] The execution progress of the construction task unit is periodically obtained. When the execution progress meets the preset performance decay threshold, the parallelism parameter is recalculated and the task is reallocated.
[0135] The integrity and consistency of the build artifacts of the build task unit are checked to obtain the check results. Based on the check results, the build artifacts are organized to generate the RPM base package of the target architecture.
[0136] In one specific implementation, the build task configuration manifest is parsed. This manifest, typically stored in JSON or YAML format, contains project build information, dependencies, and build parameters. The parsing process extracts dependency information between packages, such as package A depending on packages B and C, package B depending on package D, etc. Taking an enterprise application as an example, its configuration manifest might include core service packages depending on data access packages and toolkits, and data access packages depending on infrastructure packages, etc. Specifically, the parsing implementation can obtain project dependency information through Gradle's Project API and perform structured processing of the dependency relationships.
[0137] Based on the parsed inter-package dependency information, a dependency topology graph is generated. This graph represents the dependencies between packages in the form of a directed acyclic graph. Each node in the graph represents a package, and edges represent dependencies, with the direction of the edges pointing from the dependent to the dependent. The topology graph is generated using an adjacency list storage structure, which facilitates subsequent traversal and analysis. For the enterprise application example mentioned above, its topology graph includes nodes such as core services, data access, tools, and infrastructure, and the connection relationships between the edges reflect the dependencies between them.
[0138] The dependency topology graph is traversed to calculate the build complexity weight of each package. The weight calculation considers three factors: the number of direct dependencies, the depth of transitive dependencies, and the package size. For each package node, its out-degree (number of times it is depended on) and in-degree (number of times it depends on other packages) are calculated, while also considering the node's hierarchical position in the graph. The build complexity weight calculation formula can be expressed as: the number of direct dependencies multiplied by 0.4, plus the depth of transitive dependencies multiplied by 0.3, plus the normalized value of the package size multiplied by 0.3. For example, the core service package has 2 direct dependencies, a transitive dependency depth of 2, and a package size of 20MB (normalized value 0.8), so its complexity weight is 2×0.4 + 2×0.3 + 0.8×0.3 = 1.44. Based on the calculated complexity weights, all packages are arranged in descending order of weight to form the initial build task sequence.
[0139] The computational resource usage data for the build environment is collected through a resource monitoring module, which periodically collects metrics such as CPU utilization, memory usage, disk I / O, and network bandwidth. Resource usage overlap rate is calculated based on historical build data, analyzing the competition for various resources during the build process of different packages. The competition matrix is an n×n matrix (n is the number of packages), where matrix elements represent the degree of resource competition between two packages. The degree of competition is determined by the ratio of the intersection to the union of resource usage; a larger value indicates more intense competition. Eigenvalue decomposition is performed on the competition matrix to extract the main eigenvectors as resource competition factors. These resource competition factors reflect the resource contention patterns of each package and are used to guide parallel build strategies.
[0140] Graph coloring algorithms are used to optimize the initial sequence of build tasks by assigning competing tasks to different parallel groups. In implementation, program packages are treated as vertices in a graph. If the resource contention between two packages exceeds a threshold (e.g., 0.6), an edge is added between the vertices. A greedy coloring method is used to assign colors to the vertices in the graph, ensuring that adjacent vertices have different colors. The number of colors represents the parallelism parameter, indicating the number of build task groups that can be executed simultaneously. In practical applications, if the core service package and the data access package have intense resource contention, they will be assigned to different build groups to avoid resource conflicts caused by simultaneous execution.
[0141] Based on the parallelism parameter, file dependencies in the package are scanned, and a dependency set is obtained by analyzing source code references, resource file dependencies, etc. The scanning process can leverage Gradle's dependency analysis API, combined with in-depth analysis using static code analysis tools. The dependency set contains detailed file-level dependency information, such as class A depending on class B, configuration file C being referenced by multiple classes, etc. Analyzing the dependency set identifies the smallest independent compilation unit, i.e., the smallest set of code that can be compiled independently. The identification process uses a strongly connected component algorithm to group tightly coupled files together. For enterprise application examples, the smallest independent compilation unit might be at the functional module level, such as a user management module, an access control module, etc. The smallest independent compilation units are combined according to functional relevance and dependencies to form build task units, each task unit containing explicit input dependencies and expected outputs.
[0142] When prioritizing build task units, three metrics are considered: dependency position (those at the bottom of the dependency chain have higher priority), historical build time (longer-running tasks are executed first), and build failure rate (higher-failure-rate tasks are executed earlier to detect problems sooner). The priority evaluation formula can be expressed as: dependency level inverse value multiplied by 0.5, plus normalized build time multiplied by 0.3, plus normalized failure rate multiplied by 0.2. Task units are sorted according to the calculation results to obtain an execution priority sequence. Build task units are allocated to build queues according to the priority sequence, with each queue corresponding to a parallel build group. The build queues are implemented using a producer-consumer pattern, with Gradle worker threads acting as consumers, retrieving tasks from the queues for execution.
[0143] During the build process, the execution status of task units is collected in real time, including execution stage, resource usage, and log output. When abnormal states such as compilation errors, test failures, or resource exhaustion are detected, a task retry mechanism is initiated. Retry employs an exponential backoff strategy, with an initial retry interval of 5 seconds, subsequently increasing to 10 seconds, 20 seconds, and so on. The maximum number of retries is set to 3; after exceeding this limit, the task is migrated to a new compute node. The task migration process includes serializing the task status, transferring the build context, and restoring the execution environment on the new node. The selection of a new node is based on resource availability scoring; the node with the highest score is chosen to execute the migration task.
[0144] The execution progress of build task units is periodically retrieved, and monitoring metrics include completion percentage, build speed, and resource utilization. When the execution progress meets a preset performance degradation threshold (e.g., build speed drops by more than 30% or resource utilization falls below 50%), the parallelism parameter is recalculated. The recalculation considers the current resource status and the characteristics of remaining tasks, and may increase or decrease the parallelism. Tasks are then reallocated based on the new parallelism parameters, which may involve task merging or splitting operations. During task reallocation, a build cache of completed work is retained to avoid duplicate builds.
[0145] The build artifacts of the build task unit undergo integrity and consistency checks. Integrity checks ensure that all expected files have been generated, while consistency checks verify that the file content meets expectations. Check methods include file manifest verification, checksum verification, and functional testing. Check results include success rate, error list, and warning messages. Based on the check results, the build artifacts are organized into an RPM package structure, containing binary files, configuration files, documentation, and metadata. Finally, an RPM build tool is used to generate the base RPM package for the target architecture, such as x86_64 or arm64. The generation process specifies the package name, version number, dependencies, and installation scripts to ensure that the RPM package conforms to standards and can be correctly installed and used in the target environment.
[0146] Using the methods described above, the Gradle build system can efficiently and intelligently build software RPM packages, fully utilize parallel computing resources, handle complex dependencies, and ensure build quality and efficiency.
[0147] like Figure 2 The diagram shown illustrates the logical flowchart for building the RPM base package.
[0148] In one optional implementation, computational resource usage data of the construction environment is collected, a competition matrix is constructed based on the resource usage overlap rate, and eigenvalue decomposition is performed to determine the resource competition factor. Parallelism parameters are obtained by optimizing the initial construction task sequence using a graph coloring algorithm, including:
[0149] Collect CPU utilization data, memory usage data, disk I / O bandwidth utilization data, and disk space remaining data in the construction environment, and perform normalization processing to obtain computing resource usage data;
[0150] The computational resource usage data is converted into a competition matrix. The resource usage overlap rate of the construction task in each data dimension is calculated and weighted summation is performed to obtain the competition coefficient. The competition coefficient is used to fill the corresponding element positions of the competition matrix. Based on the construction complexity weight, the competition matrix is decomposed into eigenvalues, and the eigenvector corresponding to the largest eigenvalue is taken to obtain the resource competition factor.
[0151] The construction tasks are treated as vertices in a graph. When the resource contention factor between task pairs is greater than a preset contention threshold, an edge is established between the corresponding vertices to construct a conflict graph. The vertices are sorted in descending order according to the size of the resource contention factor of the construction tasks. An integer set of available color numbers is established, starting from zero and incrementing. Starting from the vertex with the highest degree, the smallest color number that has not been used by adjacent vertices is assigned to each vertex. The total number of colors used is recorded as a basic score. The parallelism parameter is obtained by dynamically adjusting the basic score in combination with the current load status.
[0152] In one specific implementation, the construction environment needs to collect multi-dimensional computing resource usage data, including CPU utilization, memory utilization, disk I / O bandwidth utilization, and remaining disk space. Data can be collected every 5 seconds through the resource monitoring interface provided by the operating system. For example, the `top` command can obtain the CPU utilization percentage, the `free` command the memory utilization percentage, the `iostat` command the disk I / O bandwidth utilization percentage, and the `df` command the remaining disk space percentage. The collected raw data needs to be normalized, mapping each dimension of data to a standard range between 0 and 1. Normalization can be achieved using the maximum-minimum method; that is, for the original data x, the normalized value is (x-min) / (max-min), where min and max are the minimum and maximum values of that dimension, respectively. For example, if the CPU utilization in a certain collection is 75%, and the minimum value is known to be 10% and the maximum value is 90%, then the normalized value is (75-10) / (90-10) = 0.8125.
[0153] The normalized multi-dimensional resource usage data needs to be converted into a competition matrix to measure the resource competition relationship between construction tasks. The competition matrix is an n×n square matrix, where n is the number of parallel construction tasks. Each element of the competition matrix represents the degree of resource competition between two corresponding tasks. When calculating the competition coefficient, the resource usage overlap rate of the two tasks in each data dimension must first be calculated. For task i and task j, the overlap rate in the CPU dimension can be calculated using min(CPU_i, CPU_j), where CPU_i and CPU_j are the normalized CPU utilization rates of the two tasks, respectively. Similarly, the overlap rates in the memory, disk I / O, and disk space dimensions can be calculated. The overlap rates in each dimension need to be weighted and summed to obtain the competition coefficient. The weights can be set according to the characteristics of the construction tasks, for example, CPU weight 0.4, memory weight 0.3, disk I / O weight 0.2, and disk space weight 0.1. After the competition coefficient is calculated, it is filled into the corresponding positions in the competition matrix. For example, suppose there are 3 build tasks. Task 1 and Task 2 have a CPU overlap rate of 0.7, a memory overlap rate of 0.6, a disk I / O overlap rate of 0.5, and a disk space overlap rate of 0.4. Then their competition coefficient is 0.7×0.4+0.6×0.3+0.5×0.2+0.4×0.1=0.61, which is filled into the (1,2) and (2,1) positions of the competition matrix.
[0154] After the competition matrix is constructed, eigenvalue decomposition is performed based on the construction complexity weights. The construction complexity weights can be determined by factors such as the number of lines of code for each task, the number of dependencies, and the historical construction time. Eigenvalue decomposition can be implemented using a power-law iteration method, with 100 iterations and a convergence threshold of 1e-6. The eigenvector corresponding to the largest eigenvalue is taken as the resource competition factor, and each component of this eigenvector represents the resource competition intensity of each construction task. For example, for the competition matrix of the above three tasks, the largest eigenvalue obtained after eigenvalue decomposition is 2.1, and the corresponding eigenvector is [0.58, 0.63, 0.52], indicating that task 2 has the highest competition intensity, followed by task 1, and finally task 3.
[0155] A conflict graph is constructed based on the calculated resource contention factor. Tasks are used as vertices in the graph. When the resource contention factor between a pair of tasks is greater than a preset contention threshold, an edge is created between the corresponding vertices. The preset contention threshold can be set to 0.5. For example, if the resource contention factors of task 1 and task 2 are greater than 0.5, an edge is created between vertex 1 and vertex 2, indicating that these two tasks should not be executed in parallel. After the conflict graph is constructed, the vertices are sorted in descending order according to the size of the resource contention factor, with vertices having larger contention factors appearing earlier in the order.
[0156] A graph coloring algorithm is used to assign colors to conflict graphs. An integer set {0, 1, 2, ...} is created, starting from zero and incrementing. Starting with the vertex with the highest degree, each vertex is assigned the smallest color number not used by its adjacent vertices. For example, given vertices with degrees of 3, 2, 2, and 1, ordered as v1, v2, v3, and v4, if v1 is adjacent to v2, v3, and v4, v2 is adjacent to v1 and v3, v3 is adjacent to v1 and v2, and v4 is adjacent to v1, then the coloring result is: v1 uses color 0, v2 uses color 1, v3 uses color 2, and v4 uses color 1. After coloring, the total number of colors used is recorded as a baseline value. In the example above, the baseline value is 3, indicating that at least 3 time slices are needed to complete all tasks.
[0157] The parallelism parameter is obtained by dynamically adjusting the baseline value based on the current load status. The load can be obtained by collecting the average load value of the build machines. If the current system load value is lower than the threshold (e.g., 1.0), the parallelism can be increased appropriately, such as by multiplying the baseline value by 0.8 to obtain the new parallelism parameter; if the current system load value is higher than the threshold (e.g., 3.0), the parallelism can be decreased appropriately, such as by multiplying the baseline value by 1.2 to obtain the new parallelism parameter. For example, if the baseline value is 3 and the current system load is 0.8, the parallelism parameter is 3 × 0.8 = 2.4, which, rounded down to 2, indicates that a maximum of 2 build tasks are allowed to execute in parallel.
[0158] In practical applications, the method of this embodiment can be integrated into the Gradle build system and implemented as a Gradle plugin. When building an RPM package, the dependencies between build tasks are parsed, and a task dependency graph is generated. For tasks that do not depend on each other, the resource contention factor is calculated using the method described above, a conflict graph is constructed and colored, and the parallelism parameter is obtained. This parameter controls Gradle's maxParallelForks parameter, enabling intelligent parallel scheduling of build tasks.
[0159] Compared to traditional Gradle build methods, which typically employ fixed parallelism or simple parallelism strategies based on the number of CPU cores, the parallelism cannot be dynamically adjusted according to the resource requirements of different build tasks, leading to severe resource contention or low resource utilization. The method in this embodiment, through multi-dimensional resource monitoring and intelligent scheduling based on graph coloring, fully considers the resource competition relationships between build tasks, maximizing system resource utilization and improving build efficiency while ensuring build quality.
[0160] In one optional implementation, a configuration scan is performed on the RPM base package, a depth-first search algorithm is applied to extract the configuration item hierarchy, a service parameter table is generated, and a standardized service control instruction set is constructed according to the service parameter table, including:
[0161] The configuration files in the RPM base package are preprocessed to convert the configuration information into key-value pairs to obtain the configuration content.
[0162] The configuration content is traversed using a depth-first search algorithm. A tree structure of configuration items is constructed with the configuration group as the root node. The parent-child relationship and dependency relationship between configuration items are identified. The association relationship of sibling configuration items is extracted. The hierarchical depth information of configuration items is recorded. The namespace mapping of configuration items is established to obtain the hierarchical structure of configuration items.
[0163] A service parameter table is generated based on the configuration item hierarchy to record the attribute information of the service parameters;
[0164] Based on the service parameter table, define the basic operation instructions, parameter configuration instructions, and status monitoring instructions for the service, establish instruction combination rules and execution order constraints, and generate a standardized service control instruction set.
[0165] In one specific implementation, processing the configuration files of the RPM base package is a crucial step in ensuring the correct deployment and operation of the service. When preprocessing the configuration files in the RPM base package, the configuration file content needs to be read and converted into key-value pairs. Configuration files may exist in various formats, such as ini, xml, json, and yaml. Appropriate parsers are used to parse the configuration files for different formats. For example, for ini format configuration files, the java.util.Properties class can be used to load the configuration content; for xml format, a DOM or SAX parser can be used; for json format, JsonParser can be used; and for yaml format, the SnakeYAML library can be used. During the parsing process, all configuration information is converted into standard key-value pairs, where the key is the full path of the configuration item, and the value is the specific value of the configuration item. For example, for xml format configuration fragments... <server> <port> 8080< / port> < / server> The converted key-value pair is "server.port=8080". For hierarchical configurations, dot notation is used to connect the names of each level to form a complete path. For example, the YAML-formatted configuration `server: {port: 8080, timeout: 30}` is converted to two key-value pairs: "server.port=8080" and "server.timeout=30". Through preprocessing, configuration content of different formats is unified into a standard set of key-value pairs, facilitating subsequent processing.
[0166] After preprocessing to obtain the configuration content, a depth-first search algorithm is used to traverse the configuration content and construct a tree structure of configuration items. The depth-first search begins with the configuration group as the root node. The configuration group can be identified by the configuration file name or the top-level configuration item in the configuration file. For example, for application configuration, "application" can be used as the root node; for database configuration, "database" can be used as the root node. During the construction of the tree structure, the parent-child relationship between configuration items is identified by the prefix relationship of the keys. For example, the common prefix "server" of "server.port" and "server.timeout" indicates that they have the same parent node "server". The depth information of the nodes is recorded during the construction process, with the root node having a depth of 0, and the depth increasing by 1 for each additional level of child nodes. In the previous example, the "server" node has a depth of 1, and the "server.port" and "server.timeout" nodes have a depth of 2. During the traversal, the dependency relationships between configuration items also need to be identified. Dependencies can be determined by the reference relationships in the configuration values. For example, if a configuration item value contains an expression of the form "${other.config}", it indicates that the configuration item depends on the "other.config" configuration item. The process involves traversing the system to extract relationships between sibling configuration items, which are those with the same parent node. These relationships can be identified through naming patterns or value similarities. For example, "jdbc.url", "jdbc.username", and "jdbc.password" share the prefix "jdbc" and collectively describe database connection information, making them related configuration items. In the tree structure built using depth-first search, each node contains the configuration item's key, value, depth, parent node reference, list of child nodes, list of dependent configuration items, and list of related configuration items. For this tree structure, a namespace mapping is established, mapping the full path to a specific namespace. For instance, "server.http.port" is mapped to the "http.server.port" namespace, facilitating configuration compatibility across different systems.
[0167] Based on the constructed configuration item hierarchy, a service parameter table is generated to record the attribute information of service parameters. The service parameter table is a structured data collection containing attributes such as parameter name, parameter type, default value, value range, required status, and description. The parameter name is directly taken from the configuration item's key; the parameter type is inferred from the configuration item's value, such as a pure numeric value inferred as an integer or floating-point type, a true / false value inferred as a boolean type, and others inferred as a string type; the default value is extracted from the configuration content; the value range is determined according to the parameter type and business rules, such as the port parameter's value range being 1024-65535; required status is determined based on the configuration item's comments or special markers in the configuration file; the description information is extracted from the configuration file's comments. The generated service parameter table is presented in tabular form. For example, the attribute record for the parameter "server.port" is: Name="server.port", Type="Integer", Default Value="8080", Value Range="1024-65535", Required="Yes", Description="Service Listening Port". For parameters with dependencies, add a "dependent parameter" column to the parameter table to record the names of the dependent parameters. For parameters with associations, add a "associate parameter group" column to record the group names of the associated parameters. The service parameter table provides the basic data support for the subsequent generation of service control commands.
[0168] Based on the service parameter table, define basic operation commands, parameter configuration commands, and status monitoring commands for the service. Establish command combination rules and execution order constraints to generate a standardized service control command set. Basic operation commands include start, stop, restart, and reload, corresponding to the service lifecycle management. For example, the start command "start-service" corresponds to the execution script "service-name start". Parameter configuration commands are used to modify service parameters, using the "set-param" pattern. For example, "set-param server.port 8080" sets the service port. For different types of parameters, define corresponding validation rules; for example, integer parameters need to be validated for numerical range, and string parameters need to be validated for format validity. Status monitoring commands are used to query service status and parameter values. For example, "get-status" retrieves the service running status, and "get-param server.port" retrieves the service port configuration. When establishing command combination rules, the logical relationships between commands must be considered. For example, if a service restart is required for parameter modifications to take effect, the "set-param" command should be followed by the "restart-service" command. Execution order constraints define the order in which commands are executed. For example, necessary parameters must be configured before starting a service, and the configuration should be backed up before stopping a service. Based on combination rules and execution order constraints, a standardized service control command set containing single and compound commands is generated. Single commands directly correspond to basic operations, parameter configuration, or status monitoring; compound commands are composed of multiple single commands combined in a specific order, such as the "update-config-and-restart" compound command which includes two steps: updating the configuration file and restarting the service. The standardized service control command set is saved in JSON or YAML format as configuration input for service control scripts, facilitating automated deployment and operation and maintenance tools to use it.
[0169] By processing the configuration files in the RPM base package using the method in this embodiment, the structured management of configuration information and the standardized definition of service control are realized. This provides a reliable configuration processing mechanism for building Gradle-based software RPM packages and improves the automation level of software deployment and operation.
[0170] The Gradle-based software RPM package building system of this invention includes:
[0171] The first unit is used to receive the package path and RPM package build parameters of the Java Web project and parse out the target processor architecture type.
[0172] The second unit is used to construct a two-layer cache based on the time-series eviction algorithm, calculate the frequency of use of construction dependencies to generate a feature sequence, build a data channel between the local cache and the shared cache, and form a set of construction dependency resources.
[0173] The third unit is used to initialize the Gradle environment using the build dependency resource set and generate a build task configuration list.
[0174] The fourth unit is used to calculate the parallelism parameter based on the build task configuration list using the resource competitor algorithm, split the program package into build task units according to the parallelism parameter, execute parallel build operations, and output the RPM base package of the target architecture.
[0175] The fifth unit is used to perform configuration scanning on the RPM base package, apply a depth-first search algorithm to extract the configuration item hierarchy, generate a service parameter table, and build a standardized service control instruction set according to the service parameter table;
[0176] The sixth unit is used to integrate the standardized service control instruction set into the RPM base package to build the RPM target package;
[0177] Unit 7 is used to perform RPM target package installation tests in an isolated environment, and outputs the final RPM package after successful verification.
[0178] A third aspect of the present invention provides an electronic device, comprising:
[0179] processor;
[0180] Memory used to store processor-executable instructions;
[0181] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0182] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0183] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for building software RPM packages based on Gradle, characterized in that, include: Receive the package path and RPM package build parameters of the Java Web project, and parse out the target processor architecture type; A two-layer cache is constructed based on the time-series eviction algorithm. The frequency of use of construction dependencies is calculated to generate a feature sequence. A data channel is built between the local cache and the shared cache to form a set of construction dependency resources. Initialize the Gradle environment using the build dependency resource set and generate a list of build task configurations; Based on the build task configuration list, the parallelism parameter is calculated using the resource competitor algorithm. The program package is then split into build task units according to the parallelism parameter, parallel build operations are executed, and the RPM base package of the target architecture is output. Perform a configuration scan on the RPM base package, apply a depth-first search algorithm to extract the configuration item hierarchy, generate a service parameter table, and build a standardized service control instruction set based on the service parameter table; Integrate the standardized service control instruction set into the RPM base package to build the RPM target package; Perform the RPM target package installation test in an isolated environment, and output the final RPM package after successful verification.
2. The method according to claim 1, characterized in that, Constructing a two-level cache based on a time-ordered replacement algorithm includes: A local cache area is established using off-heap memory, a shared cache area is established using a distributed file system, and a data channel is constructed between the local cache area and the shared cache area; The frequency of use of the build dependency is statistically analyzed, and the time decay value of the build dependency is calculated based on the usage time. The time decay value is multiplied by the usage frequency to obtain the usage weight of the build dependency. Slide the sampling window within a preset time period to record the usage time interval of the build dependencies and generate a feature sequence of the build dependencies. The resource activity score of the build dependency is calculated based on the weight and the feature sequence, and a first cache threshold and a second cache threshold are set based on the resource activity score. When the resource activity score of a build dependency is lower than the first cache threshold, the build dependency is moved out of the local cache; when the resource activity score is lower than the second cache threshold, the build dependency is moved to the shared cache.
3. The method according to claim 2, characterized in that, The resource activity score of the build dependency is calculated based on the weights used and the feature sequence. The first cache threshold and the second cache threshold are set based on the resource activity score, including: A base score is set based on the weight used, and a fluctuation range is calculated based on the feature sequence. The product of the base score and the fluctuation range is used as the resource activity score of the corresponding construction dependency. Collect a sample set of resource activity scores for dependencies built within a preset time window, and normalize the data to obtain a standard distribution sequence. Calculate the expected value and variance of the resource activity score corresponding to the dependency item based on the standard distribution sequence; obtain the total capacity and current occupancy rate of the cached resources, and calculate the capacity adjustment coefficient based on the current occupancy rate; The threshold base is determined by weighting the expected value, the variance value, and the capacity adjustment coefficient. A first cache threshold is set based on the product of the threshold base and a preset high-order decay factor to trigger local cache eviction; a second cache threshold is set based on the product of the threshold base and a preset low-order decay factor to trigger migration to a shared cache. The system monitors the current occupancy rate of cached resources in real time. When the current occupancy rate exceeds a preset alarm value, the capacity adjustment coefficient is dynamically adjusted according to a preset increment. The resource activity score is updated based on the update cycle to determine the latest capacity adjustment coefficient, and the first cache threshold and the second cache threshold are calibrated based on the latest capacity adjustment coefficient.
4. The method according to claim 1, characterized in that, Building data channels includes: Receive build dependency migration requests from the local cache; Obtain the version identifier of the build dependency to be migrated, and check the data integrity of the build dependency; Build dependencies are grouped according to data block size to generate a data transmission queue. A mapping table is established between the data transmission queue and the shared cache. Build dependencies in the data transmission queue are written to the corresponding storage location in the shared cache according to the mapping table. Record the storage status of the build dependencies in the shared cache and generate a resource location index; The build dependencies in the local cache and the shared cache are integrated to form a build dependency resource set.
5. The method according to claim 1, characterized in that, Based on the build task configuration list, the parallelism parameter is calculated using the resource competitor algorithm. The program package is then split into build task units according to the parallelism parameter, and parallel build operations are executed. The output RPM base package for the target architecture includes: Parse the build task configuration list to obtain package dependency information, generate a dependency relationship topology graph based on the package dependency information, traverse the dependency relationship topology graph to calculate the build complexity weight of each package, and output the initial build task sequence according to the build complexity weight; Collect computational resource usage data of the construction environment, construct a competition matrix based on resource usage overlap rate and perform eigenvalue decomposition to determine resource competition factors, and optimize the initial construction task sequence using graph coloring algorithm to obtain parallelism parameters; Based on the parallelism parameter, the file dependencies in the package are scanned to obtain the dependency set. The dependency set is analyzed to identify the smallest independent compilation unit and the smallest independent compilation unit is combined to form a build task unit. The execution priority sequence is obtained by prioritizing the build task units, and the build task units are allocated to the build queue to perform parallel build operations according to the execution priority sequence; The execution status of the construction task unit is collected in real time. When the execution status shows an abnormality, the task is retried and the task that has exceeded the number of retries is migrated to a new computing node. The execution progress of the construction task unit is periodically obtained. When the execution progress meets the preset performance decay threshold, the parallelism parameter is recalculated and the task is reallocated. The integrity and consistency of the build artifacts of the build task unit are checked to obtain the check results. Based on the check results, the build artifacts are organized to generate the RPM base package of the target architecture.
6. The method according to claim 5, characterized in that, The computational resource usage data of the construction environment is collected. A competition matrix is constructed based on the resource usage overlap rate, and eigenvalue decomposition is performed to determine the resource competition factor. The initial construction task sequence is optimized using a graph coloring algorithm to obtain parallelism parameters, including: Collect CPU utilization data, memory usage data, disk I / O bandwidth utilization data, and disk space remaining data in the construction environment, and perform normalization processing to obtain computing resource usage data; The computational resource usage data is converted into a competition matrix. The resource usage overlap rate of the construction task in each data dimension is calculated and weighted summation is performed to obtain the competition coefficient. The competition coefficient is used to fill the corresponding element positions of the competition matrix. Based on the construction complexity weight, the competition matrix is decomposed into eigenvalues, and the eigenvector corresponding to the largest eigenvalue is taken to obtain the resource competition factor. The construction tasks are treated as vertices in a graph. When the resource contention factor between task pairs is greater than a preset contention threshold, an edge construction conflict graph is established between the corresponding vertices. The vertices are sorted in descending order according to the size of the resource contention factor of the construction tasks. An integer set of available color numbers is established, starting from zero and increasing. Starting from the vertex with the highest degree, the smallest color number that has not been used by adjacent vertices is assigned to each vertex. The total number of colors used is recorded as a baseline value. The parallelism parameter is obtained by dynamically adjusting the baseline value in combination with the current load status.
7. The method according to claim 1, characterized in that, A configuration scan is performed on the RPM base package, and a depth-first search algorithm is applied to extract the configuration item hierarchy, generate a service parameter table, and construct a standardized service control instruction set based on the service parameter table, including: The configuration files in the RPM base package are preprocessed to convert the configuration information into key-value pairs to obtain the configuration content. The configuration content is traversed using a depth-first search algorithm. A tree structure of configuration items is constructed with the configuration group as the root node. The parent-child relationship and dependency relationship between configuration items are identified. The association relationship of sibling configuration items is extracted. The hierarchical depth information of configuration items is recorded. The namespace mapping of configuration items is established to obtain the hierarchical structure of configuration items. A service parameter table is generated based on the configuration item hierarchy to record the attribute information of the service parameters; Based on the service parameter table, define the basic operation instructions, parameter configuration instructions, and status monitoring instructions for the service, establish instruction combination rules and execution order constraints, and generate a standardized service control instruction set.
8. A Gradle-based software RPM package building system for implementing the method of any one of claims 1-7, characterized in that, include: The first unit is used to receive the package path and RPM package build parameters of the Java Web project and parse out the target processor architecture type. The second unit is used to construct a two-layer cache based on the time-series eviction algorithm, calculate the frequency of use of construction dependencies to generate a feature sequence, build a data channel between the local cache and the shared cache, and form a set of construction dependency resources. The third unit is used to initialize the Gradle environment using the build dependency resource set and generate a build task configuration list. The fourth unit is used to calculate the parallelism parameter based on the build task configuration list using the resource competitor algorithm, split the program package into build task units according to the parallelism parameter, execute parallel build operations, and output the RPM base package of the target architecture. The fifth unit is used to perform configuration scanning on the RPM base package, apply a depth-first search algorithm to extract the configuration item hierarchy, generate a service parameter table, and build a standardized service control instruction set according to the service parameter table; The sixth unit is used to integrate the standardized service control instruction set into the RPM base package to build the RPM target package; Unit 7 is used to perform RPM target package installation tests in an isolated environment, and outputs the final RPM package after successful verification.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Application program packaging method and device, equipment and storage medium
CN117055969A
Cache detection method of dependent file and related device
CN117555543A