Package installation processing method, device, equipment, medium and product

CN122777142APending Publication Date: 2026-09-18BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610942917.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-18

AI Technical Summary

Benefits of technology

[0008] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program or instructions that, when executed by a processor, implement the steps of the described method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122777142A_ABST
    Figure CN122777142A_ABST
Patent Text Reader

Abstract

This disclosure provides a method, apparatus, device, medium, and product for processing an installation package, relating to the fields of computer technology and software technology. The processing of the installation package includes: responding to a received processing request for the installation package, performing structural parsing on the installation package to obtain a file content structure, wherein the file content structure represents the relationships between multiple files in the installation package and the content characteristics of each file; generating a reuse identifier based on the configuration information indicated in the processing request and the file content structure, wherein the reuse identifier is used to match with a mapping to obtain a matching result, the matching result indicating whether a reusable installation package or reusable file corresponding to the processing request exists; and processing the installation package according to the processing method corresponding to the matching result to obtain a target installation package.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of computer technology and software technology, and more specifically, to a method, apparatus, device, medium, and product for processing an installation package. Background Technology

[0002] During the distribution and deployment of applications, the installation package typically adopts a full repetitive execution work mode. Full repetitive execution treats the installation package processing as an indivisible full atomic operation. Whenever there is a new build, hardening, or channel release task, the entire installation package needs to be re-executed through a complete series of unpacking, scanning, hardening, packaging, and signing processes. Summary of the Invention

[0003] In view of this, the present disclosure provides a method, apparatus, device, medium, and product for processing installation packages.

[0004] According to one aspect of this disclosure, a method for processing an installation package is provided, comprising: responding to receiving a processing request for the installation package, performing structural parsing on the installation package to obtain a file content structure, wherein the file content structure represents the association relationship between multiple files in the installation package and the content characteristics of each file; generating a reuse identifier based on configuration information indicated by the processing request and the file content structure, wherein the reuse identifier is used to match with a mapping to obtain a matching result, the matching result indicating whether there is a reusable installation package or reusable file corresponding to the processing request; and processing the installation package according to a processing method corresponding to the matching result to obtain a target installation package.

[0005] According to another aspect of this disclosure, an installation package processing apparatus is provided, comprising: a structure parsing module, configured to, in response to receiving a processing request for the installation package, perform structure parsing on the installation package to obtain a file content structure, wherein the file content structure characterizes the association relationship between multiple files in the installation package and the content characteristics of each file; a first generation module, configured to generate a reuse identifier based on configuration information indicated by the processing request and the file content structure, wherein the reuse identifier is used to match with a mapping to obtain a matching result, the matching result indicating whether there is a reusable installation package or reusable file corresponding to the processing request; and a processing module, configured to process the installation package according to a processing method corresponding to the matching result to obtain a target installation package.

[0006] According to another aspect of this disclosure, an electronic device is provided, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0007] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores a computer program or instructions thereon, which, when executed by a processor, implement the steps of the above-described method.

[0008] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program or instructions that, when executed by a processor, implement the steps of the described method.

[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0011] Figure 1 The illustration schematically shows a system architecture to which an installation package processing method can be applied according to an embodiment of the present disclosure;

[0012] Figure 2 A flowchart illustrating a method for processing an installation package according to an embodiment of the present disclosure is shown schematically.

[0013] Figure 3 An example schematic diagram of a front-end interface according to an embodiment of the present disclosure is shown;

[0014] Figure 4A This illustration schematically shows an example of the process of performing structural parsing on an installation package to obtain the file content structure according to an embodiment of the present disclosure;

[0015] Figure 4B This illustration schematically shows an example diagram of a document content structure according to an embodiment of the present disclosure;

[0016] Figure 5 This illustration schematically shows an example of a process for generating a reuse identifier based on configuration information and file content structure indicated by a processing request, according to an embodiment of the present disclosure.

[0017] Figure 6 This illustration shows an example of a process in which an installation package is processed according to a processing method corresponding to a matching result to obtain a target installation package, according to an embodiment of the present disclosure.

[0018] Figure 7 A block diagram schematically illustrates a processing apparatus for an installation package according to an embodiment of the present disclosure; and

[0019] Figure 8 A block diagram of an electronic device suitable for implementing a package processing method according to an embodiment of the present disclosure is shown schematically. Detailed Implementation

[0020] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.

[0021] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0022] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0023] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0024] In the full-scale repetitive execution mode, because installation packages with the same or similar content, unpacked structured code files, resource files, etc. are repeatedly unpacked and hardened during multiple builds, channel packaging, or version rollbacks, there is a lack of reuse mechanism, resulting in redundant calculations and resource waste.

[0025] Furthermore, due to the high complexity of the hardening algorithm and the differences in the size of different installation packages, the execution time of a single task fluctuates greatly, and the processing latency is difficult to predict, which affects the stability of the automated installation package release pipeline.

[0026] Therefore, this disclosure proposes a processing scheme for an installation package. For example, in response to receiving a processing request for an installation package, the installation package is structurally parsed to obtain a file content structure, wherein the file content structure represents the association relationship between multiple files in the installation package and the content characteristics of each file; based on the configuration information indicated by the processing request and the file content structure, a reuse identifier is generated, wherein the reuse identifier is used to match with a mapping to obtain a matching result, and the matching result indicates whether there is a reusable installation package or reusable file corresponding to the processing request; the installation package is processed according to the processing method corresponding to the matching result to obtain the target installation package.

[0027] According to embodiments of this disclosure, by performing structural parsing on the installation package, unstable binary installation packages can be transformed into standardized file content structures, eliminating non-functional noise and ensuring that subsequent judgments on whether installation packages or files are identical are based on functional equivalence. By generating reuse identifiers based on configuration information and the file content structure, and using reuse identifier matching mappings to determine the existence of reusable content, the optimal processing method can be selected based on different matching results. This reduces resource waste caused by redundant calculations, shortens the average response time of installation package processing requests, and improves the processing efficiency of installation packages.

[0028] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution of this invention all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0029] In the technical solution of the present invention, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.

[0030] Figure 1 The illustration schematically depicts a system architecture to which an installation package processing method can be applied according to embodiments of the present disclosure. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.

[0031] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0032] Users can interact with server 105 via network 104 using at least one of the first terminal device 101, second terminal device 102, and third terminal device 103 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, second terminal device 102, and third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0033] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0034] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0035] It should be noted that the installation package processing method provided in this embodiment can generally be executed by server 105. Correspondingly, the installation package processing device provided in this embodiment can generally be located in server 105. The installation package processing method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the installation package processing device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0036] Alternatively, the installation package processing method provided in this embodiment of the present disclosure can also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or by other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the installation package processing apparatus provided in this embodiment of the present disclosure can also be disposed in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or in other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103.

[0037] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0038] It should be noted that the sequence numbers of the operations in the following methods are for descriptive purposes only and should not be considered as indicating the execution order of the operations. Unless explicitly stated otherwise, the method does not need to be executed in the exact order shown.

[0039] Figure 2 A flowchart illustrating a method for processing an installation package according to an embodiment of the present disclosure is shown schematically.

[0040] like Figure 2 As shown, the processing method 200 for the installation package may include operations S210 to S230.

[0041] In operation S210, in response to receiving a processing request for the installation package, the installation package is structurally parsed to obtain the file content structure, wherein the file content structure represents the relationship between multiple files in the installation package and the content characteristics of each file.

[0042] In operation S220, a reuse identifier is generated based on the configuration information and file content structure indicated by the processing request. The reuse identifier is used to match with the mapping to obtain a matching result. The matching result indicates whether there is a reusable installation package or reusable file corresponding to the processing request.

[0043] In operation S230, the installation package is processed according to the processing method corresponding to the matching result to obtain the target installation package.

[0044] In the embodiments of this disclosure, the Application Programming Interface (API) gateway can be monitored in real time. If a user or a Continuous Integration (CI) / Continuous Deployment (CD) pipeline submits a task containing an installation package via the API, the installation package processing flow can be initiated. The installation package is a binary application package (Android Package Kit, APK) file submitted by the user that requires security hardening, signing, or repackaging. After the installation package is decompressed, a global configuration file, an executable file, compiled binary resource files, a signing file, resource files, configuration files, etc., can be obtained.

[0045] Structural parsing refers to the process of disassembling and analyzing a binary installation package according to its internal standard format, extracting key components, and ultimately organizing it into a standardized and normalized file content structure. The purpose of structural parsing is to eliminate noise caused by non-functional differences such as compression methods, build times, and file ordering, and to generate a file content structure that reflects its functional components.

[0046] The file content structure is the output obtained after parsing the installation package. The file content structure can include the relationship between multiple files in the installation package and the content characteristics of each file. It is used to describe the hierarchical relationship and core file composition within the installation package, and removes unstable factors such as timestamps from the original package body.

[0047] The specific method of structure parsing can be configured according to actual business needs and is not limited here. For example, an APK parsing tool can be used to decompress the installation package into multiple files, and then these files can be normalized, such as converting all file paths to lowercase. The file content structure can be generated by calculating the hash values ​​of the .dex and .so files. Alternatively, when the installation package is large, it can be divided into blocks, and different blocks can be decompressed in parallel using a distributed computing framework, and finally aggregated and normalized.

[0048] Configuration information refers to the customized requirements specified by the user when submitting the installation package, in addition to uploading the installation package itself. Configuration information can define the scope, strength, and objectives of the processing applied to the installation package. For example, configuration information could be "Obfuscate only file A, do not harden file B," or "Apply advanced encryption to file C," or "Use signature scheme D and harden file E to prevent repackaging," etc.

[0049] After obtaining the file content structure, a reuse identifier can be generated based on the configuration information and the file content structure. This reuse identifier can be used to match mappings to determine whether the same or similar processing requests have been executed. The mapping is essentially an index database, which maintains the correspondence between candidate identifiers and candidate objects.

[0050] The specific method for generating the reuse identifier can be configured according to actual business needs and is not limited here. For example, the content hash of the core file can be extracted from the file content structure, concatenated into a string, and then combined with the hash value of the configuration information to generate the final reuse identifier through secondary hashing. Alternatively, Locality-Sensitive Hashing (LSH) or Fuzzy Hashing algorithms can be used. In this case, even if the code of two installation packages is slightly different, as long as the similarity exceeds the threshold, they can still be matched to the same approximate reusable file or reusable installation package.

[0051] A reusable installation package is the final signed product stored after all hardening, packaging, and signing processes have been completed for installation packages with the same content structure and processing strategies during historical processing. Reusable files are intermediate products generated and stored during historical processing. They can be the result of hardening / processing some files in the reusable installation package and can be reused in subsequent repackaging processes, thereby skipping the time-consuming hardening calculation stage.

[0052] After matching the mapping using the reuse identifier, the optimal processing method can be selected from multiple candidate processing methods to execute the processing request based on the hit status of the reuse identifier in the mapping, thereby minimizing redundant calculations to obtain the target installation package. The target installation package is the final version of the installation package delivered to the user after processing by the installation package processing method provided in this disclosure. It is a hardened, standardized, reorganized, signed, and integrity-verified installable installation package.

[0053] In the embodiments of this disclosure, by performing structural parsing on the installation package, unstable binary installation packages can be transformed into standardized file content structures, eliminating non-functional noise and ensuring that subsequent judgments on whether installation packages or files are identical are based on functional equivalence. By generating reuse identifiers based on configuration information and file content structures, and using reuse identifier matching mapping to determine the existence of reusable content, the optimal processing method can be selected based on different matching results. This reduces resource waste caused by redundant calculations, shortens the average response time of installation package processing requests, and improves the processing efficiency of installation packages.

[0054] According to embodiments of this disclosure, the processing request can be submitted through a front-end interface, which is used to receive configuration information and the source information of the installation package. The front-end interface is a visual operation window provided to the user for human-computer interaction. The front-end interface allows the user to input data, issue commands, and provide feedback on task progress and results. The front-end interface may include an installation package upload box, checkboxes for configuration options, a task start button, and a status bar displaying task progress, etc.

[0055] For example, when a user clicks the "Start Hardening" button on a webpage, this click constitutes a processing request. This processing request can instruct the processor to "generate a processed installation package for me using the installation package I uploaded and the configuration information I entered."

[0056] Source information refers to data indicating the location of the installation package that needs to be processed. For example, source information can be a directly uploaded file stream, where the user selects a local file and submits it via the "Upload" button on the front-end interface. Alternatively, source information can also be an object storage address, where the user enters a Uniform Resource Locator (URL) on the interface, and the installation package will be downloaded based on this address.

[0057] In response to the verification of source information and configuration information, a request identifier for processing the request is generated; the request identifier and processing status are displayed through the front-end interface, where the processing status represents the execution status corresponding to the current execution stage of the processing request.

[0058] After obtaining the source and configuration information, these can be validated. This includes checking if the installation package's source is reachable, if the file format is supported, and if the parameters in the configuration are valid and do not conflict with each other. If all validations pass, the processing request can be accepted, and a unique request identifier (taskId) can be created for it, usable for full lifecycle traceability. If any validations fail, a request identifier can be omitted, and a failure reason can be returned.

[0059] A request identifier is a globally unique and unforgeable string assigned to a processing request after it has been successfully accepted. It serves as the representation of the processing request throughout its entire lifecycle and is used to associate all subsequent operations, queries, and results of the request.

[0060] After a request is accepted, its processing status can be pushed to the front end and displayed to the user in real time or at time intervals. The processing status is a textual or graphical description of the current stage of the entire process from start to finish, reflecting the progress of the request in the pipeline. For example, the processing status may include: request accepted, structure parsing in progress, calculating reuse identifiers, processing in progress, mapping hit and skipping processing, repackaging in progress, signature verification in progress, processing completed and downloadable, and processing failed and the reason can be viewed.

[0061] The specific display method for the processing status can be configured according to actual business needs and is not limited here. For example, the front-end interface can be more than just simple text, but a dynamic visual pipeline diagram showing nodes such as "structure parsing -> identifier calculation -> processing -> packaging -> signing". When the processing request reaches a certain node, that node is highlighted and displays an animation of "in progress", turns green upon completion to "completed", and turns red upon failure, etc.

[0062] Alternatively, in addition to displaying the processing status, users can expand to view the detailed execution status of that stage. For example, the "Processing Execution Status" will not only show "Processing in progress" but also "Processing file 2 / 5, estimated time remaining 30 seconds." If a stage fails, its "Processing Status" can not only display "Failed" but also provide an interactive button, such as "Retry this stage," "Re-execute after adjusting hardening parameters," or "Ignore risks and continue," allowing users to manually intervene based on fine-grained execution status.

[0063] In the embodiments of this disclosure, by generating a request identifier after verifying the source information and configuration information, it is ensured that the identifier can be assigned after verification, preventing invalid requests from being created due to invalid source information or conflicting configuration information, thereby saving subsequent computation and scheduling resources. Based on this, displaying the processing status through a front-end interface transforms the highly complex multi-stage pipeline operation into user-understandable status information, allowing users to know the request processing progress in real time. This achieves standardized request submission, effective filtering of acceptance, and transparency of the processing process, thereby improving the success rate of request processing and the efficiency of user-side response interaction.

[0064] Figure 3 An example schematic diagram of a front-end interface according to an embodiment of the present disclosure is shown.

[0065] like Figure 3 As shown, in embodiment 300 where a processing request is submitted via a front-end interface, the front-end interface 310 may include a selection control 320, a source indicator control, and a configuration indicator control.

[0066] Regarding the selection control 320, after the user clicks on the selection control 320, various candidate processing solutions can be displayed through the display area 322 of the selection control 320. For example, candidate processing solution 1, candidate processing solution 2, ..., candidate processing solution M, where M is a positive integer. In addition, the display area 322 of the selection control 320 can be scrolled to display other candidate processing solutions.

[0067] In response to detecting a user's selection of a target processing scheme from the candidate processing schemes, the target processing scheme can be displayed in drop-down box 321. For example, if the user selects "Candidate Processing Scheme 2" as the target processing scheme, then the target processing scheme can be displayed in drop-down box 321.

[0068] The source indicator control can include a dropdown source indicator control 331 and a path selection source indicator control 332. For example, after a user clicks the source indicator control 331, various candidate sources can be displayed in the display area of ​​the source indicator control 331. In response to detecting a user's selection of a target source among the candidate sources, the target source can be displayed in the dropdown list of the source indicator control 331. Alternatively, after a user clicks the source indicator control 332, the target source can be specified by selecting a hierarchical path.

[0069] The configuration indicator control can include a dropdown menu-style configuration indicator control 341 and a path selection-style configuration indicator control 342. For example, after a user clicks the configuration indicator control 331, various candidate configurations can be displayed in the display area of ​​the configuration indicator control 341. In response to detecting a user's selection of a target configuration among the candidate configurations, the target configuration can be displayed in the dropdown menu of the configuration indicator control 341. Alternatively, after a user clicks the configuration indicator control 342, the target configuration can be specified by selecting a hierarchical path.

[0070] According to an embodiment of this disclosure, operation S210 may include the following operations: among a plurality of candidate files obtained by standardizing the installation package, the candidate files whose influence on the functionality of the installation package is greater than a predetermined threshold are identified as files; and a file content structure is generated based on the content characteristics of each file.

[0071] Standardization is the process of decompressing, disassembling, and formatting the raw, binary installation package. Standardization eliminates non-essential differences caused by different build environments and packaging tools, presenting the package content in a uniform format. For example, decompression tools can be used to extract files and directories from the package. During this process, path separators can be uniformly converted to forward slashes, file timestamps and other metadata differences can be ignored, and core code files and resource files can be separated.

[0072] Candidate files are all the individual files extracted from the standardized installation package. After obtaining all candidate files from the decompressed installation package, those files that truly define the application's behavior can be identified and selected based on predetermined thresholds. Files to be processed are those selected from the candidate files that are deemed to have a substantial impact on the core functionality of the installation package. For example, since any modification to dex, so, or manifest type files directly affects the application's functionality, candidate files may include these types of files.

[0073] The method for determining files from candidate files can be configured according to actual business needs and is not limited here. For example, core file type rules can be pre-configured, and during the filtering process, all candidate files are traversed. Files that match these rules are identified as files to be processed. Alternatively, a classification model can be pre-trained, taking the pathname, file header, and dependencies of candidate files as input, and outputting the probability that the file affects the functionality of the installation package. The core code and resource files with the highest probability are then identified as files to be processed.

[0074] Once the files are identified, a file content structure that reflects their organization and essential attributes can be created. Content characteristics are computable numerical or string representations used to uniquely identify the attributes and content of a file to be processed. Content identifiers can be used to determine whether two files are functionally equivalent. For example, content characteristics may include strings calculated using the Secure Hash Algorithm 256-bit (SHA256) algorithm, file type, file size, package name, and version number.

[0075] In the embodiments of this disclosure, by filtering files among multiple candidate files, filtering conditions that affect functionality are introduced. This proactively eliminates noisy files unrelated to the core application logic. Since the subsequently generated file content structure discards non-functional differences, two installation packages with the same functionality but different build environments can be abstracted into the same file content structure, avoiding duplicate processing caused by byte-level differences and improving the hit rate of identifying and reusing identical or similar installation packages. Furthermore, by generating file content structures based on the content characteristics of each file, the specific location level of a file can be bound to features that uniquely identify its content. This allows for determining compositional similarity based on structure and functional consistency based on content features, enabling quick and accurate retrieval of reusable products in the mapping, thus improving the efficiency and accuracy of installation package processing.

[0076] According to embodiments of this disclosure, content features can be obtained by: filtering out construction noise in a file based on preset fields to obtain an intermediate file, wherein construction noise is content generated during the construction of the installation package that causes differences between bytes but does not affect functionality; and generating content features based on the content of the intermediate file.

[0077] After obtaining the file, the data stream can be matched against preset fields. If noisy data is found, a removal operation is performed to obtain an intermediate file. For example, removal operations may include zeroing, removal, or replacement with fixed values. Preset fields are predefined data identifiers, format rules, or byte ranges that indicate which unstable or functionally irrelevant data needs to be found, removed, or ignored from the original file before generating content features. For example, preset fields may include timestamps, builder information, signature blocks, and compression levels.

[0078] Build noise refers to non-deterministic data automatically injected into files during the compilation and packaging of an installation package by external factors such as build tools, environment, or time. Although build noise may manifest as differences at the byte level, it does not affect the application's runtime logic, interface, or functionality.

[0079] An intermediate file is a transitional, clean file generated after the original file has undergone content filtering. The intermediate file retains all the file's functional content but has had all build noise removed. For example, consider a classes.dex file whose header embeds a compilation timestamp. You can preset fields to locate the offset and length of this timestamp and fill them all with zero values. The resulting file is the intermediate file. This intermediate file is functionally equivalent to the original file.

[0080] After obtaining the intermediate file, which has eliminated all construction noise, a one-way summarization or feature extraction algorithm can be performed on it. Since the intermediate file has eliminated the construction noise in the original file, the calculated content features will match regardless of how many times the original file has been signed or when it was built, as long as the functionality remains unchanged. Therefore, the content features can serve as the unique functional identifier of the file.

[0081] Content features can be configured according to actual business needs and are not limited here. For example, the entire byte stream of the intermediate file can be read and passed as input to SHA256 or Message Digest Algorithm 5 (MD5) to calculate a fixed-length hash string, which serves as the content feature of the file. Alternatively, locality-sensitive hashing can be used, which involves reading the content of the intermediate file, extracting features using a sliding window approach, and generating hash values ​​that are resistant to minor differences. In this case, the file features can be hash vectors.

[0082] In the embodiments of this disclosure, by filtering content based on preset fields, non-functional content that fluctuates with the build environment can be extracted from files, thereby standardizing the input source. This eliminates the possibility of misidentifying files with the same functionality built at different times and on different machines as different files due to byte differences caused by external operations such as secondary signing and repackaging. Furthermore, by generating content features based on the content of intermediate files, these features can accurately focus on the essential functional parts such as code and resources, improving the hit rate of accurate identification and reuse of installation packages with the same functionality, and reducing the resource waste of repeatedly performing heavy computational tasks such as security hardening.

[0083] According to embodiments of this disclosure, generating a file content structure based on the content features of each file includes: constructing a hierarchical structure based on the association relationships between the files, wherein the hierarchical structure includes multiple nodes, and each node represents a file; associating the content features of each file with the corresponding nodes in the hierarchical structure to obtain the file content structure.

[0084] After obtaining each file, the association relationships for each file can be read. Association relationships refer to the ownership and organizational relationships of various files within the installation package, based on their storage paths. Association relationships can describe the directory location of a file within the installation package and the inclusion relationship between files and directories.

[0085] By parsing the path strings with hierarchical separators in the association relationships, root nodes, intermediate nodes, and leaf nodes are dynamically or logically created, and connected according to parent-child relationships to form a hierarchical structure. A hierarchical structure, or tree-like data structure, can transform association relationships into a three-dimensional model with hierarchical and parent-child relationships. In the hierarchical structure, the top level is the root node, which can represent the entire installation package, and below it are child nodes representing directories or direct files at each level. Nodes are the basic units that constitute the hierarchical structure. Each node represents a file in the installation package. After obtaining the hierarchical structure, pre-calculated content characteristics for the corresponding files can be used as attribute values ​​and bound to the corresponding nodes to form the file content structure.

[0086] The specific construction method of the hierarchical structure can be configured according to actual business needs and is not limited here. For example, all files can be traversed, the path can be split based on the delimiter, and directory nodes can be searched or created layer by layer starting from the root node. Under the directory nodes, file nodes can be created to form a tree structure. Alternatively, the hierarchical structure can be constructed as a directed graph. For example, the code of .dex and .so files can be analyzed to extract class calls, method calls, and library function dependencies across files, and each file can be a node with the entry point as the starting point, forming directed edges for the call relationships.

[0087] In the embodiments of this disclosure, by constructing a hierarchical structure based on relationships, the internal organizational logic of the installation package can be made explicit, thereby enabling the accurate restoration of the original file locations during subsequent repackaging based on the hierarchical structure. Furthermore, by injecting content features representing the essential function of each file in the hierarchical structure, the query speed and accuracy for reusable files or reusable installation packages are improved, while saving computing resources and storage overhead.

[0088] Figure 4A The illustration shows an example diagram of the process of performing structural parsing on an installation package to obtain the file content structure according to an embodiment of the present disclosure.

[0089] like Figure 4A As shown, in embodiment 400A of obtaining the file content structure, after obtaining the installation package 410, the installation package 410 can be standardized to obtain multiple candidate files, such as candidate file 421, candidate file 422, ..., candidate file 42P, where P is a positive integer. Among the multiple candidate files, candidate files that affect the functionality of the installation package 410 are selected to obtain file 431, file 432, ..., file 43Q, where Q is a positive integer.

[0090] After obtaining multiple files, content filtering can be performed on the files according to preset field 440 to eliminate construction noise and obtain intermediate files. For example, content filtering can be performed on file 431 according to preset field 440 to obtain intermediate file 451; content filtering can be performed on file 432 according to preset field 440 to obtain intermediate file 452; and so on, content filtering can be performed on file 43Q according to preset field 440 to obtain intermediate file 45Q.

[0091] After obtaining multiple intermediate files, content features can be generated based on the content of the intermediate files. For example, content feature 461 can be generated based on the content of intermediate file 451; content feature 462 can be generated based on the content of intermediate file 452; and so on, content feature 46Q can be generated based on the content of intermediate file 45Q.

[0092] After obtaining the content features of each file, a hierarchical structure can be constructed based on the relationship between the files; content features 461, 462, ..., 46Q are associated with the corresponding nodes in the hierarchical structure to obtain the file content structure 470.

[0093] Figure 4B An example schematic diagram illustrating a document content structure according to an embodiment of the present disclosure is shown.

[0094] like Figure 4BAs shown, in embodiment 400B of the file content structure, an example is given with the file content structure including four levels.

[0095] The first level of the file content structure may include node 480; the second level may include node 481 and node 482, which are subordinate nodes of node 480; the third level may include node 4811 and node 4812, which are subordinate nodes of node 4811; and the fourth level may include node 4811_1 and node 4811_2, which are subordinate nodes of node 4811, and node 4812_1, which are subordinate nodes of node 4812.

[0096] According to embodiments of this disclosure, the configuration information indicates at least one configuration item. A configuration item is a single processing request entry included in the configuration information, which can be specifically configured by the user or the system. The configuration item can be used to indicate which part of the installation package the current processing request will operate on and in what manner.

[0097] According to an embodiment of this disclosure, operation S220 may include the following operations: obtaining target file features and attribute information from the file content structure according to the configuration item, wherein the attribute information represents the attributes of the application corresponding to the installation package to be processed; and fusing the target file features, attribute information and configuration item to obtain a reuse identifier.

[0098] After obtaining the file content structure, configuration items can be used as query conditions to perform matching within the constructed file content structure. For successfully matched nodes, their attached content features are extracted as target content features; simultaneously, attribute information that does not change with the processing scope can also be extracted from the root node or list node of the file content structure.

[0099] Target content features are content features extracted from the constructed file content structure, based on configuration settings, specifically targeting the files that need to be processed. For example, if a configuration setting indicates hardening only SO libraries, then all file nodes with the .so extension in the hierarchical structure can be located in the file content structure, and their content features can be extracted as target file content features. Attribute information is metadata extracted from the constructed file content structure, representing the installation package as a whole and corresponding to the application's basic identity and functional declarations.

[0100] After obtaining the target content features and attribute information, these features and configuration items can be merged to generate an identifier that comprehensively reflects the input content and processing requirements. Target content features represent the content, attribute information represents the application, and configuration items represent the operation to be performed. The specific fusion method can be configured according to actual business needs and is not limited here.

[0101] For example, the feature list within the target can be sorted lexicographically and concatenated into a string using a delimiter; the policy identifier and toolchain version number can be extracted from the configuration items and serialized; the package name and version number in the attribute information can be concatenated; based on this, these three can be concatenated into the target string in a fixed order, and a SHA256 hash calculation can be performed on it to obtain the reuse identifier.

[0102] Alternatively, the target file features, attribute information, and configuration items can be encoded into a set of basis vectors, and then these basis vectors can be weighted and summed according to preset weights to obtain vector identifiers. During matching, cosine similarity can be calculated to find nearest neighbor requests that have similar core processing even if the configuration items are not exactly the same. For example, the preset weights can be 0.7 for target content features, 0.2 for attribute information, and 0.1 for configuration items.

[0103] In the embodiments of this disclosure, by obtaining target file features and attribute information based on configuration items, feature extraction can be focused on the processing intent, thereby improving the hit rate in targeted reuse scenarios. Furthermore, by fusing target file features, attribute information, and configuration items, the accuracy, security, and flexibility of reuse at the semantic level are ensured.

[0104] According to embodiments of this disclosure, fusing target content features, attribute information, and configuration items to obtain a reuse identifier may include the following operations: obtaining version information of the tool used to execute the processing request; fusing the version information, target content features, attribute information, and configuration items to obtain a reuse identifier.

[0105] During request execution, the software version numbers of each tool that will actually be invoked can be actively read and recorded to obtain version information. This version information can be determined by built-in version constants in the code, deployment package manifest files, or system environment variables.

[0106] Version information refers to the specific version identifiers of the various tools used to perform the package processing task. Version information records which version of the tool will actually be used when this processing request is processed. Tools are software programs or libraries that are invoked at various stages of the request processing flow to perform specific calculations, transformations, or operations. For example, tools may include an unpacking engine tool for structure parsing, a repacking engine tool for reassembly, and a signature verification tool for signing.

[0107] After obtaining the version information, we can determine which tools to use to generate the reuse identifier. For example, the obtained version information can be treated as a new independent segment, and then concatenated with the target content features, attribute information, and configuration items at fixed positions to form the final string. Then, its SHA256 hash can be calculated as the reuse identifier.

[0108] In the embodiments of this disclosure, by obtaining version information and fusing the version information with target content features, attribute information and configuration items, the reuse identifier becomes sensitive to the version of the processing tool, preventing security risks caused by incorrect reuse of defective or non-compliant processing results of old tools, and improving the security of installation package processing while ensuring processing efficiency.

[0109] According to embodiments of this disclosure, configuration items may include at least one of the following: whether to harden the installation package, whether to encrypt the installation package, whether to rewrite the data in the installation package, whether to reorganize the file, and whether to change the file format.

[0110] Whether to harden the installation package refers to whether to perform security enhancement operations on the entire installation package or parts thereof, making it difficult to reverse engineer, tamper with, or repackage. When this configuration option is enabled, modules such as program packing, code virtualization, and anti-tampering verification can be invoked to harden the selected files and generate hardened files.

[0111] Whether to encrypt the installation package refers to whether to perform cryptographic encryption operations on the data content of specific files within the installation package. When this configuration option is enabled, symmetric or asymmetric encryption algorithms can be performed on specified data files using package-level or application-level keys, and the decryption logic can be embedded in the application code or shell.

[0112] Whether to rewrite the data in the installation package refers to whether to perform string obfuscation, numerical constant replacement, control flow flattening, or inserting junk instructions on the data content of specific files within the installation package. When this configuration option is enabled, non-encrypted deterministic transformations can be performed on code or data to hide the original logic and data readability, increasing the difficulty of reverse engineering.

[0113] Whether to reorganize the files refers to whether the collection of files after disassembling the installation package is reorganized and repackaged according to new or specific rules. When this configuration option is enabled, the files constituting the package can be rearranged and repackaged according to a specified non-original order, alignment, or compression strategy before the final package is output.

[0114] Whether to change the file format refers to whether to convert a specific file within the installation package from one file format or encoding to another. When this configuration option is enabled, the corresponding format converter can be invoked to transcode the specified file from its original format or convert it to another optimized or standardized format.

[0115] In the embodiments of this disclosure, by providing configuration items such as hardening, encryption, and rewriting, users can precisely control the scope and intensity of processing, achieving refined processing on demand. By providing configuration items such as reorganization and format change, operations at the build engineering level, such as packaging structure optimization and resource format optimization, can be integrated with security processing within the same configuration system. This allows a single processing request to simultaneously complete security protection and delivery optimization, ensuring the independent identification and accurate reuse of products for different purposes. Thus, while ensuring security, it is possible to improve processing flexibility and the diversity of final products.

[0116] Figure 5 The illustration shows an example diagram of a process for generating a reuse identifier based on configuration information and file content structure indicated by a processing request, according to an embodiment of the present disclosure.

[0117] like Figure 5 As shown, in embodiment 500 of generating a reuse identifier, the configuration information in processing request 510 indicates at least one configuration item 511.

[0118] Based on configuration item 511, target content features 531 and attribute information 532 are obtained from the file content structure 520. Version information 540 for executing processing request 510 is obtained. Based on this, version information 540, target content features 531, attribute information 532 and configuration item 511 can be merged to obtain reuse identifier 550.

[0119] According to embodiments of this disclosure, the mapping can indicate the correspondence between candidate identifiers and the storage addresses of candidate objects. A candidate identifier is a reuse identifier corresponding to all historically successfully processed requests. A storage address is a physical or logical path pointer associated with a candidate identifier, pointing to a specific reusable installation package or reusable file in storage. The storage address is used to indicate from which location the corresponding artifact should be retrieved.

[0120] According to an embodiment of this disclosure, operation S230 may include the following operations: matching multiple candidate identifiers in the mapping according to the reuse identifier to obtain a matching result; and processing the installation package according to the processing method corresponding to the matching result to obtain the target object.

[0121] After obtaining the reuse identifier, the reuse identifier calculated in this processing request can be used as a query condition to perform an exact or approximate search among all candidate identifiers stored in the mapping. The matching result is the conclusion output after performing the matching, which may contain structured information such as the hit level, the type of the hit object, and the specific storage reference, to guide the selection of subsequent processing methods.

[0122] For example, if a candidate identifier that exactly matches the query conditions is found, the matching result is returned, including information such as the storage address associated with that candidate identifier; if no candidate identifier is found after traversing all candidate identifiers, a result indicating a match is returned.

[0123] The specific method for matching reused identifiers with candidate identifiers can be configured according to actual business needs and is not limited here. For example, both reused identifiers and candidate identifiers can be treated as high-dimensional vectors, and a vector database can be used to perform an approximate nearest neighbor search based on cosine similarity.

[0124] Alternatively, the prefix generated from the attribute information in the reuse identifier can be used for matching to quickly locate the candidate application family. Then, the target content feature part can be used for secondary matching within the family. Finally, the configuration item and version information parts can be used for fine matching. This progressive matching method can quickly lock the change point when there are partial misses.

[0125] In the embodiments of this disclosure, by matching the mapping according to the reuse identifier, it is possible to instantly determine whether the current processing request contains reusable content that has already been processed. Based on this, by processing according to the processing method corresponding to the matching result, the semantics of the matching result can be directly converted into an execution path, achieving installation package processing path diversion based on historical experience, and improving the targeting and efficiency of installation package processing.

[0126] According to embodiments of this disclosure, candidate objects may include at least one of candidate installation packages and candidate files. A candidate installation package is a complete installation package that has completed all processing steps and is stored in the Final Artifact Store, ready for direct installation and use. A candidate file is a processed file generated during a historical processing step and stored in the CAS Blob Store.

[0127] File storage is a storage system used to store candidate files. File storage can use the content hash of the candidate file as the key and the candidate file as the value. Installation package storage is a storage system used to store candidate installation packages. Installation package storage can use the reuse identifier of the candidate installation package as the key and the candidate installation package as the value.

[0128] The installation package to be processed is processed according to the processing method corresponding to the matching result to obtain the target object. This may include the following operations: if there is a candidate file corresponding to the reuse identifier in the mapping, the installation package to be processed is processed according to the candidate file obtained from the file storage according to the storage address to obtain the target installation package; if there is a candidate installation package corresponding to the reuse identifier in the mapping, the candidate installation package obtained from the installation package storage according to the storage address is used as the target installation package.

[0129] If a matching process reveals that a historical request has processed the same file, but no complete installation package is available due to differences in signature, other resources, or user requirements, then the processed candidate files can be retrieved directly from the file storage. These can be used as new intermediate products and, along with other files in the current installation package that do not require processing, are then fed into the subsequent repackaging and signing process to ultimately generate the target installation package.

[0130] Furthermore, when the candidate files to be retrieved are distributed across different storage nodes, multiple read requests can be initiated concurrently. The retrieved files can be passed to the reassembly engine in a streaming pipeline manner, allowing the reassembly engine to immediately begin calculating the file's index structure. This enables parallel data flow and computation, thereby reducing the reassembly time.

[0131] If the matching process finds that a previous request has processed the same installation package, then all calculation steps can be skipped, and the existing installation package can be read from the installation package storage according to the storage address, and then directly returned to the user as the target object for processing this request.

[0132] Furthermore, the storage address of the candidate installation package can be more than just an HTTP URL; it may also be a structured object with a digital signature or expected hash value. Before delivering the acquired candidate installation package as the target installation package, the integrity and signature of the acquired file stream can be verified locally or through a Trusted Execution Environment (TEE) to ensure that the files stored in the repository have not been tampered with or corrupted, and to consider them valid after the verification passes.

[0133] In the embodiments of this disclosure, by retrieving candidate files from the file storage when a candidate file is hit and then continuing processing, previously completed partial processing files can be utilized, saving core computing resources. By retrieving candidate installation packages from the installation package storage as the target installation package when a candidate installation package is hit, processing efficiency is improved. Thus, by distinguishing reusable historical candidate objects into candidate installation packages and candidate files, and setting corresponding retrieval and reprocessing paths for different hit types, fine-grained adaptive reuse is achieved.

[0134] According to embodiments of this disclosure, the installation package to be processed is processed according to the processing method corresponding to the matching result to obtain the target object. The process may further include the following operations: if no candidate file or candidate installation package corresponding to the reuse identifier exists in the mapping, each file is processed according to the configuration information to obtain the target installation package; the content features and reuse identifier of each processed file are associated and stored in file storage; the target installation package and reuse identifier are associated and stored in installation package storage; the mapping is updated using the reuse identifier, the storage address of the processed file, and the storage address of the target installation package.

[0135] If a query of the mapping fails to find any historical processing record corresponding to the reuse identifier of this processing request, indicating that no historical artifacts can be reused, then a full execution is required. In this case, the orchestrator (such as Orchestrator) can sequentially schedule processing tools to process the files that need to be processed, then send the processed files to the rebuilding engine for packaging, and finally sign them to generate a brand new target installation package. The processed file is a newly generated file containing the processing results after the operations indicated by the configuration information have been applied to the original file in this new execution process.

[0136] It should be noted that if the installation package contains a large number of files, the orchestrator does not assign all files to a single processing node for serial execution. Instead, it distributes the processing requests to a processing cluster, where multiple worker nodes in the cluster process the files assigned to them in parallel and independently. Finally, all processed files are aggregated and reassembled to shorten the overall processing time of large applications.

[0137] After obtaining the processed file and the target installation package, on the one hand, the processed file can be stored in the file storage using the content hash of the corresponding processed file as the index key; on the other hand, the target installation package can be stored in the installation package storage using the reuse identifier of this processing request as the index key.

[0138] Furthermore, when processed files and target installation packages are stored, not only can reuse identifiers be associated, but lifecycle tags can also be associated based on configuration information. For example, a lifecycle tag can be 30d, indicating that the candidate object is valid for 30 days, after which it will be automatically downgraded or cleaned up.

[0139] After storing the processed file in the file storage and the target installation package in the installation package storage, the mapping can be updated using the storage address of the processed file in the file storage, the storage address of the target installation package in the installation package storage, and the reuse identifier, so that the new processing results can be queried and reused by all subsequent identical or similar requests.

[0140] It's important to note that before updating the mapping, a reuse flag can be used to query the mapping again to check if another concurrent similar request has preemptively completed the same work and updated the mapping during the current request execution. If a conflict is found, the current request can abort its write; otherwise, the write is executed to ensure repeated writes and data consistency under high concurrency.

[0141] In the embodiments of this disclosure, full processing is performed after determining that the input does not exist in the mapping, ensuring that full computational resources are allocated when faced with entirely new combinations of inputs. Secondly, by storing the processed files with a reuse identifier associated with the target installation package, these artifacts are given a precise, traceable identity label. By updating the mapping using the reuse identifier and storage address, any subsequent user or pipeline initiating the same task can reuse the results of this processing request, thereby improving the reuse hit rate and reducing average processing time and resource consumption.

[0142] According to embodiments of this disclosure, processing each file to be processed according to configuration information to obtain a target installation package may include the following operations: sorting each processed file according to the order of each file in the installation package to obtain a processed file sequence; filtering out construction noise in the processed file sequence according to preset fields to obtain the target installation package.

[0143] Before performing the actual binary packaging, the physical order of the files in the original installation package, which was recorded during the unpacking and parsing phase, can be read. Then, all the files to be used in this packaging are sorted according to this order, thus forming a deterministic sequence of processed files.

[0144] The order of the files to be processed within the installation package refers to the specific arrangement of the internal files that make up the installation package within the original physical storage structure of the installation package. This order can be determined by the installation package developer or the build tool when generating the package. For example, taking the file entries in the central directory of an APK in ZIP format as an example, the order could be: first .xml files, then .dex files, followed by all resource files in the res directory arranged lexicographically, and finally library files in the lib directory.

[0145] The processed file sequence is an ordered list of files formed by arranging all the processed files and the original files that were not processed but are part of the installation package in sequence.

[0146] Even when using the same set of files, generating an installation package twice with the same packaging tool may still result in build noise in the final package, even if the file content remains unchanged. Build noise consists of variable, non-functional data automatically introduced into the final installation package file by the packaging tool or library due to its own implementation principles. Build noise does not affect application execution, but it can cause inconsistencies in the byte order of the two packaged packages.

[0147] To address the aforementioned build noise, content filtering of the processed file sequence can be performed using preset fields to accurately locate newly generated, uncertain metadata fields during the packaging process, and replace them all with fixed, meaningless, or reproducible values ​​to eliminate build noise.

[0148] Predefined fields are fields that need to be processed in advance during the repackaging stage. Predefined fields can be used to indicate which metadata fields, which may be automatically generated by the packaging tool or are variable, need to be filtered or set to fixed values ​​during the generation of the target installation package. For example, when generating a ZIP format APK, the field recording the last modified time of the file in the header of each file entry can be used as a preset field. Alternatively, the ZIP file compression method identifier, the location of the compressed checksum, etc., can also be used as preset fields.

[0149] In the embodiments of this disclosure, by sorting the processed files according to the original order of the files within the installation package, the various processed files can be restored to their original physical arrangement before packaging, thereby eliminating file order changes that may be introduced by the processing logic itself or the file aggregation process. Furthermore, by filtering the processed file sequence to eliminate build noise introduced during the reassembly process, fields such as timestamps and variable compression parameters that fluctuate with the build environment can be actively eliminated. Thus, even if the exact same input sequence is used to trigger reassembly and packaging on different servers at different times, the resulting target installation package, before and after signature removal, will have byte-level or functional-level consistency in its code and resources. This forms a closed loop with the input denoising, avoiding duplicate processing caused by incorrect identification of different packages due to output drift during the reassembly stage, and improving the processing efficiency of the installation package.

[0150] Figure 6 The illustration shows an example of a process in which an installation package is processed according to a matching result to obtain a target installation package, based on an embodiment of the present disclosure.

[0151] like Figure 6As shown, in embodiment 600 of obtaining the target installation package, mapping 620 stores the correspondence between candidate identifiers and the storage addresses of candidate objects. Multiple candidate identifiers in mapping 620 can be matched according to reuse identifier 610 to obtain matching result 630. After obtaining matching result 630, operation S610 can be executed. In operation S610, does a candidate identifier match reuse identifier 610?

[0152] If not, operations S611 and S612 can be executed. In operation S611, each file is processed according to the configuration information to obtain processed files; in operation S612, the processed files are sorted according to their order in the installation package to obtain a processed file sequence. The processed file sequence is filtered according to preset fields to eliminate build noise introduced by the reassembly process, resulting in the target installation package 640.

[0153] If so, operation S620 can be executed. In operation S620, is there a candidate file corresponding to multiplexing identifier 610?

[0154] If so, operations S621 and S622 can be executed. In operation S621, candidate files are retrieved from the file storage; in operation S622, the candidate files are used to replace the corresponding files in the installation package to obtain the target installation package 640.

[0155] If not, operation S623 can be executed. In operation S623, the candidate installation package obtained from the installation package storage is used as the target installation package 640.

[0156] According to embodiments of this disclosure, the above-described installation package processing method may further include the following operations: classifying candidate objects into levels based on reuse statistics for each candidate object in the mapping to obtain the level of each candidate object; and managing the candidate objects in the mapping, file storage, and installation package storage according to the level.

[0157] For each candidate object in the mapping, reuse statistics for each candidate object within the most recent time window can be periodically collected. These statistics are then compared with a predefined level classification strategy to assign a specific level to each candidate object. The reuse statistics are continuously collected and recorded during runtime, serving as a data set to quantify the degree to which each candidate object is queried and used by subsequent tasks.

[0158] In one embodiment, reuse statistics may include at least one of reuse frequency, most recent reuse time, and reuse count. For example, reuse frequency may be that a candidate file has been hit 50 times in the past 30 days, most recent reuse time may be that a candidate installation package was last hit on April 15, 2026, and reuse count may be that it has been reused 120 times since it was stored.

[0159] A level is a rating label assigned to each candidate object to guide its subsequent storage strategy and lifecycle operations. The level classification method can be configured according to actual business needs and is not limited here. For example, levels can be classified based on preset thresholds. Specifically, if a candidate object has been reused arbitrarily in the past 7 days, the level is "hot"; if it has been reused in the past 30 days but not in the past 7 days, the level is "warm"; if there are no reuse records in the past 90 days, the level can be "cold".

[0160] Alternatively, a level classification can be based on a machine learning prediction model. This involves not only looking at historical statistics but also using a time series prediction model. The input to this model is the historical reuse statistics of the candidate object and the type and version information of the installation package to which it belongs. The output is a prediction of the reuse statistics for a future period of time. Thus, levels can be classified based on the predicted future popularity.

[0161] After obtaining the levels of each candidate object, management measures that balance access efficiency and storage costs can be automatically executed based on the assigned level. For example, management may include migrating data between storage media with different performance / cost profiles, downgrading data retention methods, and thoroughly cleaning up data.

[0162] In the embodiments of this disclosure, by classifying candidates according to reuse statistics, the storage value of candidate objects can be quantified into comparable level labels, thereby establishing an objective and dynamic basis for subsequent differentiated management. Based on this, by managing candidate objects according to levels, it is possible to suppress the linear growth of storage scale over time while ensuring that candidate objects with high reuse statistics are retained to maintain the ability to rebuild installation packages, thus achieving proactive control of storage costs.

[0163] According to embodiments of this disclosure, managing candidate objects in mapping, file storage, and installation package storage according to levels may include the following operations: in the case of level 1, retaining candidate files in file storage and candidate installation packages in installation package storage; in the case of level 2, retaining candidate files in file storage; and in the case of level 3, retaining candidate identifiers in mapping and structural units used to reconstruct the file content structure.

[0164] Level 1 is the highest level of a candidate object in a hierarchy based on reuse statistics. Level 1 indicates that the candidate object has been frequently queried and reused recently. When a candidate object is determined to be at Level 1, all its data in file storage and installation package storage can be maintained completely and without deletion or modification to ensure that it is in a state that can be accessed at any time.

[0165] The second level is an intermediate level between the highest and lowest. The second level indicates that the candidate still has reuse value, but its popularity has decreased. When the popularity of a candidate drops to the second level, candidate installation packages that occupy a lot of space but are not essential can be actively deleted from the installation package storage, but the core raw materials that make up the package, namely the candidate files in the file storage, will be completely retained, thus ensuring that the finished product can still be quickly assembled from the effective parts.

[0166] Level 3 is the lowest level a candidate object reaches in the hierarchy. Level 3 indicates that the direct reuse value of the candidate object is low or that it has passed its lifecycle. When a candidate object is determined to be at Level 3, all physical artifacts in the physical storage can be completely deleted, leaving only the candidate identifier representing the task and the structural unit used to reconstruct the file content structure.

[0167] The structural units used to reconstruct the file content structure are minimal data fragments that remain in the mapping or metadata database after almost all physical storage data has been cleaned up. These structural units are not directly reusable, but they contain logical information such as what the task processed and how it was restored. For example, a structural unit may include the name, path, and content hash tree of the original file to be processed, the identifier and version number of the applied strategy, version information, etc.

[0168] In the embodiments of this disclosure, for first-level candidate objects, a full-scale strategy of retaining candidate files and candidate installation packages is implemented to ensure the response speed of frequently reused tasks and reduce resource overhead and waiting time. For second-level candidate objects, a strategy of retaining candidate files while discarding candidate installation packages is implemented. This strategy can delete large complete signature packages and release the main storage space while retaining candidate files obtained from intermediate processing. This allows subsequent reuse requests to skip the processing step, even though they need to re-signature, thus reducing storage costs while maintaining processing efficiency superior to full recalculation. For third-level candidate objects, a strategy of retaining candidate identifiers and structural units used to reconstruct the file content structure can compress storage usage to a negligible metadata level, reclaiming storage space for useless data, while the retained structural units still provide the ability to trace and theoretically reconstruct historical products. Thus, by performing differentiated management operations on candidate objects of different levels, a balance between storage costs and reuse efficiency is achieved.

[0169] The above are merely exemplary embodiments, but are not limited thereto. Other methods for processing installation packages known in the art may also be included, as long as they can reduce resource waste caused by repetitive calculations and improve the processing efficiency of installation packages.

[0170] Figure 7 A block diagram of an installation package processing apparatus according to an embodiment of the present disclosure is shown schematically.

[0171] like Figure 7 As shown, the installation package processing device 700 may include a structure parsing module 710, a first generation module 720, and a processing module 730.

[0172] The structure parsing module 710 is used to parse the structure of the installation package in response to a processing request for the installation package, and obtain the file content structure, wherein the file content structure represents the relationship between multiple files in the installation package and the content characteristics of each file.

[0173] The first generation module 720 is used to generate a reuse identifier based on the configuration information and file content structure indicated by the processing request. The reuse identifier is used to match with the mapping to obtain a matching result, and the matching result indicates whether there is a reusable installation package or reusable file corresponding to the processing request.

[0174] The processing module 730 is used to process the installation package according to the processing method corresponding to the matching result to obtain the target installation package.

[0175] According to embodiments of this disclosure, the structure parsing module 710 may include a determining submodule and a generating submodule.

[0176] The determination submodule is used to identify candidate files among multiple candidate files obtained after standardizing the installation package, those whose impact on the functionality of the installation package is greater than a predetermined threshold, as files.

[0177] The generation submodule is used to generate the file content structure based on the content characteristics of each file.

[0178] According to embodiments of this disclosure, content features are obtained by: filtering out construction noise in a file based on preset fields to obtain an intermediate file, wherein construction noise is content generated during the construction of the installation package that causes differences between bytes but does not affect functionality; and generating content features based on the content of the intermediate file.

[0179] According to embodiments of this disclosure, the generated submodule may include a building unit and an associated unit.

[0180] The building unit is used to construct a hierarchical structure based on the relationships between files. The hierarchical structure includes multiple nodes, and each node represents a file.

[0181] The association unit is used to associate the content features of each file with the corresponding nodes in the hierarchical structure to obtain the file content structure.

[0182] According to embodiments of this disclosure, the configuration information indicates at least one configuration item; the first generation module 720 may include an acquisition submodule and a fusion submodule.

[0183] The acquisition submodule is used to obtain target content features and attribute information from the file content structure based on configuration items. The attribute information represents the attributes of the application corresponding to the installation package.

[0184] The fusion submodule is used to merge the target content features, attribute information and configuration items to obtain a reuse identifier.

[0185] According to embodiments of this disclosure, the fusion submodule may include an acquisition unit and a fusion unit.

[0186] The acquisition unit is used to obtain version information of the tool used to execute the processing request.

[0187] The fusion unit is used to merge version information, target content features, attribute information and configuration items to obtain a reuse identifier.

[0188] According to embodiments of this disclosure, the configuration items include at least one of the following: whether to harden the installation package, whether to encrypt the installation package, whether to rewrite the data in the installation package, whether to reorganize the file, and whether to change the file format.

[0189] According to embodiments of this disclosure, the mapping indicates the correspondence between candidate identifiers and the storage addresses of candidate objects; the processing module 730 may include a matching submodule and a processing submodule.

[0190] The matching submodule is used to match multiple candidate identifiers in the mapping based on the reuse identifier to obtain the matching result.

[0191] The processing submodule is used to process the installation package according to the processing method corresponding to the matching result to obtain the target object.

[0192] According to embodiments of this disclosure, the candidate objects include at least one of candidate installation packages and candidate files; the processing submodule may include a first processing unit and a second processing unit.

[0193] The first processing unit is used to replace the corresponding file in the installation package with the candidate file obtained from the file storage according to the storage address when there is a candidate file corresponding to the reuse identifier in the mapping, so as to obtain the target installation package.

[0194] The second processing unit is used to, when there is a candidate installation package corresponding to the reuse identifier in the mapping, use the candidate installation package obtained from the installation package storage according to the storage address as the target installation package.

[0195] According to embodiments of this disclosure, the processing submodule may further include a third processing unit, a storage unit, and an update unit.

[0196] The third processing unit is used to process each file according to the configuration information to obtain the target installation package when there is no candidate file or candidate installation package corresponding to the reuse identifier in the mapping.

[0197] The storage unit is used to associate and store the content characteristics and reuse identifiers of the processed files to the file storage, and to associate and store the target installation package and reuse identifiers to the installation package storage.

[0198] The update unit is used to update the mapping using the reuse identifier, the storage address of the processed file, and the storage address of the target installation package.

[0199] According to embodiments of this disclosure, the third processing unit may include a sorting subunit and a filtering subunit.

[0200] The sorting subunit is used to sort the processed files according to their order in the installation package to obtain a sequence of processed files.

[0201] The filtering subunit is used to filter out build noise in the processed file sequence according to preset fields to obtain the target installation package.

[0202] According to embodiments of this disclosure, the package processing apparatus 700 may further include a level division module and a management module.

[0203] The level division module is used to classify the candidate objects into levels based on the reuse statistics of each candidate object in the mapping, and obtain the level of each candidate object. The reuse statistics include at least one of reuse frequency, most recent reuse time, and reuse frequency.

[0204] The management module is used to manage candidate objects in mapping, file storage, and installation package storage according to their levels.

[0205] According to embodiments of this disclosure, the management module may include a first management submodule, a second management submodule, and a third management submodule.

[0206] The first management submodule is used to retain candidate files in the file storage and candidate installation packages in the installation package storage when the level is first level.

[0207] The second management submodule is used to retain candidate files in the file storage when the level is second level.

[0208] The third management submodule is used to retain candidate identifiers in the mapping and structural units used to reconstruct the file content structure when the level is third. The first level is higher than the second level, and the second level is higher than the third level.

[0209] According to embodiments of this disclosure, the processing request is submitted through a front-end interface, which is used to receive configuration information and source information of the installation package; the installation package processing device 700 may further include a second generation module and a display module.

[0210] The second generation module is used to generate a request identifier for processing the request in response to the verification of the source information and configuration information.

[0211] The display module is used to display the request identifier and processing status through the front-end interface. The processing status represents the execution status corresponding to the current execution stage of the processing request.

[0212] Figure 8 A block diagram schematically illustrates an electronic device suitable for implementing a processing method for an installation package according to embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0213] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0214] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0215] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the package processing method. For example, in some embodiments, the package processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the package processing method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the package processing method by any other suitable means (e.g., by means of firmware).

[0216] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0217] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0218] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0219] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0220] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0221] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.

[0222] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0223] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for processing an installation package, comprising: In response to receiving a processing request for the installation package, the installation package is structurally parsed to obtain a file content structure, wherein the file content structure represents the relationship between multiple files in the installation package and the content characteristics of each file; Based on the configuration information indicated by the processing request and the file content structure, a reuse identifier is generated, wherein the reuse identifier is used to match with the mapping to obtain a matching result, and the matching result indicates whether there is a reusable installation package or reusable file corresponding to the processing request; and The installation package is processed according to the processing method corresponding to the matching result to obtain the target installation package.

2. The method according to claim 1, wherein, The process of parsing the installation package to obtain the file content structure includes: From the multiple candidate files obtained after standardizing the installation package, the candidate file whose impact on the functionality of the installation package exceeds a predetermined threshold is determined as the file; and The file content structure is generated based on the content characteristics of each file.

3. The method according to claim 2, wherein, The content features are obtained through the following methods: Based on preset fields, build noise in the files is filtered out to obtain intermediate files. The build noise refers to content generated during the construction of the installation package that causes differences between bytes but does not affect functionality. The content features are generated based on the content of the intermediate file.

4. The method according to claim 2, wherein, The step of generating the file content structure based on the content characteristics of each file includes: Based on the relationships between the files, a hierarchical structure is constructed, wherein the hierarchical structure includes multiple nodes, and each node represents a file; and The content features of each file are associated with the corresponding nodes in the hierarchical structure to obtain the file content structure.

5. The method according to any one of claims 1 to 4, wherein, The configuration information indicates at least one configuration item; The step of generating a reuse identifier based on the configuration information indicated by the processing request and the file content structure includes: Based on the configuration items, target content features and attribute information are obtained from the file content structure, wherein the attribute information characterizes the attributes of the application corresponding to the installation package; and The target content features, the attribute information, and the configuration items are fused to obtain the reuse identifier.

6. The method according to claim 5, wherein, The step of fusing the target content features, the attribute information, and the configuration items to obtain the reuse identifier includes: Obtain version information of the tool used to execute the processing request; and The reuse identifier is obtained by fusing the version information, the target content features, the attribute information, and the configuration items.

7. The method according to claim 5 or 6, wherein, The configuration item includes at least one of the following: Whether to harden the installation package, whether to encrypt the installation package, whether to rewrite the data in the installation package, whether to reorganize the file, and whether to change the format of the file.

8. The method according to any one of claims 1 to 7, wherein, The mapping indicates the correspondence between candidate identifiers and the storage addresses of candidate objects; The step of processing the installation package according to the processing method corresponding to the matching result to obtain the target installation package includes: Based on the reuse identifier, multiple candidate identifiers in the mapping are matched to obtain the matching result; as well as The installation package is processed according to the processing method corresponding to the matching result to obtain the target object.

9. The method according to claim 8, wherein, The candidate objects include at least one of the candidate installation packages and candidate files; The step of processing the installation package according to the processing method corresponding to the matching result to obtain the target object includes: If a candidate file corresponding to the reuse identifier exists in the mapping, the corresponding file in the installation package is replaced with the candidate file obtained from the file storage according to the storage address to obtain the target installation package; as well as If a candidate installation package corresponding to the reuse identifier exists in the mapping, the candidate installation package obtained from the installation package storage according to the storage address shall be used as the target installation package.

10. The method of claim 9, further comprising: If no candidate file or candidate installation package corresponding to the reuse identifier exists in the mapping, each file is processed according to the configuration information to obtain the target installation package; The content features of each processed file and the reuse identifier are associated and stored in the file storage; the target installation package and the reuse identifier are associated and stored in the installation package storage. as well as The mapping is updated using the reuse identifier, the storage address of the processed file, and the storage address of the target installation package.

11. The method according to claim 10, wherein, The step of processing each of the files according to the configuration information to obtain the target installation package includes: The processed files are sorted according to their order in the installation package to obtain a sequence of processed files; Based on preset fields, the construction noise in the processed file sequence is filtered out to obtain the target installation package.

12. The method according to any one of claims 1 to 11, further comprising: Based on the reuse statistics of each candidate object in the mapping, the candidate objects are classified into levels to obtain the level of each candidate object, wherein the reuse statistics include at least one of reuse frequency, most recent reuse time, and reuse frequency; and Based on the stated level, candidate objects in the mapping, file storage, and installation package storage are managed.

13. The method according to claim 12, wherein, The management of candidate objects in the mapping, file storage, and installation package storage according to the level includes: When the level is the first level, the candidate files in the file storage and the candidate installation packages in the installation package storage are retained; In the case of the second level, candidate files in the file storage are retained; and In the case of the third level, the candidate identifiers in the mapping and the structural units used to reconstruct the file content structure are retained; The first level is higher than the second level, and the second level is higher than the third level.

14. The method according to claim 1, wherein, The processing request is submitted through a front-end interface, which is used to receive the configuration information and the source information of the installation package; The method further includes: In response to the verification of the source information and the configuration information, a request identifier for the processing request is generated; and The front-end interface displays the request identifier and processing status, wherein the processing status represents the execution status corresponding to the current execution stage of the processing request.

15. An apparatus for processing an installation package, comprising: The structure parsing module is used to respond to a processing request for the installation package, perform structure parsing on the installation package, and obtain the file content structure, wherein the file content structure represents the relationship between multiple files in the installation package and the content characteristics of each file; A first generation module is configured to generate a reuse identifier based on the configuration information indicated by the processing request and the file content structure. The reuse identifier is used to match a mapping to obtain a matching result, which indicates whether a reusable installation package or reusable file corresponds to the processing request. The processing module is used to process the installation package according to the processing method corresponding to the matching result to obtain the target installation package.

16. An electronic device comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 14.

17. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 14.

18. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 14.