Code compression method and device
By building module dependency graphs in code compression, analyzing and optimizing the syntax analysis tree, eliminating redundancy and dead loops, and using polymorphic mutation technology for obfuscation and compression, the problem of insufficient code compression performance trade-offs and effects in the existing technology is solved, and the ultimate compression and security guarantee of the code is achieved.
Patent Information
- Application Number
- CN202510063090.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-09
AI Technical Summary
Existing code compression technologies have shortcomings in performance trade-offs and compression effects, especially in modern front-end engineering methods, where excessive compression rates may affect the development experience, compression effects are limited to different types of files, and there are error handling and security risks.
By detecting the startup of the code packaging construction process, the dependencies of the project module are identified layer by layer from the entry function, and the module dependency graph is built. Use loaders and syntax analyzers to analyze the source code, construct the syntax analysis tree, eliminate redundant logic through tree-shaking, detect and eliminate duplicate or similar function definitions, eliminate logical dead loops, and use polymorphic mutation technology to perform code obfuscation and compression.
It realizes the ultimate compression of the code, improves loading speed, and while ensuring code security, it solves redundancy problems, improving code optimization rate and security.
Smart Images

Figure CN119960763A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of code processing technology, and in particular to a code compression method and device. Background Art
[0002] In the field of front-end engineering, common code compression algorithms and tools include: Gzip, Brotli, fixed bit length algorithm, Run-Length Encoding (RLE) and LZ77 and its derivative algorithms (such as LZSS).
[0003] Specifically, Gzip is the abbreviation of GNUzip, which is a widely used file compression format (it can also be said to be several file compression programs). It combines the LZ77 algorithm with Huffman coding and is a lossless compression algorithm; Brotli is a new generation compression algorithm developed by Google, which aims to provide a higher compression rate than Gzip. However, compared with Gzip, Brotli takes a long time to compress, but the decompression speed is faster; the fixed bit length algorithm uses the least bits to represent text data by optimizing text encoding. It is suitable for the compression of binary data and is mainly suitable for binary data scenarios; RLE is a simple data compression algorithm that achieves compression by recording consecutive characters and their number of occurrences. It works better for text or images with a large amount of continuous repeated data; LZ77 and its derivative algorithms (such as LZSS) are dictionary-based compression algorithms that achieve compression by finding repeated substrings in the text. LZSS is a derivative algorithm of LZ77 and uses a smaller sliding window. Among them, the common compression algorithms in front-end engineering mainly include Gzip, Brotli, etc., which are mainly used for text resource compression.
[0004] However, although the existing code compression technologies have their strengths in the field of general text compression, they still have shortcomings when combined with modern front-end engineering methods, mainly manifested in: Performance trade-off: Although the compression algorithm can significantly reduce the file size, the compression and decompression process itself takes time. Especially for algorithms with high compression rates such as Brotli, the compression time is relatively long, which may increase the build time or server processing time. For example, in real-time development and preview scenarios, if the compression rate is too high, it will directly affect the development experience; Limited compression effect: Different types of files (such as text, pictures, videos, etc.) respond to compression algorithms to different degrees. For some already highly compressed files (such as JPEG pictures), further compression may not achieve significant results, and may even cause file quality to deteriorate; Error handling: If the compressed file is damaged or decompressed during transmission, it may cause resource loading failure or page display abnormality. Developers need to properly handle these abnormal situations to ensure that users can access the page smoothly; Security risks: Although the compression algorithm itself does not directly involve data security, in some cases, if the compressed file is accessed or leaked without authorization, sensitive information may be exposed. In addition, for data involving user privacy, special attention should be paid to encryption and transmission security.
[0005] In view of this, how to provide a code compression method to achieve extreme compression of the code, improve loading speed, and ensure the security of the code has become a technical problem that urgently needs to be solved. Summary of the invention
[0006] The embodiments of the present application provide a code compression method, a code compression device, an electronic device and a computer storage medium, which are used to solve the problem of how to eliminate code redundancy and improve loading speed while ensuring code security.
[0007] In a first aspect of an embodiment of the present application, a code compression method is provided, comprising:
[0008] When the code packaging and building process is detected to be started, starting from the entry function, the dependencies of each module in the project are identified layer by layer, and a module dependency graph is constructed, wherein the module dependency graph carries the target address table of the module and the dependency relationship between each sub-module;
[0009] Through the loader and syntax analyzer, the source code file type of the module is analyzed and a syntax analysis tree is constructed;
[0010] Eliminate redundant logic branches of the grammar analysis tree through tree-shaking, eliminate duplicate or similar function definitions through similar grammar analysis tree detection, obtain a target grammar analysis tree, and analyze the target grammar analysis tree through deep traversal routes to determine and eliminate the existence of logical dead loop code flows;
[0011] The optimized source code is automatically mutated using polymorphic mutation technology to generate a target code text, which is written into a file in a target compression format and output, wherein the target code has the same function as the source code.
[0012] In a second aspect of an embodiment of the present application, a code compression device is provided, including:
[0013] The loading module is configured to, when detecting that the code packaging and building process is started, identify the dependencies of each module in the project layer by layer starting from the entry function, and build a module dependency graph, wherein the module dependency graph carries the module's target address table and the dependency relationship between each sub-module;
[0014] The analysis module is configured to analyze the source code file type of the module through the loader and the syntax analyzer, and construct a syntax analysis tree;
[0015] A construction module is configured to eliminate redundant logic branches of the grammar analysis tree by tree-shaking, eliminate duplicate or similar function definitions by similar grammar analysis tree detection, obtain a target grammar analysis tree, and analyze the target grammar analysis tree by deep traversal route, determine the existence of logic dead loop code flow, and eliminate it;
[0016] The compression module is configured to automatically mutate the optimized source code using polymorphic mutation technology to generate a target code text, write the target code text into a file in a target compression format, and output it, wherein the target code has the same function as the source code.
[0017] In a third aspect of an embodiment of the present application, a computing device is provided, including:
[0018] Memory and processor;
[0019] The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions. When the computer executable instructions are executed by the processor, the steps of the above-mentioned code compression method are implemented.
[0020] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned code compression method are implemented.
[0021] The present application provides a code compression method, comprising: first, when it is detected that the code packaging construction process is started, starting from the entry function, the dependencies of each module in the project are identified layer by layer, and a module dependency graph is constructed, wherein the module dependency graph carries the module's target address table and the dependencies between each sub-module; then, through a loader and a syntax analyzer, the source code file type of the module is analyzed to construct a syntax analysis tree; secondly, redundant logical branches of the syntax analysis tree are eliminated through tree-shaking, and repeated or similar function definitions are eliminated through similar syntax analysis tree detection to obtain a target syntax analysis tree, and the target syntax analysis tree is analyzed through a deep traversal route to judge the existence of a logical dead loop code flow and eliminate it; finally, the optimized source code is automatically mutated using a polymorphic mutation technology to generate a target code text, and the target code text is written into a file through a target compression format and output, wherein the target code has the same function as the source code.
[0022] The code compression method provided in the embodiment of the present application is applied in combination with modern front-end development engineering methods. Based on the code AST syntax tree, code logic redundancy detection such as repeated function implementation and dead loop of program execution process is completed, the code volume is deeply trimmed, and the obfuscation and compression algorithms are used to finally complete the extreme compression of the code and improve the loading speed.
[0023] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, which can be implemented in accordance with the contents of the specification, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0025] Figure 1 A flowchart of a code compression method provided in an embodiment of the present application;
[0026] Figure 2 A schematic diagram of a loop-dependent scenario in a code compression method provided in an embodiment of the present application;
[0027] Figure 3 A schematic diagram of another flow chart of a code compression method provided in an embodiment of the present application;
[0028] Figure 4A schematic diagram of a module index representation in a code compression method provided in an embodiment of the present application;
[0029] Figure 5 A schematic diagram of the structure of a code compression device provided in an embodiment of the present application;
[0030] Figure 6 A structural block diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0031] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0032] See also Figure 1 , Figure 1 A schematic diagram of a code compression method provided in an embodiment of the present application. Figure 1 As shown, the specific steps include the following steps.
[0033] Step S102: When the code packaging and building process is detected to be started, starting from the entry function, the dependencies of each module in the project are identified layer by layer, and a module dependency graph is constructed, wherein the module dependency graph carries a target address table of the module and dependencies between each sub-module;
[0034] Step S104: Analyze the source code file type of the module through the loader and the syntax analyzer, and construct a syntax analysis tree;
[0035] Step S106: Eliminate redundant logic branches of the grammar analysis tree by tree-shaking, eliminate duplicate or similar function definitions by similar grammar analysis tree detection, obtain a target grammar analysis tree, and analyze the target grammar analysis tree by deep traversal route, determine the existence of logic dead loop code flow, and eliminate it;
[0036] Step S108: using polymorphic mutation technology to automatically mutate the optimized source code to generate a target code text, writing the target code text into a file in a target compression format, and outputting it, wherein the target code has the same function as the source code.
[0037] In the embodiment of the present application, the source code file type of the module is analyzed by the loader and the syntax analyzer to construct a syntax analysis tree, including:
[0038] According to the source code file type of the module, during the loading process using the loader, the analyzer will scan each character in the source code to generate a character sequence, wherein the character sequence carries the words corresponding to each recognized character;
[0039] Based on the generated words, a grammar dictionary corresponding to the module is generated.
[0040] The grammar dictionary is analyzed using a syntax analyzer to generate various phrases, and a syntax analysis tree is constructed based on the phrases.
[0041] In the embodiment of the present application, the method of eliminating duplicate or similar function definitions by similar syntax parse tree detection to obtain a target syntax parse tree includes:
[0042] In the case where it is detected that there are duplicate imported functions between modules, determining that addresses corresponding to the duplicate imported functions are the same address, and storing the addresses in the module dependency graph;
[0043] When it is detected that the execution processes of a first process function and a second process function are similar between modules, a first syntax parse tree is constructed for the first process function, a second syntax parse tree is constructed for the second process function, and the first syntax parse tree and the second syntax parse tree are compared through a similar tree algorithm to identify a target function, wherein the first process function and the second process function are execution functions in the execution process, and the structures of the first syntax parse tree and the second syntax parse tree are similar.
[0044] The identification of repeatedly implemented function methods mentioned in the embodiments of the present application includes two situations: the first situation is the repeatedly imported functions between modules: the functions in this scenario will use the same target address in the module function dependency table, thereby simplifying the dependency graph and reducing the repeated import of imported module functions; the second situation is the functions with similar execution processes in the module: the functions in this scenario will form an AST syntax tree with a similar structure during the syntax analysis process. At this time, the functions with similar execution processes can be identified by comparing the similar tree algorithm, and then the code optimization processing can be performed according to the context.
[0045] In an embodiment of the present application, the target syntax analysis tree is analyzed through a deep traversal route to determine whether the analysis process of the target syntax analysis tree is a loop, and if the analysis process is a loop, the loop logic is removed; if the analysis process is not looped, the execution nodes are reduced.
[0046] In an embodiment of the present application, the method of eliminating redundant logical branches of the grammar analysis tree by tree-shaking includes:
[0047] Collecting export variables, and recording the export variables into module variables corresponding to the syntax analysis tree;
[0048] The syntax analysis tree is traversed to mark whether the module export variable is used, and if the module export variable is not used, the unused export variable is eliminated.
[0049] The embodiment of the present application mainly uses the Tree-Shaking optimization technology to identify and process useless logical branches.
[0050] Specifically, the Tree-Shaking optimization technology used in the embodiment of the present application mainly includes two steps: 1) marking unused module export values; 2) using Terser to delete unused export statements.
[0051] The marking process can be divided into three steps: 1) Make stage, collecting export variables and recording them in the module dependency graph ModuleGraph variable; 2) Seal stage, traversing ModuleGraph to mark whether the module export variables are used; 3) When generating products, if the variable is not used by other modules, delete the corresponding export statement.
[0052] In the embodiment of the present application, the target syntax analysis tree is analyzed by deep traversal route, and the existence of logic dead loop code flow is judged and eliminated, including:
[0053] Analyze the target syntax analysis tree by deep traversal route, randomly select unsearched nodes to start depth-first search, and mark the unsearched nodes as being in the searching state;
[0054] Traversing the adjacent nodes of the unsearched node, and when there is a first adjacent node in the adjacent nodes that is in an unsearched state, starting a depth-first search from the first adjacent node;
[0055] When there is a second adjacent node in the adjacent nodes that is in the searching state, determining that the analysis process corresponding to the second adjacent node forms a loop;
[0056] When each of the adjacent nodes is completed, it is determined that the analysis process of the target analysis tree has no loop.
[0057] like Figure 2 As shown, Figure 2 A schematic diagram of a dependent loop scenario in a code compression method provided in an embodiment of the present application, wherein module1, module2, and module3 are circularly dependent, resulting in the risk of the code output falling into an infinite loop.
[0058] This application mainly traverses the dependency tree through the depth-first search method to avoid the problem of circular dependencies.
[0059] The specific process is as follows: First, take any "unsearched" node x to start a depth-first search, and mark the node as "searching"; then, traverse each adjacent node y of the node x. If y is an "unsearched" node, then start the depth-first search from the y node; if y is a "searching" node, it means there is a cycle in the graph; if y is a "completed" node, it means that the branches and sub-branches of the node have been confirmed to be cycle-free and no operation is required; finally, when all nodes of x are "completed", it means that the graph is cycle-free.
[0060] This method can be used to effectively identify loop scenarios where modules depend on each other, and perform "loop breaking" processing before outputting code to avoid dead loop problems in code execution.
[0061] In summary, the implementation of the code compression method provided by the embodiment of the present application is as follows: first, a module dependency graph is constructed, that is, starting from the entry function, the dependencies of each module in the project are identified layer by layer, and thus a module dependency graph is constructed. The module dependency graph not only records the target address table of the dependent module, but also the dependency relationship between each sub-module; then module loading and lexical analysis are performed, that is, after identifying the module dependency, different Loaders are used to load the module according to the source code file type. During the Loader loading process, the analyzer will scan each character of the source code, identify each word, and generate a grammar dictionary corresponding to the module; secondly, grammatical analysis and AST syntax tree construction are performed, that is, according to the token sequence output by the lexical analyzer in the previous step, a grammatical analyzer is used for analysis. It can generate various phrases and construct a syntax analysis tree to prepare for code optimization and output; secondly, code optimization is performed, that is, based on the AST syntax tree, logical redundancy is eliminated through tree-shaking, and repeated or similar function definitions are eliminated through similar AST syntax tree detection. Whether the tree analysis process is a cycle is analyzed by deep traversal. If it is a cycle, the loop logic is removed; if there is no cycle, the execution nodes are minimized and the depth of the directed graph is reduced, so as to achieve the purpose of code optimization; finally, code obfuscation and compression are performed, that is, in the obfuscation stage, polymorphic mutation technology is used so that the code will automatically mutate each time it is called, and change into a completely different code from before, but the function remains unchanged. In the compression stage, the text is written to the file in gzip compression format, and the code with compressed file size is finally output.
[0062] The embodiment of the present application points out that by detecting the AST syntax tree structure obtained by lexical analysis and syntax analysis, it is determined whether there are repeated or similar function definitions in the source code. Redundant and similar function definitions are found at the grammatical level, and they are merged in combination with the execution context to deeply trim the code volume, which can improve the code optimization rate.
[0063] The embodiment of the present application points out that by checking whether the module dependencies form a cycle, the circular dependency problem between modules can be detected in advance, and the "loop breaking" processing is performed through independent module splitting and asynchronous reference, so as to avoid errors caused by circular dependencies during program operation.
[0064] The embodiments of the present application point out that in the code obfuscation and compression stage, polymorphic mutation technology is used so that the code will automatically mutate each time it is called, changing into a completely different code from before, but the function remains unchanged, so as to prevent malicious users from performing dynamic analysis, reverse engineering, copying or tampering, and protect the security of the optimized code.
[0065] See also Figure 3 , Figure 3 Another flow chart of a code compression method provided in an embodiment of the present application. Figure 3 shown.
[0066] First, execute entry-option to initialize option.
[0067] That is, start the packaging and building process, read the building configuration, and arrange the building tasks according to the building configuration.
[0068] Then, execute run to start compiling. Execute make to recursively analyze dependencies starting from entry and build each dependent module.
[0069] That is, determine the entry corresponding to the build configuration, recursively analyze the dependencies from the entry, and prepare to start the build subtask of the submodule.
[0070] Then, execute before-resolve to resolve the module location, and execute build-module to start building a module.
[0071] That is, the module positions are parsed and a module index table is constructed, wherein the module index table carries the relative position dependencies between the modules.
[0072] Then, execute mormal-module-loader to compile the module loaded by the loader and generate an AST tree.
[0073] That is, the construction subtask of the submodule is started, and according to the construction configuration and the precompiled code block type, the specified loader is called, the module loaded by the loader is compiled, and an AST syntax tree is generated.
[0074] Then, execute the program, traverse the AST, and collect dependencies when encountering some call expressions such as require.
[0075] That is, traverse the AST syntax tree and collect related dependencies when encountering some call expressions such as require.
[0076] Then, execute seal, all dependencies are built, and optimization begins.
[0077] That is, after all dependencies are built, start the code optimization task.
[0078] Then, perform analysis, AST syntax tree analysis, to find logical redundancy and process dead loops.
[0079] That is, according to the AST syntax tree, simplify the dependency graph, wherein simplifying the dependency graph includes: eliminating useless logic branches according to tree-shaking; analyzing the repeatedly implemented function methods through the dependency registry to simplify the dependency graph; searching whether the directed graph is a cycle according to the deep traversal route analysis, and removing the loop logic if it is a cycle; if there is no cycle, reducing the execution nodes and reducing the depth of the directed graph.
[0080] Then, gzip is executed to obfuscate and compress the code.
[0081] That is, the optimized AST syntax tree is output as optimized code, and a code obfuscation process is prepared to be started, wherein the code obfuscation process is started with the help of an open source obfuscation tool.
[0082] Finally, execute emit and output to the dist directory.
[0083] That is, the text is written into the file in a template compression format, and the template code with the file size compressed is output.
[0084] It should be noted that the build configuration here refers to the packaging process configuration through JSON files or chain calls before packaging. This configuration mainly makes personalized configuration modifications based on the general configuration. For example: specify which Loader to use for loading analysis of the module; whether to use UglifyJS or JScambler for code compression, etc.
[0085] The module index table here is a table that is serialized and stored in a table structure after module dependency analysis (refer to the "module name / function name-target address" table in the upper right corner). For example, a project contains three blocks of code: entry, module1, and module2. The dependency relationship between the three blocks of code can be analyzed through the import keyword. Figure 4 , Figure 4 A schematic diagram of a module index representation in a code compression method provided in an embodiment of the present application.
[0086] The code compression method provided in the embodiment of the present application is applied in combination with modern front-end development engineering methods. Based on the code AST syntax tree, code logic redundancy detection such as repeated function implementation and dead loop of program execution process is completed, the code volume is deeply trimmed, and the obfuscation and compression algorithms are used to finally complete the extreme compression of the code and improve the loading speed.
[0087] Corresponding to the above method embodiment, this specification also provides a code compression device embodiment, Figure 5 A schematic diagram of the structure of a code compression device provided in an embodiment of the present application. Figure 5 As shown, the device comprises:
[0088] The loading module 502 is configured to, when detecting that the code packaging and building process is started, identify the dependencies of each module in the project layer by layer starting from the entry function, and build a module dependency graph, wherein the module dependency graph carries the module's target address table and the dependency relationship between each sub-module;
[0089] The analysis module 504 is configured to analyze the source code file type of the module through the loader and the syntax analyzer, and construct a syntax analysis tree;
[0090] The construction module 506 is configured to eliminate redundant logic branches of the syntax analysis tree by tree-shaking, eliminate duplicate or similar function definitions by similar syntax analysis tree detection, obtain a target syntax analysis tree, and analyze the target syntax analysis tree by deep traversal route, determine the existence of logic dead loop code flow, and eliminate it;
[0091] The compression module 508 is configured to automatically mutate the optimized source code using polymorphic mutation technology to generate a target code text, write the target code text into a file in a target compression format, and output it, wherein the target code has the same function as the source code.
[0092] In an optional embodiment, the analysis module 504 is further configured to:
[0093] According to the source code file type of the module, during the loading process using the loader, each character in the source code is scanned by the analyzer to generate a character sequence, wherein the character sequence carries the words corresponding to each recognized character;
[0094] Based on the generated words, generate the grammar dictionary corresponding to the module;
[0095] The grammar dictionary is analyzed using a syntax analyzer to generate various phrases, and a syntax analysis tree is constructed based on the phrases.
[0096] In an optional embodiment, the construction module 506 is further configured to:
[0097] In the case where it is detected that there are duplicate imported functions between modules, determining that addresses corresponding to the duplicate imported functions are the same address, and storing the addresses in the module dependency graph;
[0098] When it is detected that the execution processes of a first process function and a second process function are similar between modules, a first syntax parse tree is constructed for the first process function, a second syntax parse tree is constructed for the second process function, and the first syntax parse tree and the second syntax parse tree are compared by a similar tree algorithm to identify a target function, wherein the structures of the first syntax parse tree and the second syntax parse tree are similar.
[0099] In an optional embodiment, the construction module 506 is further configured to:
[0100] Whether the analysis process of the target syntax analysis tree forms a loop is analyzed by deep traversal route. If the analysis process forms a loop, the loop logic is removed; if the analysis process does not form a loop, the execution nodes are reduced.
[0101] In an optional embodiment, the construction module 506 is further configured to:
[0102] Collecting export variables, and recording the export variables into module variables corresponding to the syntax analysis tree;
[0103] The syntax analysis tree is traversed to mark whether the module export variable is used, and if the module export variable is not used, the unused export variable is eliminated.
[0104] In an optional embodiment, the construction module 506 is further configured to:
[0105] Analyze the target syntax analysis tree by deep traversal route, randomly select unsearched nodes to start depth-first search, and mark the unsearched nodes as being in the searching state;
[0106] Traversing the adjacent nodes of the unsearched node, and when there is a first adjacent node in the adjacent nodes that is in an unsearched state, starting a depth-first search from the first adjacent node;
[0107] When there is a second adjacent node in the adjacent nodes that is in the searching state, determining that the analysis process corresponding to the second adjacent node forms a loop;
[0108] When each of the adjacent nodes is completed, it is determined that the analysis process of the target analysis tree has no loop.
[0109] The code compression device provided in the embodiment of the present application is applied in combination with modern front-end development engineering methods. Based on the code AST syntax tree, code logic redundancy detection such as repeated function implementation and dead loop of program execution process is completed, the code volume is deeply trimmed, and the obfuscation and compression algorithms are used to finally complete the extreme compression of the code and improve the loading speed.
[0110] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the code compression device, since it is basically similar to the code compression method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the code compression method embodiment.
[0111] Figure 6 This is a block diagram of a computing device provided in an embodiment of the present application. The components of the computing device 600 include but are not limited to a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and the database 650 is used to store data.
[0112] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).
[0113] In one embodiment of the present specification, the above components of the computing device 600 and Figure 6 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 6 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0114] The computing device 600 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 600 may also be a mobile or stationary server.
[0115] The processor 620 is used to execute the following computer executable instructions, which implement the steps of the above-mentioned code compression method when executed by the processor.
[0116] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the computing device embodiment, since it is basically similar to the code compression method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the code compression method embodiment.
[0117] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which can implement the steps of the above-mentioned code compression method when executed by a processor.
[0118] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the computer-readable storage medium embodiment, since it is basically similar to the code compression method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the code compression method embodiment.
[0119] An embodiment of the present specification further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned code compression method.
[0120] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the computer program embodiment, since it is basically similar to the code compression method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the code compression method embodiment.
[0121] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0122] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0123] It should be noted that the above is a description of a specific embodiment of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of the present specification.
[0124] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0125] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A code compression method, characterized in that: include: When the code packaging and building process is detected to be started, starting from the entry function, the dependencies of each module in the project are identified layer by layer, and a module dependency graph is constructed, wherein the module dependency graph carries the target address table of the module and the dependency relationship between each sub-module; Through the loader and syntax analyzer, the source code file type of the module is analyzed and a syntax analysis tree is constructed; Eliminate redundant logic branches of the grammar analysis tree through tree-shaking, eliminate duplicate or similar function definitions through similar grammar analysis tree detection, obtain a target grammar analysis tree, and analyze the target grammar analysis tree through deep traversal routes to determine and eliminate the existence of logical dead loop code flows; The optimized source code is automatically mutated using polymorphic mutation technology to generate a target code text, which is written into a file in a target compression format and output, wherein the target code has the same function as the source code.
2. The method according to claim 1, characterized in that The module source code file type is analyzed by the loader and the syntax analyzer to construct a syntax analysis tree, including: According to the source code file type of the module, during the loading process using the loader, each character in the source code is scanned by the analyzer to generate a character sequence, wherein the character sequence carries the words corresponding to each recognized character; Based on the generated words, generate the grammar dictionary corresponding to the module; The grammar dictionary is analyzed using a syntax analyzer to generate various phrases, and a syntax analysis tree is constructed based on the phrases.
3. The method according to claim 1, characterized in that The method of eliminating duplicate or similar function definitions by detecting similar grammar analysis trees to obtain a target grammar analysis tree includes: In the case where it is detected that there are duplicate imported functions between modules, determining that addresses corresponding to the duplicate imported functions are the same address, and storing the addresses in the module dependency graph; When it is detected that the execution processes of a first process function and a second process function are similar between modules, a first syntax parse tree is constructed for the first process function, a second syntax parse tree is constructed for the second process function, and the first syntax parse tree and the second syntax parse tree are compared through a similar tree algorithm to identify a target function, wherein the first process function and the second process function are execution functions in the execution process, and the structures of the first syntax parse tree and the second syntax parse tree are similar.
4. The method according to claim 1, characterized in that: The target syntax analysis tree is analyzed by deep traversal route, and the code flow with logical dead loop is judged and eliminated, including: Whether the analysis process of the target syntax analysis tree forms a loop is analyzed by deep traversal route. If the analysis process forms a loop, the loop logic is removed; if the analysis process does not form a loop, the execution nodes are reduced.
5. The method according to claim 1, characterized in that The method of eliminating redundant logical branches of the grammar analysis tree by tree-shaking includes: Collecting export variables, and recording the export variables into module variables corresponding to the syntax analysis tree; The syntax analysis tree is traversed to mark whether the module export variable is used, and if the module export variable is not used, the unused export variable is eliminated.
6. The method according to claim 1, characterized in that The target syntax analysis tree is analyzed by deep traversal route, and the code flow with logical dead loop is judged and eliminated, including: Analyze the target syntax analysis tree by deep traversal route, randomly select unsearched nodes to start depth-first search, and mark the unsearched nodes as being in the searching state; Traversing the adjacent nodes of the unsearched node, and when there is a first adjacent node in the adjacent nodes that is in an unsearched state, starting a depth-first search from the first adjacent node; When there is a second adjacent node in the adjacent nodes that is in the searching state, determining that the analysis process corresponding to the second adjacent node forms a loop; When each of the adjacent nodes is completed, it is determined that the analysis process of the target analysis tree has no loop.
7. A code compression device, characterized in that: include: The loading module is configured to, when detecting that the code packaging and building process is started, identify the dependencies of each module in the project layer by layer starting from the entry function, and build a module dependency graph, wherein the module dependency graph carries the module's target address table and the dependency relationship between each sub-module; The analysis module is configured to analyze the source code file type of the module through the loader and the syntax analyzer, and construct a syntax analysis tree; A construction module is configured to eliminate redundant logic branches of the grammar analysis tree by tree-shaking, eliminate duplicate or similar function definitions by similar grammar analysis tree detection, obtain a target grammar analysis tree, and analyze the target grammar analysis tree by deep traversal route, determine the existence of logic dead loop code flow, and eliminate it; The compression module is configured to automatically mutate the optimized source code using polymorphic mutation technology to generate a target code text, write the target code text into a file in a target compression format, and output it, wherein the target code has the same function as the source code.
8. A computing device comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an implementation program for information transmission, and when the program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Source code security detection large model dimensionality reduction method based on homomorphic merging
CN119807721A
Industrial control system configuration cloud compiling method and system
CN120743284A
An industrial control system configuration cloud compiling method and system
CN120743284B