CodeBERT-based JavaScript global identifier conflict static detection and repair method and system
Through the CodeBERT model combined with static analysis technology, JavaScript global identifier conflicts are identified and resolved, which solves the problem of difficult to identify and resolve global identifier conflicts in the existing technology, and realizes a method of detecting and repairing code conflicts earlier.
Patent Information
- Application Number
- CN202510270211.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-17
AI Technical Summary
JavaScript global identifier conflicts often lead to unpredictable results and performance problems in web applications, and prior art is difficult to effectively identify, infer and resolve these conflicts.
The static detection method based on CodeBERT is adopted to identify and resolve global identifier conflicts by extracting JavaScript code, parsing abstract syntax trees, and analyzing the scope and type inference rules of global identifiers.
It enables detection and repair of global identifier conflicts during code writing, reducing performance overhead and code maintenance costs, and improving code quality.
Smart Images

Figure CN120162783A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer software engineering, and relates to a method and system for static detection and repair of JavaScript global identifier conflicts, specifically to a method and system for static detection and repair of JavaScript global identifier conflicts based on CodeBERT. Background Art
[0002] The characteristics of the JavaScript language include weak typing and dynamic features, which are mainly reflected in the following aspects: dynamically adding and deleting object properties, dynamically generating code, and implicit type conversion. First, in terms of dynamically managing object properties, different from other programming languages, JavaScript allows adding or deleting object properties and methods at any time during runtime without pre-defining them during object initialization. Although this feature provides great flexibility in development, it is also prone to errors, especially when misoperating on object properties and methods in actual applications. Second, in terms of dynamically generating code, JavaScript can represent code as a string and use the eval function to parse and execute the string as an actual code snippet. In program analysis, treating eval as an ordinary function will lead to misunderstandings. The correct approach is to treat the parameter of the eval function as the code to be parsed and handle it specifically to avoid errors during the analysis process. Finally, in terms of implicit type conversion, JavaScript allows the type of variables to be automatically converted during program runtime without explicit declaration. For example, in string operations, if a variable of number type is passed in, the program will automatically convert it to string type. Although this implicit conversion mechanism provides flexibility, it is also prone to type-related problems and may cause potential defects during program runtime. Generally speaking, although these features of JavaScript enhance the convenience and flexibility of development, they are also prone to confusion and errors in actual applications and program analysis. Therefore, special care is needed when writing and maintaining code.
[0003] Web applications usually introduce some JavaScript code to implement partial functions. While the extensive introduction of JavaScript facilitates the construction of Web applications, it also brings some problems. In the browser environment, JavaScript code within the same framework shares the same namespace during runtime. Therefore, these pieces of code will affect each other, leading to unexpected results or even crashes in the program. If multiple JavaScript files or libraries define variables, functions, or objects with the same name in the same namespace, naming conflicts will occur, which may lead to unpredictable results and, in severe cases, code overwriting or mutual contamination. Additionally, it may also result in the rewriting of built-in properties and methods, affecting the operation of other scripts.
[0004] In previous research work, Phung et al. embedded security policies into the code by transforming the code and intercepted security-related API calls to control and modify the execution of JavaScript (Reference 1). The server-driven JavaScript sandbox framework JSand proposed by Agten et al. (Reference 2) restricts the permissions of JavaScript code through isolation but cannot effectively solve the problem of global identifier conflicts. Patra et al. studied the global identifier conflicts between JavaScript libraries (Reference 3). This method generates a test client to evaluate whether there will be differences when two libraries are loaded under various settings. In practice, a Web application contains many JavaScript files, and the conflict detection workload for pairwise library pairs is too large. JSObserver proposed by Zhang Mingxue et al. detects global identifier conflicts based on the browser's record of JavaScript code's memory access (Reference 4). This solution requires running the code after the development work is completed to detect conflicts. Developers need to modify the completed code, which will bring a considerable amount of work. Since this method is a dynamic analysis and can detect the code executed during runtime, it cannot detect conflicts in the code that is not triggered by test cases. JSISOLATE studied by Zhang Mingxue et al. prevents conflicts between JavaScript codes by analyzing the dependencies between scripts and isolating non-dependent scripts to different environments for execution (Reference 5). This solution increases the page loading time and introduces additional performance overhead. In addition, these works only detect or isolate conflicts and do not fundamentally fix conflicts.
[0005] References
[0006] [1] Phung P H, Sands D, Chudnov A. Lightweight self - protecting JavaScript[C] / / Proceedings of the 4th International Symposium on Information,
[0007] Computer, and Communications Security. 2009:47 - 60.
[0008] [2] Agten P, Van Acker S, Brondsema Y, et al. JSand: complete client - side sandboxing of third - party JavaScript without browser modifications[C].
[0009] Proceedings of the 28th Annual Computer Security Applications Conference. 2012:1 - 10.
[0011] [3] Patra J, Dixit P N, Pradel M. Conflictjs: finding and understanding conflicts between javascript libraries[C]. Proceedings of the 40th International Conference on Software Engineering. 2018:741 - 751.
[0012] [4] Zhang M, Meng W. Detecting and understanding JavaScript global identifier conflicts on the web[C] / / Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 2020:38 - 49.
[0013] [5]Zhang M,Meng W.JSISOLATE: lightweight in-browser JavaScript isolation[C].
[0014] Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering. 2021:193-204. Summary of the Invention
[0016] To solve one or several technical problems existing in the prior art such as the recognition of JavaScript global identifiers, the inference of global identifier types, and conflict resolution, the present invention provides a static detection and repair method and system for JavaScript global identifier conflicts based on CodeBERT.
[0017] The technical solution adopted by the method of the present invention is as follows: A static detection and repair method for JavaScript global identifier conflicts based on CodeBERT, comprising the following steps:
[0018] Step 1: Extract JavaScript code;
[0019] Step 2: Extract function type information from the API documentation corresponding to the JavaScript code;
[0020] Step 3: Analyze global identifiers;
[0021] Parse the extracted JavaScript code into an abstract syntax tree. By analyzing the abstract syntax tree, find all global identifiers according to the JavaScript identifier scope rules, record the definition and read / write situations of the JavaScript file for global identifiers, and perform global identifier type inference according to the type inference rules and the function type information obtained in Step 2 to form a log file;
[0022] Step 4: Analyze the conflicts between each JavaScript code according to the log file;
[0023] Step 5: Resolve conflicts.
[0024] Preferably, in step 1, all ways of embedding JavaScript code in the web page are analyzed, JavaScript code recognition rules are formulated, all JavaScript code is recognized from the web page, a unique id is given to each JavaScript file, and it is classified according to the namespace.
[0025] Preferably, for the JavaScript code recognition rules, htmlparser2 is used as the core parsing engine, and by constructing a multi-level document object model, several script-related attributes are accurately captured to capture comprehensive JavaScript code.
[0026] Preferably, in step 2, a pre-trained CodeBERT model is used to automatically extract function names, parameter names, return value types, and parameter type information from the corresponding API documentation;
[0027] The specific implementation of the training process of the pre-trained CodeBERT model includes the following sub-steps:
[0028] Step 2.1: Obtain the original data from the target API documentation through web crawler technology, and implement data standardization through multi-level processing, specifically including: deduplication optimization, non-functional text filtering, key metadata annotation, and special symbol normalization processing, to obtain the processed structured data, including function prototypes, parameter type constraints, and return value semantic descriptions;
[0029] Step 2.2: Convert the processed structured data into a sequence of tokens through the BPE tokenization algorithm to form a model input matrix with a unified dimension;
[0030] Step 2.3: Input the sequence of tokens into the CodeBERT model for training. The CodeBERT model adopts a 12-layer Transformer stacking architecture, establishes cross-token associations through the self-attention mechanism, generates context-sensitive feature vectors, and accurately encodes the syntax tree structure and functional semantics of API elements;
[0031] Step 2.4: Perform feature decoding, project the high-dimensional vector space to structured labels using the Softmax layer, and optimize the CodeBERT model through transfer learning and fine-tuning techniques. After training for a preset number of rounds, a trained CodeBERT model is obtained.
[0032] Preferably, in step 3, for the JavaScript identifier scope rule, for the method of defining global identifiers in JavaScript, if it is defined outside the function for and if code blocks, or defined using the var keyword inside the for and if code blocks, then use Babel to parse the JavaScript code into an abstract syntax tree; by checking the scope.bindings object of the root node, determine whether the identifier belongs to a global identifier; for the method of defining global identifiers in JavaScript, if directly assigning a value to an undefined variable or adding a property to the window object through the dot operator, first check whether the scope.bindings object of the parent node contains the identifier. If not, continue to search upward. If the root node is reached, then determine that the identifier is a global identifier; at the same time, record that the identifier is defined by this file; if the identifier or its alias appears on the left side of an assignment statement, record that the script has written to the identifier; otherwise, as long as the identifier appears, record it as a read operation; record the operations of the identifier in a log file.
[0033] Preferably, in step 3, the type inference rules include:
[0034] Expression: Literal assignment: a = 123 / '123' / true / [1, 2, 3] / {1, 2, 3}; Inference result: The type of a is directly determined by the literal.
[0035] Expression: a++ / a--; Inference result: Increment and decrement operations, then a is of type number.
[0036] Expression: function a() {}; Inference result: Function definition, then a is of type function.
[0037] Expression: a = b + - * / % c; Inference result: If both b and c are of type number, then a is of type number. If one of b and c is of type string, then b + c is of type string, and the result of b - * / % c is still of type number;
[0038] Expression: a += -= *= / = b; Inference result: If both a and b are of type number, then a is of type number after the operation. If one of a and b is of type string, then a is of type string after the += operation, and a is of type number after the -= *= / = operations;
[0039] Expression: a = b; Inference result: Infer that the type of a is the same as the type of b;
[0040] Expression: a < / <= / > / >= b; Inference result: If the type of one of a and b can be determined, then it is inferred that the other is also of that type;
[0041] Expression: a = myFunction(b, c); Inference result: If a is equal to the return value of a function, then the type of a is the same as the return value of that function. b and c are function parameters, so the types of b and c are the same as the parameter types of the function. The return value and parameter type information of the function are obtained by parsing the corresponding API documentation through CodeBERT.
[0042] Preferably, in step 5, conflicts are resolved according to the conflict detection results. First, rename the conflicting identifiers, replace the conflicting identifiers in the code with the [MASK] tag, and use the original pre-training objective MLM of the pre-trained model CodeBERT to predict the correct identifier that the developer actually wants to reference at the conflict position to fix the conflict.
[0043] Preferably, for fixing the conflict, input the code snippet with [MASK] into the pre-trained model CodeBERT. The model CodeBERT predicts the correct identifier at the [MASK] position based on the context, replaces the prediction result back into the code, and generates the repaired version. Re-run the conflict detection to ensure that there are no conflicts after repair.
[0044] The technical solution adopted by the system of the present invention is: A static detection and repair system for JavaScript global identifier conflicts based on CodeBERT, including:
[0045] One or more processors;
[0046] A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the static detection and repair method for JavaScript global identifier conflicts based on CodeBERT.
[0047] The technical solution adopted by the product of the present invention is: A computer program product including computer program instructions, which when running on a computer, cause the computer to execute the static detection and repair method for JavaScript global identifier conflicts based on CodeBERT.
[0048] The present invention has the following advantages over the existing methods for detecting JavaScript global identifier conflicts:
[0049] (1) The present invention performs conflict detection and repair based on static analysis, allowing developers to detect and repair global identifier conflicts during the coding process, rather than waiting until the project development is completed and enters the testing phase to discover the existence of conflicts. It provides a method for front-end development to detect and resolve code conflicts earlier, which helps to reduce performance overhead and code maintenance costs, thereby promoting the improvement of code quality in front-end development.
[0050] (2) The type inference of the present invention is more accurate. The present invention utilizes information such as function names, parameter names, return value types, and parameter types included in API documents in type inference to perform type inference, improving the type inference rules and enhancing the accuracy of type inference.
[0051] (3) The present invention fundamentally resolves global identifier conflicts, rather than merely detecting or isolating conflicts. Brief Description of the Drawings
[0052] The following uses embodiments and specific implementation manners to further illustrate the technical solution of the present invention. Additionally, during the process of explaining the technical solution, some drawings are also used. For those skilled in the art, without creative efforts, other drawings and the intent of the present invention can also be obtained based on these drawings.
[0053] Figure 1 It is a flowchart of the method for an embodiment of the present invention. Detailed Description of the Preferred Embodiment
[0054] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0055] The key technologies involved in this embodiment include program static analysis technology and type inference technology.
[0056] Program static analysis is a code analysis technology, corresponding to dynamic analysis, and its main feature is that it does not actually execute the program. Static analysis discovers potential program hazards by automatically scanning the code. In the early stage of development, static analysis can detect many vulnerability problems caused by coding defects (such as buffer overflows), thereby significantly reducing the cost of later testing and shortening the product's time to market.
[0057] Unlike dynamic analysis which requires executing the program in a simulated or real environment, static analysis only relies on static scanning of the source code to complete program analysis. Therefore, static analysis is usually used for code quality checking in the early stage, while dynamic analysis is more often used in scenarios such as functional testing, performance testing, and memory leak testing. The advantage of static analysis lies in its fast execution speed and high efficiency. Currently, relatively mature static analysis tools can scan tens of thousands of lines of code per second when dealing with complex code. This efficient analysis ability is of great significance for the development and maintenance of large-scale Internet applications. However, static analysis also has certain limitations, and its false positive rate is usually high. This is because static analysis scans and matches the program based on preset rule patterns to identify potentially high-risk code, rather than necessarily code with actual errors. The level of the false positive rate depends on whether the rules are reasonably formulated. To improve the detection accuracy, static analysis can be combined with dynamic analysis, which can reduce false positives while more comprehensively identifying potential problems in the program.
[0058] Type Inference Technology: A type system can be understood as representing a set of values. If the security attributes of variables are corresponded to their types, operations that violate security specifications can be identified through type checking and type inference. A complete type system usually consists of multiple parts such as type syntax, static semantics, and dynamic semantics. Among them, static defect detection mainly focuses on type static semantics, including typing assertions, checking rules, and inference rules. A typing assertion means that in a static typing environment, a certain variable has a specific type; checking rules are used to determine whether an operation is a legal behavior; while inference rules are used to draw conclusion B when premise A is satisfied. Since different inference rules may derive different types for the same expression, a type inference algorithm needs to be designed to determine how to apply these inference rules. Most type systems are context-insensitive and flow-insensitive, that is, in the whole program, each variable can only have one type. Context-sensitive type inference can utilize the context information at the function call point to determine the types of functions and their internal variables; flow-sensitive type inference allows variables to have different types at different program points. In a type system, there can be a subtype relationship between types. If type A is a subtype of type B, then variables of type A can replace variables of type B at any time, but this replacement relationship is one-way. To extend the existing type system without the need to rebuild a completely new system, type qualifiers are introduced into defect detection based on type inference. Type qualifiers divide the set of values of a variable type into several non-overlapping subsets, that is, subtypes of the original type, which are used to represent different security levels. During the static detection process, the user first assigns type qualifiers to some variables, then uses the type inference algorithm for inference, and performs type checking on the inference results to identify those operations that violate the predefined rules. This type inference method is simple and efficient, and is especially suitable for quickly detecting software security vulnerabilities.
[0059] Features of the JavaScript language: The weak typing and dynamic nature of JavaScript are mainly reflected in the following aspects: dynamically adding and deleting object properties, dynamically generating code, and implicit type conversion. First, in terms of dynamically managing object properties, different from other programming languages, JavaScript allows properties and methods of an object to be added or deleted at any time during runtime, without the need to be predefined during object initialization. Although this feature provides great flexibility in development, it is also prone to errors, especially when misoperating on the properties and methods of an object in practical applications. Second, in terms of dynamically generating code, JavaScript can represent code as a string and use the eval function to parse and execute the string as an actual code snippet. In program analysis, if eval is treated as an ordinary function, it will lead to misunderstandings. The correct approach is to treat the parameter of the eval function as the code to be parsed and handle it specifically, which can avoid errors during the analysis process. Finally, in terms of implicit type conversion, JavaScript allows the type of a variable to be automatically converted during program execution without explicit declaration. For example, in string operations, if a variable of number type is passed in, the program will automatically convert it to string type. Although this implicit conversion mechanism provides flexibility, it is also prone to type-related problems and may cause potential defects during program runtime. Generally speaking, although these features of JavaScript enhance the convenience and flexibility of development, they are also prone to confusion and errors in practical applications and program analysis. Therefore, special care needs to be taken when writing and maintaining code.
[0060] The technical problems to be solved in this embodiment include:
[0061] (1) Identification of JavaScript global identifiers. To identify global identifiers in JavaScript, first, the JavaScript code in the Web page must be extracted. The Web page can write the JavaScript code directly in the script tag through inline scripts, or specify the path of the JavaScript code through the src attribute of the script tag, or dynamically add and modify JavaScript through Dom (Document Object Model) operations. When performing conflict detection, all JavaScript code must be comprehensively covered. Second, global identifiers need to be identified. There are various ways to define global identifiers in JavaScript. Using the var keyword to define identifiers, adding properties to the window object through the dot operator, directly assigning values to variables that have not been defined, etc. can all define a global identifier. The present invention needs to design global identifier detection rules for global identifier identification.
[0062] (2) Global identifier type inference. JavaScript is a weakly typed programming language. When declaring a variable, it is not necessary to specify the variable type, and a value of a different type from the original can be assigned to a variable during assignment without causing errors or warnings. Therefore, it is difficult to determine the type of a variable through static analysis. During conflict detection, it is necessary to infer the type of the global identifier to detect type conflicts.
[0063] (3) Resolving conflicts. When a global identifier conflict is detected, it is necessary to analyze the developer's intention and determine the correct identifier that the developer actually wants to reference at the conflict location based on the context of the code.
[0064] See Figure 1 , a static detection and repair method for JavaScript global identifier conflicts based on CodeBERT provided in this embodiment includes the following steps:
[0065] Step 1: Extract JavaScript code;
[0066] In one implementation, obtain the source code of the Web page to be analyzed, including HTML files, external JavaScript file links, inline scripts, and dynamically generated scripts. Use the htmlparser2 library of Python to parse the HTML document. Extract all <script>标签内容(内联脚本)及src属性指向的外部JS文件。监听DOM动态操作(如document.createElement('script')),捕获运行时加载的JavaScript代码。
[0067] 分析web页面嵌入JavaScript代码的所有方式,制定JavaScript代码识别规则,从web页面中识别出所有的JavaScript代码,给每个JavaScript文件一个独一无二的id,并按照命名空间分类。
[0068] 在构建网页时,JavaScript代码的包含机制分为静态包含和动态包含。其中静态包含方法主要包括以下两种实现方式:(1)基于<script>标签的嵌入式引入,即通过标签内联方式直接在HTML文档中嵌入可执行代码;(2)外部文件引用模式,利用<script>标签的src属性指定本地或远程的JavaScript资源路径。动态包含通过DOM API实现运行时脚本注入,通过动态修改DOM树结构实现脚本的按需加载。
[0069] 在一种实施方式中,设计并实现了一套JavaScript代码识别规则。采用htmlparser2作为核心解析引擎,通过构建多层级文档对象模型,精确捕获11种脚本相关属性(详见表1和表2),从而能够捕获更加全面的JavaScript代码。
[0070] 表1Html页面静态包含JS代码的方法
[0071]
[0072] 表2Html页面动态包含JS代码的方法
[0073]
[0074]
[0075] 步骤2:从JavaScript代码对应的API文档中提取函数类型信息;
[0076] 在一种实施方式中,使用预训练的CodeBERT模型,从相应的API文档中自动提取函数名称,参数名称,返回值类型和参数类型信息,以帮助下一步做类型推断;
[0077] 在一种实施方式中,预训练CodeBERT模型的数据来源于从网络上爬取的10个流行的结构化JS库的API文档。爬取的数据包含大量无关信息,例如示例代码片段和其他非必要内容。因此,第一步是去除这些无关数据,只保留包含函数名称和类型信息的条目。其次,API文档中某些函数重复,导致爬取到冗余条目,因此需要剔除重复样本,以确保数据集仅包含相关的不重复的信息,以供分析和模型训练。
[0078] 由于预测实体边界需要同时预测实体类型,因此有9个标签需要预测,分别是:"O”、"B-FUNC”、"I-FUNC”、"B-RET-TYPE”、"I-RET-TYPE”、"B-ARG”、"I-ARG”、"B-ARG-TYPE”和"I-ARG-TYPE”。在测试过程中,只有当实体的边界和类型都完全准确时,才认为该实体的预测正确。
[0079] 在一种实施方式中,采用分层处理架构实现API文档智能解析。首先通过爬虫技术从目标API文档中获取原始数据,通过多级处理实现数据标准化,具体包括:去重优化、非功能性文本过滤(如版本号、版权声明)、关键元数据标注(函数名、参数类型等)、特殊符号规范化(如统一的制表符和空格)。
[0080] 接下来,处理后的结构化数据包含函数原型、参数类型约束和返回值语义描述,通过BPE词汇切分算法转化为词元序列,形成统一维度的模型输入矩阵。
[0081] 第三步,将预处理后的词元序列输入到CodeBERT多模态理解框架中。该模型采用12层Transformer堆叠架构,通过自注意力机制建立跨词元关联,生成上下文敏感的特征向量,准确编码API元素的语法树结构和功能语义。
[0082] 最后进行特征解码,利用Softmax层将高维向量空间投影到结构化标签集(包括功能类别、参数属性等细粒度标签),并通过迁移学习和微调技术对所述CodeBERT模型进行优化,训练预设轮次后,获得训练好的CodeBERT模型。
[0083] 训练过程中,从API文档提取类型信息(CodeBERT微调)。训练模型从API文档中提取函数类型信息。收集JavaScript API文档(如MDN、React官方文档),构建标注数据集。在CodeBERT基础上,使用标注数据训练分类任务(参数类型分类、返回值类型分类)。提升模型对JavaScript API语义的理解能力。
[0084] 使用微调后的CodeBERT模型,从JavaScript API文档(如MDN Web Docs)提取函数参数类型、返回值类型等信息在结合类型推断规则推断全局标识符的类型(如string、number、object),并分析其读写操作。例如:
[0085] 提取Array.prototype.map的返回值类型为Array。
[0086] 上下文关联:结合函数调用链(如fetch().then()推断返回值为Promise)。
[0087] 读写分析:记录标识符的修改操作(如user.name="Alice")及引用位置。
[0088] 步骤3:分析全局标识符;
[0089] 将提取的JavaScript代码解析为抽象语法树(AST),通过分析抽象语法树,根据JavaScript标识符作用域规则找出所有的全局标识符,记录JavaScript文件对全局标识符的定义和读写情况,并根据类型推断规则和步骤2中得到的函数类型信息进行全局标识符类型推断,形成日志文件;
[0090] 在一种实施方式中,使用开源JavaScript解析器Babel将代码解析为AST。
[0091] 在一种实施方式中,所述JavaScript标识符作用域规则,存在多种定义全局标识符的方式,针对表3中的前两种定义方式,我们首先使用Babel将JavaScript代码解析成抽象语法树。通过检查根节点的scope.bindings对象,我们能够判断一个标识符是否属于全局标识符。对于后两种定义方式,由于它们并非变量定义语句而只是赋值语句,我们需要先检查父节点的scope.bindings对象是否包含该标识符,如果不包含,我们将继续向上查找,若一直找到根节点,则判断该标识符是一个全局标识符。同时记录该标识符被该文件定义。如果标识符或者它的别名出现在赋值语句的左边,则记录该脚本对该标识符进行了写操作,除此之外,只要标识符出现,都记录为读操作。将标识符的操作记录到日志文件中。
[0092] 表3JavaScript中定义全局标识符的方法
[0093]
[0094] 所述类型推断规则,JavaScript中标识符可分为以下类型:number、string、boolean、function、object、undefined以及symbol。在一种实施方式中,根据表4中的类型推断规则来推导全局标识符的类型。使用表4展示的格式来记录标识符的定义、读写操作的检测以及类型推断的结果。
[0095] 表4类型推断规则表
[0096]
[0097]
[0098] 识别所有全局作用域中定义的变量、函数及对象属性。遍历AST节点,检测通过var声明的变量(未在函数或块级作用域内)。识别直接赋值给window对象的属性(如window.user={})。捕获未声明直接赋值的变量(如counter=0)。输出结果:生成全局标识符列表,记录其定义位置(文件ID、行号)及初始值。
[0099] 步骤4:根据日志文件分析各个JavaScript代码之间的冲突;
[0100] 在一种实施方式中,检测全局标识符在名称或类型上的冲突。定义冲突:检查不同文件中是否定义了同名全局标识符(如file1.js和file2.js均定义config)。值冲突:分析同名标识符是否被不同的JavaScript文件多次赋值(如libA.version="1.0";libB.version="2.0")。类型冲突:不同的文件中同名全局标识符的类型不同(如file1.js中logger为函数,file2.js中logger为对象)。
[0101] 输出日志:生成冲突报告,标注冲突类型、位置及上下文代码片段。
[0102] 步骤5:解决冲突;
[0103] 在一种实施方式中,根据冲突检测结果解决冲突;首先重命名互相冲突的标识符,将代码中冲突的标识符替换成[MASK]标记(如将window.user替换为[MASK].user),利用预训练模型CodeBERT的原始预训练目标MLM来预测开发者在冲突位置实际想要引用的正确标识符来修复冲突;所述修复冲突,输入带[MASK]的代码片段至预训练模型CodeBERT,模型CodeBERT基于上下文预测[MASK]位置的正确标识符(如预测为libA.user),将预测结果替换回代码,生成修复后的版本;重新运行冲突检测,确保修复后无冲突。
[0104] 例如:假设检测到以下冲突:
[0105] 文件A:var config={env:"dev"};
[0106] 文件B:window.config={debug:true};
[0107] 修复过程:
[0108] 替换冲突标识符:[MASK].config={...}
[0109] CodeBERT预测[MASK]为window(基于上下文高频引用window对象)。
[0110] 自动修复为window.config={debug:true};,并在文件A中保留原定义或提示开发者调整作用域。
[0111] 本实施例还提供了一种基于CodeBERT的JavaScript全局标识符冲突静态检测与修复系统,包括:
[0112] 一个或多个处理器;
[0113] 存储装置,用于存储一个或多个程序,当所述一个或多个程序被所述一个或多个处理器执行时,使得所述一个或多个处理器实现所述的基于CodeBERT的JavaScript全局标识符冲突静态检测与修复方法。
[0114] 本实施例还提供了一种计算机程序产品,包括计算机程序指令,当所述计算机程序指令在计算机上运行时,使得计算机执行所述的基于CodeBERT的JavaScript全局标识符冲突静态检测与修复方法。
[0115] 本发明的创新点包括:
[0116] 1.Web页面中JavaScript代码提取规则与JavaScript全局标识符的识别规则;
[0117] 2.利用CodeBERT模型从API文档中提取的函数类型信息做类型推断以检测类型冲突。将预训练语言模型(CodeBERT)用于JavaScript类型推断,突破传统静态分析的局限性。通过API文档的语义信息补充,显著提升类型推断准确性;
[0118] 3.利用预训练模型(CodeBERT)的原始预训练目标(MLM)来预测开发者在冲突位置实际想要引用的正确标识符以修复全局标识符冲突。首次将MLM任务应用于代码冲突修复,通过上下文语义理解实现精准预测。修复过程无需人工干预,且与开发阶段无缝集成。
[0119] 为了证明本发明相对于现有技术的技术优势,以下通过具体实验对本发明与与文献"Zhang M,Meng W.Detecting and understanding JavaScript global identifierconflicts on the web[C] / / Proceedings of the 28th ACM Joint Meeting onEuropean Software Engineering Conference and Symposium on the Foundations ofSoftware Engineering.2020:38-49.”提出的方法进行比较分析。
[0120] 本实验在1000个真实网站样本集上对两类工具的检测准确率、执行效率、环境兼容性等进行了系统的对比分析。对比分析结果如表5所示。实验数据表明,本发明共检测到2618个潜在冲突,而JSOBSERVER则识别出2485个冲突,其中2353个(占比85.56%)为两种工具均检测到的有效冲突,说明两种工具在基本检测能力上具有较高的一致性。通过工具差异分析可知,本发明多检测到的265个冲突均是由非测试用例触发的JavaScript代码段引起的,这类静态但未执行的代码超出了JSOBSERVER的动态监测范围。
[0121] 表5
[0122]
[0123] 应当理解的是,上述描述的实施例是本发明一部分实施例,而不是全部的实施例。另外,本发明提供的各个实施例或单个实施例中的技术特征可以相互任意结合,以形成可行的技术方案,这种结合不受步骤先后次序和 / 或结构组成模式的约束,但是必须是以本领域普通技术人员能够实现为基础,当技术方案的结合出现相互矛盾或无法实现时,应当认为这种技术方案的结合不存在,也不在本发明要求的保护范围之内。
[0124] 应当理解的是,上述针对较佳实施例的描述较为详细,并不能因此而认为是对本发明专利保护范围的限制,本领域的普通技术人员在本发明的启示下,在不脱离本发明权利要求所保护的范围情况下,还可以做出替换或变形,均落入本发明的保护范围之内,本发明的请求保护范围应以所附权利要求为准。< / script>
Claims
1. A static detection and repair method for JavaScript global identifier conflicts based on CodeBERT, characterized in that: The following steps are involved: Step 1: Extract JavaScript code; Step 2: Extract function type information from the API documentation corresponding to the JavaScript code; Step 3: Analyze the global identifier; Parse the extracted JavaScript code into an abstract syntax tree. By analyzing the abstract syntax tree, find all global identifiers according to the JavaScript identifier scope rules, record the definition and reading and writing of global identifiers in the JavaScript file, and perform global identifier type inference according to the type inference rules and the function type information obtained in step 2 to form a log file. Step 4: Analyze the conflicts between various JavaScript codes based on the log files; Step 5: Resolve conflicts.
2. The CodeBERT-based JavaScript global identifier conflict static detection and repair method according to claim 1, characterized in that: In step 1, all the ways of embedding JavaScript code in web pages are analyzed, and JavaScript code identification rules are formulated to identify all JavaScript codes from web pages, give each JavaScript file a unique id, and classify them according to namespaces.
3. The CodeBERT-based JavaScript global identifier conflict static detection and repair method according to claim 2, characterized in that: The JavaScript code recognition rule adopts htmlparser2 as the core parsing engine, and accurately captures several script-related attributes and comprehensive JavaScript codes by building a multi-level document object model.
4. The CodeBERT-based JavaScript global identifier conflict static detection and repair method according to claim 1, characterized in that: In step 2, the pre-trained CodeBERT model is used to automatically extract function names, parameter names, return value types, and parameter type information from the corresponding API documents; The training process of the pre-trained CodeBERT model includes the following sub-steps: Step 2.1: Use crawler technology to obtain raw data from the target API document, and implement data standardization through multi-level processing, including: deduplication optimization, non-functional text filtering, key metadata annotation and special symbol normalization processing, to obtain processed structured data, including function prototypes, parameter type constraints and return value semantic descriptions; Step 2.2: Convert the processed structured data into word sequence through the BPE lexical segmentation algorithm to form a model input matrix of uniform dimension; Step 2.3: Input the word sequence into the CodeBERT model for training. The CodeBERT model uses a 12-layer Transformer stacking architecture to establish cross-word associations through a self-attention mechanism, generate context-sensitive feature vectors, and accurately encode the syntax tree structure and functional semantics of the API elements. Step 2.4: Perform feature decoding, use the Softmax layer to project the high-dimensional vector space to the structured label, and optimize the CodeBERT model through transfer learning and fine-tuning technology. After training for a preset number of rounds, a trained CodeBERT model is obtained.
5. The CodeBERT-based JavaScript global identifier conflict static detection and repair method according to claim 1, characterized in that: In step 3, the JavaScript identifier scope rule, for the method of defining global identifiers in JavaScript, if it is defined outside the function for, if code block, or defined using the var keyword within the for, if code block, then use Babel to parse the JavaScript code into an abstract syntax tree; by checking the scope.bindings object of the root node, determine whether the identifier is a global identifier; for the method of defining global identifiers in JavaScript, if you directly assign a value to an undefined variable, or add an attribute to the window object through the dot operator, first check whether the scope.bindings object of the parent node contains the identifier, if not, continue to search upward, if the root node is found, it is determined that the identifier is a global identifier; at the same time, record that the identifier is defined by the file; if the identifier or its alias appears on the left side of the assignment statement, then record that the script has performed a write operation on the identifier; in addition, as long as the identifier appears, it is recorded as a read operation; Logs the identifier's actions to a log file.
6. The CodeBERT-based JavaScript global identifier conflict static detection and repair method according to claim 1, characterized in that: In step 3, the type inference rules include: Expression: Literal assignment: a=123 / '123' / true / [1,2,3] / {1,2,3}; Inference result: The type of a is directly determined by the literal; Expression: a++ / a--; Inference result: self-increment and self-decrement operation, so a is of type number; Expression: function a(){}; Inference result: function definition, then a is of function type; Expression: a=b+-* / %c; Inference result: If b and c are both numbers, then a is a number. If one of b and c is a string, then b+c is a string, and b-* / %c is still a number. Expression: a+=-=*= / =b; Inference result: If a and b are both numbers, then a is a number after the operation. If one of a and b is a string, then a is a string after the += operation, and a is a number after the -=*= / = operation. Expression: a=b; Inference result: The type of a is inferred to be the same as the type of b; Expression: a < / <= / > / >=b; Inference result: If the type of one of a and b can be determined, then the other one is also inferred to be of the same type; Expression: a = myFunction(b,c); Inference result: If a is equal to the return value of a function, then the type of a is the same as the return value of the function. If b and c are function parameters, then the types of b and c are the same as the function parameter types. The function's return value and parameter type information are obtained by parsing the corresponding API document through CodeBERT.
7. The CodeBERT-based JavaScript global identifier conflict static detection and repair method according to any one of claims 1 to 6, characterized in that: In step 5, the conflict is resolved based on the conflict detection results. First, the conflicting identifiers are renamed, and the conflicting identifiers in the code are replaced with [MASK] tags. The original pre-trained target MLM of the pre-trained model CodeBERT is used to predict the correct identifier that the developer actually wants to reference at the conflicting location to fix the conflict.
8. The CodeBERT-based JavaScript global identifier conflict static detection and repair method according to claim 7, characterized in that: The conflict repair method inputs the code snippet with [MASK] into the pre-trained model CodeBERT. The model CodeBERT predicts the correct identifier of the [MASK] position based on the context, replaces the predicted result back into the code, and generates a repaired version. Rerun conflict detection to ensure there are no conflicts after the repair.
9. A JavaScript global identifier conflict static detection and repair system based on CodeBERT, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the CodeBERT-based JavaScript global identifier conflict static detection and repair method as described in any one of claims 1 to 8.
10. A computer program product comprising computer program instructions, characterized in that: When the computer program instructions are executed on a computer, the computer is enabled to execute the CodeBERT-based JavaScript global identifier conflict static detection and repair method according to any one of claims 1 to 8.