Component packaging method and device
By extracting semantic feature information from programming language and natural language development documents, automatically identifying and encapsulating duplicate code, it solves the problem of redundant code in the project, improves code reuse rate and maintenance efficiency, and is suitable for component-based transformation of large front-end projects.
Patent Information
- Application Number
- CN202510894933.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-14
AI Technical Summary
In the existing technology, there is a large amount of redundant and duplicate code in the project, resulting in low code reuse rate, low maintenance efficiency, and difficulty in efficiently identifying and encapsulating duplicate code.
By extracting semantic feature information from development documents written in programming languages and natural languages, judging module and node similarity based on structural similarity and semantic similarity, it automatically identifies and encapsulates code snippets with the same functionality to form reusable components.
It significantly improves code reuse and maintenance efficiency, reduces code redundancy, ensures packaging quality and applicability, and supports component-based transformation of large front-end projects.
Smart Images

Figure CN120780294A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular to a component packaging method and device. Background Art
[0002] As project requirements continue to iterate, the project scale and code volume are also expanding. As a result, more and more redundant and duplicate codes appear in the project, resulting in a lot of time required to deal with similar or duplicate codes during later project maintenance.
[0003] Therefore, how to identify and encapsulate duplicate code in projects and improve code reuse is an urgent problem that needs to be solved. Summary of the Invention
[0004] The present invention provides a component packaging method and device for identifying and packaging repeated codes in a project, thereby improving code reuse rate.
[0005] In a first aspect, the present invention provides a component encapsulation method, which is applicable to multiple development documents with different functional modules, the method comprising: for any functional module, extracting first semantic feature information corresponding to the functional module from a first development document, and extracting second semantic feature information corresponding to the functional module from a second development document; the first development document is a development document recorded in a programming language; the second development document is a development document recorded in a natural language; based on the first semantic feature information and the second semantic feature information corresponding to the first functional module and the second functional module respectively, determining the module similarity between the first functional module and the second functional module; if the module similarity meets a first set condition, determining the node similarity between the first node in the first functional module and the second node in the second functional module; if the node similarity is higher than the module similarity, determining the first node and the second node as encapsulated nodes with the same functionality; wherein the first functional module and the second functional module are any two of different functional modules; and performing component encapsulation on any encapsulated node.
[0006] This invention converts highly repetitive code modules into reusable components and automatically encapsulates code nodes with identical functionality, significantly reducing code redundancy and improving code reuse and maintenance efficiency. Compared to traditional manual identification and encapsulation methods, this automated approach based on multi-source semantic analysis not only improves component encapsulation efficiency but also ensures encapsulation quality through similarity judgment criteria, providing strong support for the component-based transformation of large-scale front-end projects.
[0007] The present invention extracts semantic feature information from different types of development documents. By simultaneously analyzing development documents recorded in programming languages (such as code libraries) and development documents recorded in natural languages (such as requirement documents or UI design documents), it establishes a semantic bridge between code implementation and business needs, thereby more comprehensively understanding the purpose and implementation method of functional modules and significantly improving the accuracy and applicability of component encapsulation.
[0008] Optionally, determining the module similarity between the first functional module and the second functional module includes: determining the structural similarity and semantic similarity between the first functional module and the second functional module; the module similarity meeting the first set condition includes: the structural similarity meeting the first threshold and the semantic similarity meeting the second threshold.
[0009] Through the above scheme, the structural similarity and semantic similarity between functional modules are taken into consideration at the same time, and the corresponding threshold is set as the judgment standard, which significantly improves the accuracy and practicality of component encapsulation. Among them, structural similarity focuses on the formal characteristics of the code, such as document structure, tag nesting and style rules, while semantic similarity focuses on the functional implementation, business logic and interactive behavior of the code. This dual-dimensional similarity evaluation method avoids the incorrect encapsulation caused by only being based on a single dimension. In addition, the present application can also identify code fragments that are "different in structure but similar in function" or "similar in structure but different in function", providing developers with more encapsulation options, making component encapsulation more flexible and adaptable to the needs of different scenarios.
[0010] Optionally, node similarity is the attention weight between any two nodes; module similarity is a comprehensive similarity determined based on structural similarity and semantic similarity.
[0011] By quantifying node similarity as attention weights, this approach captures the strength of associations between code snippets. This allows similarity calculations to move beyond simple syntactic comparisons and instead reflect the actual importance and relevance of nodes in functional implementation. This attention weighting mechanism automatically strengthens connections between key nodes while simultaneously deemphasizing the influence of less important nodes, thereby improving the accuracy of identifying functional boundaries during component encapsulation.
[0012] Optionally, structural similarity and semantic similarity are determined in the following manner: for any functional module, global semantic feature information is determined based on the first semantic feature information and the second semantic feature information; structural similarity is determined based on the cosine similarity of the global semantic feature information of the first functional module and the global semantic feature information of the second functional module; semantic similarity is determined based on the weighted value of the cosine similarity of the first semantic feature information of the first functional module and the first semantic feature information of the second functional module and the cosine similarity of the second semantic feature information of the first functional module and the second semantic feature information of the second functional module.
[0013] By the above scheme, using cosine similarity to calculate similarity can accurately reflect the similarity degree of different function modules, and a two-dimensional similarity evaluation mechanism is adopted to provide a more accurate judgment standard for component packaging, which considers both the formal similarity of code structure and the semantic consistency of code function implementation, thereby more accurately identifying code modules suitable for packaging as components.
[0014] Optionally, component packaging for any packable node includes: packaging the same code part of the packable node in at least two function modules as a component constant; and abstracting a component parameter from different code parts of the packable node in at least two function modules, thereby obtaining a packaged component.
[0015] By the above scheme, the same code part in multiple function modules is packaged as a component constant, which reduces the amount of repeated code and improves the overall quality and maintainability of project code. When common functions need to be modified, only one packaged component needs to be modified, which can achieve unified updating of all places using the component, significantly improving code maintenance efficiency; the mechanism of abstracting a component parameter from different code parts ensures the flexibility and adaptability of the component. By parameterizing the changing parts (such as data sources, configuration options, style themes, etc.), the same component can exhibit different behaviors and appearances through different parameters in different scenarios. This parameterized design not only improves the reuse range of the component, but also preserves the unique needs of each function module, realizing the "one development, multiple use" mode.
[0016] Optionally, abstracting a component parameter from different code parts of the packable node in at least two function modules includes: extracting local semantic feature information from context information of the packable node; extracting global semantic feature information from at least two function modules; and determining the component parameter based on the local semantic feature information and the global semantic feature information.
[0017] By the above scheme, the local semantic feature information can capture the unique needs and usage scenarios of the packable node according to the context environment around the packable node; and the global semantic feature information identifies the parameter configurations commonly needed in different scenarios from the overall perspective of multiple function modules. This dual semantic analysis method of local + global makes the abstracted component parameter meet the needs of specific scenarios and have cross-scene versatility. Secondly, the present application significantly improves the intelligence of component packaging. By semantic feature analysis rather than simple code structure analysis, the actual function and business meaning of the code are understood, thereby more accurately identifying which variables should be component parameters. This semantic-based parameter abstraction method avoids errors that may be caused by relying only on variable names or positions in traditional methods, making the packaged component more consistent with actual business needs.
[0018] Optionally, determining that the first node and the second node are encapsulated nodes with the same functionality includes: determining the encapsulated node based on the abstract syntax tree of the first development document; or determining the encapsulated node based on the abstract syntax tree of the first development document and a knowledge graph formed based on multiple development documents.
[0019] Through the above scheme, the use of abstract syntax can capture the grammatical structure and logical relationship of the code. Compared with simple text matching or keyword search, the present invention can more accurately identify code fragments with the same functional characteristics, reduce the possibility of misjudgment, and improve the accuracy of component encapsulation. Secondly, the introduction of knowledge graph as an optional scheme further enhances the ability to identify complex nested relationships. Due to the limitations of relying solely on abstract syntax tree analysis, the abstract syntax tree method cannot cover complex component nested relationships, while the knowledge graph can capture implicit nested relationships and cross-modal scenarios that are difficult for the abstract syntax tree to identify; therefore, this dual recognition mechanism can cope with code structures of various complexities, that is, for simple nested relationships, only the abstract syntax tree can be used for rapid recognition; and for complex nested relationships, a more comprehensive analysis can be performed in combination with the knowledge graph, thus expanding the scope of application of component encapsulation and improving the coverage and accuracy of component encapsulation.
[0020] Optionally, a knowledge graph formed based on multiple development documents includes: forming entities and relationships between entities in the knowledge graph based on multiple development documents; determining entities corresponding to containment relationships in the relationships between entities; for each entity in the entities with containment relationships: determining the number of child nodes of the entity and the maximum path length from the entity to the deepest child node; determining a first score of the entity based on the number of child nodes and the maximum path length; if the first score is greater than a third threshold, determining that the node corresponding to the entity is an encapsulated node.
[0021] Through the above scheme, by extracting entities and their relationships from multiple development documents to form a knowledge graph, it is possible to evaluate the complexity and nesting depth of the code, identify the nesting relationship of the code, and provide a scientific basis for component encapsulation decisions.
[0022] Optionally, the first development document is a code library; the first semantic feature information corresponding to the function module is extracted from the first development document, including: abstract syntax tree feature information corresponding to the function module is extracted from the first development document, data dependency feature information based on a data flow graph, and parameter feature information based on interface description; the second development document is a UI design diagram; the second semantic feature information corresponding to the function module is extracted from the second development document, including: element image feature information, element layout feature information, and interaction feature information corresponding to the function module are extracted from the second development document; the second development document is a requirement document; the second semantic feature information corresponding to the function module is extracted from the second development document, including: function feature information, business process feature information, and requirement priority feature information corresponding to the function module are extracted from the second development document.
[0023] Through the above scheme, the extraction of multi-dimensional feature information provides a basis for subsequent determination of encapsulatable nodes, improves the accuracy, comprehensiveness and applicability of component encapsulation, and makes the encapsulated component meet code specifications and UI design requirements while meeting business requirements.
[0024] In a second aspect, the present application provides a component encapsulation device suitable for a plurality of development documents with different function modules, the device comprising:
[0025] An extraction module is configured to extract, for any function module, first semantic feature information corresponding to the function module from a first development document, and second semantic feature information corresponding to the function module from a second development document; the first development document is a development document recorded by a programming language; the second development document is a development document recorded by a natural language;
[0026] A determination module is configured to determine module similarity between first and second function modules based on first and second semantic feature information corresponding to the first and second function modules, respectively; if the module similarity meets a first set condition, determine node similarity between a first node in the first function module and a second node in the second function module; if the node similarity is higher than the module similarity, determine that the first node and the second node are encapsulatable nodes with the same functionality; wherein the first and second function modules are any two of different function modules.
[0027] An encapsulation module is configured to perform component encapsulation on any encapsulatable node.
[0028] In a possible implementation, the determination module is specifically configured to determine structural similarity and semantic similarity between the first and second function modules; and the determination module is specifically configured to determine that the structural similarity meets a first threshold value and the semantic similarity meets a second threshold value.
[0029] In one possible implementation, node similarity is the attention weight between any two nodes; module similarity is the comprehensive similarity determined based on structural similarity and semantic similarity.
[0030] In one possible implementation, structural similarity and semantic similarity are determined as follows: for any functional module, global semantic feature information is determined based on the first semantic feature information and the second semantic feature information; structural similarity is determined based on the cosine similarity of the global semantic feature information of the first functional module and the global semantic feature information of the second functional module; semantic similarity is determined based on the cosine similarity of the first semantic feature information of the first functional module and the first semantic feature information of the second functional module and the weighted value of the cosine similarity of the second semantic feature information of the first functional module and the second semantic feature information of the second functional module.
[0031] In one possible implementation, the determination module is also used to: for any functional module, determine the mixed feature information based on the first semantic feature information and the second semantic feature information; the structural similarity includes the cosine similarity of the mixed information of the first functional module and the mixed information of the second functional module; the semantic similarity includes the weighted value of the cosine similarity of the first semantic feature information of the first functional module and the first semantic feature information of the second functional module and the cosine similarity of the second semantic feature information of the first functional module and the second semantic feature information of the second functional module.
[0032] In one possible implementation, the encapsulation module is specifically used to: encapsulate the same code parts of the encapsulable node in at least two functional modules as component constants; extract component parameters from the different code parts of the encapsulable node in at least two functional modules, thereby obtaining an encapsulated component.
[0033] In one possible implementation, the encapsulation module is specifically used to: extract local semantic feature information from the context information of the encapsulated node; extract global semantic feature information from at least two functional modules; and determine component parameters based on the local semantic feature information and the global semantic feature information.
[0034] In one possible implementation, the determination module is specifically used to: determine the encapsulated nodes based on the abstract syntax tree of the first development document; or determine the encapsulated nodes based on the abstract syntax tree of the first development document and a knowledge graph formed based on multiple development documents.
[0035] In one possible implementation, the determination module is specifically used to: form entities and relationships between entities in a knowledge graph based on multiple development documents; determine entities corresponding to inclusion relationships among the relationships between entities; for each entity among the entities with inclusion relationships: determine the number of child nodes of the entity and the maximum path length from the entity to the deepest child node; determine a first score of the entity based on the number of child nodes and the maximum path length; if the first score is greater than a third threshold, determine that the node corresponding to the entity is an encapsulated node.
[0036] In one possible implementation, the first development document is a code library; the extraction module is specifically used to extract abstract syntax tree feature information corresponding to the functional module, data dependency feature information based on the data flow diagram, and parameter feature information based on the interface description from the first development document; the second development document is a UI design drawing; the extraction module is specifically used to extract element image feature information, element layout feature information, and interaction feature information corresponding to the functional module from the second development document; the second development document is a requirement document; the extraction module is specifically used to extract functional feature information, business process feature information, and requirement priority feature information corresponding to the functional module from the second development document.
[0037] In a third aspect, the present invention further provides a component packaging device, which includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements the method described in the various possible designs of the first aspect.
[0038] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program or instruction stored thereon. When the computer program or instruction is executed by a processor, the method described in the various possible designs of the first aspect is implemented.
[0039] In a fifth aspect, the present invention further provides a computer program product, which, when executed on a computer, enables the computer to execute any of the methods described in the first aspect above.
[0040] These implementations or other implementations of the present application will be more concise and understandable in the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0042] Figure 1A schematic diagram of a flow chart of a component packaging method provided by an embodiment of the present invention;
[0043] Figure 2 A schematic diagram of a component packaging device provided by an embodiment of the present invention;
[0044] Figure 3 A schematic diagram of a front-end code structure provided by an embodiment of the present invention;
[0045] Figure 4 A schematic diagram of a component packaging device provided by an embodiment of the present invention;
[0046] Figure 5 A schematic diagram of another component packaging device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solutions and beneficial effects of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0048] The following is an explanation of some of the terms used in this application. It should be noted that these explanations are for the purpose of facilitating understanding by those skilled in the art and do not limit the scope of protection claimed in this application.
[0049] 1. Large language models (LLMs).
[0050] Large models are AI models that are large in scale and have a large number of parameters. Trained with massive amounts of data, they can understand and generate human language and perform a variety of complex natural language processing tasks. These models, trained using deep learning techniques, possess powerful text understanding, generation, and reasoning capabilities, and can be applied in a variety of scenarios, including text creation, question-answering systems, and language translation.
[0051] 2. Bidirectional encoder representations from transformers (BERT).
[0052] BERT is a pre-trained natural language processing model. Its innovation lies in its bidirectional training approach, which considers both the left and right context of a text to more accurately understand the semantics of language. BERT excels in tasks such as text classification, question-answering, and sentiment analysis. It can extract key functional terms from requirement documents and help understand their core content.
[0053] 3. Abstract syntax tree (AST).
[0054] An AST is a tree-like representation of source code, displaying the grammatical structure of a programming language. During code analysis and compilation, the AST transforms the code into a tree structure that is easy for computers to understand and manipulate, allowing for clear analysis of features such as variable dependencies and component lifecycles.
[0055] 4. Front-end components.
[0056] A front-end component is a reusable, independent unit of code in front-end development that encapsulates a specific user interface element and its functional logic. In modern front-end frameworks such as Vue and React, components are the basic units for building user interfaces.
[0057] Web front-end development refers to creating the user interface portion of a website or web application. In this part of the web development process, developers use the JavaScript programming language to develop interactive features. That is, JavaScript, as a scripting language, allows developers to create dynamic content on web pages, process user input, communicate with servers, etc.
[0058] However, with the continuous iteration of web front-end development project requirements, project scale and code size have exploded, leading to an increase in redundant and duplicated code. Therefore, duplicated code needs to be encapsulated to increase code reuse. Traditional component encapsulation methods rely on manual identification, manually summarizing similar code snippets, and then subjectively evaluating which code is suitable for transformation into reusable components. This approach, when faced with a large amount of similar code and widespread distribution, not only requires developers to invest significant time in encapsulation and global replacement, but also suffers from significant drawbacks such as high maintenance costs, low efficiency, limited coverage, and difficulty achieving cross-platform adaptation. These issues are particularly prominent in large front-end projects, severely impacting development efficiency and code quality, and introducing unnecessary complexity to subsequent project maintenance. Therefore, how to efficiently identify and encapsulate duplicated code in front-end projects, improve code reuse, and reduce maintenance costs, is an urgent issue that needs to be addressed.
[0059] Based on this, the present invention provides a component encapsulation method, which determines the encapsulated nodes through the semantic feature information of the functional modules and the node similarity in the functional modules, and performs component encapsulation on the encapsulated nodes to improve code reuse rate and reduce maintenance costs.
[0060] The present application scheme is described in detail below with reference to the accompanying drawings.
[0061] See also Figure 1, which shows a flow chart of a component packaging method, which is applicable to multiple development documents with different functional modules, and specifically includes the following steps:
[0062] Step 110: For any functional module, extract first semantic feature information corresponding to the functional module from the first development document, and extract second semantic feature information corresponding to the functional module from the second development document.
[0063] The first development document is a development document recorded in a programming language, and the second development document is a document recorded in a natural language.
[0064] Here, a functional module corresponds to a first development document and a second development document. It can be understood that the first development document and the second development document respectively have descriptions of the same functional module from different angles. For example, the first development document is a code module, and the second development document is a UI design drawing or a requirement document. Then, the first semantic feature information is extracted from the first development document, and the second semantic feature information is extracted from the second development document.
[0065] Optionally, the first development document and the second development document may be input into the big model respectively, and the first semantic feature information of the first development document and the second semantic feature information of the second development document may be parsed by the big model.
[0066] Optionally, the first development document is a code library, and first semantic feature information corresponding to the functional module is extracted from the first development document, including: extracting abstract syntax tree feature information corresponding to the functional module from the first development document, data dependency feature information based on the data flow graph, and parameter feature information based on the interface description.
[0067] Optionally, the second development document is a UI design drawing, and second semantic feature information corresponding to the functional module is extracted from the second development document, including: extracting element image feature information, element layout feature information, and interaction feature information corresponding to the functional module from the second development document.
[0068] In a specific example, element image feature information includes component position, hierarchical relationship, color space distribution, etc.; element layout feature information includes spacing constraints, grid system, responsive breakpoints, etc.; interaction feature information includes click area, sliding path, etc.
[0069] Optionally, the second development document is a requirement document, and second semantic feature information corresponding to the functional module is extracted from the second development document, including: extracting functional feature information, business process feature information, and requirement priority feature information corresponding to the functional module from the second development document.
[0070] Generally speaking, the code base is iteratively developed based on the interface elements and layout in the UI design drawings and the functional descriptions in the requirements documents. Here, by performing relevant semantic analysis of the UI design drawings and requirements documents, the large model can better understand the code structure and characteristics in the code base, thereby improving the component encapsulation effect.
[0071] Through the above solution, the extraction of multi-dimensional feature information provides a basis for the subsequent determination of encapsulated nodes, improving the accuracy, comprehensiveness and applicability of component encapsulation, so that the encapsulated components comply with both code specifications and UI design requirements while aligning with business needs.
[0072] Step 120: Determine module similarity between the first functional module and the second functional module based on the first semantic feature information and the second semantic feature information corresponding to the first functional module and the second functional module respectively.
[0073] Here, the first functional module and the second functional module are any two different functional modules. In step 110, the first semantic feature information and the second semantic feature information of the first functional module and the first semantic feature information and the second semantic feature information of the second functional module are obtained. Based on the above feature information, the module similarity between the first functional module and the second functional module is determined, that is, whether the first functional module and the second functional module have duplicate code.
[0074] Optionally, before step 120, the first semantic feature information of the first functional module and the first fusion vector of the second semantic feature information, and the second fusion vector of the first semantic feature information of the second functional module and the second semantic feature information can be calculated, and the module similarity between the first functional module and the second functional module can be determined based on the first fusion vector and the second fusion vector.
[0075] Through the above solution, different semantic feature information can be fused first, that is, feature information of different modalities can be mapped into the same space to enhance the semantic understanding of the code by the large model.
[0076] Optionally, module similarity may include structural similarity and semantic similarity. Structural similarity refers to the degree of similarity between functional modules in grammatical structure, component layout, and code organization, while semantic similarity refers to the degree of similarity between functional modules in functional implementation, business logic, and user interaction behavior.
[0077] Furthermore, optionally, module similarity can be based on a comprehensive similarity determined based on structural similarity and semantic similarity. Here, module similarity can be evaluated using both independent and combined methods. The independent method involves determining whether structural similarity meets a first preset threshold and semantic similarity meets a second preset threshold. The combined method involves setting weights for structural similarity and semantic similarity, respectively, and determining module similarity by calculating the weighted values of the structural and semantic similarities.
[0078] In practical applications, the weights of structural and semantic similarity can be adjusted based on specific project requirements to optimize component encapsulation. For example, for projects that prioritize UI consistency, the weight of structural similarity can be appropriately increased; while for enterprise-level applications with complex functions, the weight of semantic similarity can be increased to ensure accurate encapsulation of functional logic.
[0079] The above approach, which separately assesses structural and semantic similarity, can provide more refined componentization recommendations based on similarity across different dimensions, offering developers more packaging options. A combined approach allows for a more comprehensive and rigorous assessment of functional module similarity, avoiding inappropriate packaging due to similarity across a single dimension and improving the quality and practicality of component packaging.
[0080] Optionally, structural similarity is determined in the following manner: for any functional module, global semantic feature information is determined based on the first semantic feature information and the second semantic feature information; structural similarity is determined based on the cosine similarity of the global semantic feature information of the first functional module and the global semantic feature information of the second functional module.
[0081] Here, for any functional module, the semantic feature information from different sources (first semantic feature information and second semantic feature information) is fused to generate a global semantic feature information. These semantic features from different sources may come from code libraries, UI design drawings or requirement documents, etc. By calculating the cosine similarity between the global semantic feature information of two different functional modules, the cosine similarity value is used as an indicator to measure the structural similarity of the two functional modules.
[0082] Furthermore, optionally, the calculation formula for structural similarity is as follows:
[0083]
[0084] Among them, C i , C j Code snippets for different functional modules, is the fusion vector of different functional modules (global semantic feature information).
[0085] Optionally, the semantic similarity is determined as follows: the semantic similarity is determined based on the cosine similarity of the first semantic feature information of the first functional module and the first semantic feature information of the second functional module, and the weighted value of the cosine similarity of the second semantic feature information of the first functional module and the second semantic feature information of the second functional module.
[0086] Unlike structural similarity, which first calculates the fused features and then the similarity, semantic similarity first calculates the similarity of features from different dimensions separately and then performs a weighted fusion. This dimensional calculation followed by a weighted fusion approach allows for flexible adjustment of the influence of different feature dimensions in semantic similarity assessment, adapting to the component packaging requirements of different project types and improving the accuracy and adaptability of semantic similarity assessment.
[0087] Furthermore, optionally, the calculation formula for semantic similarity is as follows:
[0088]
[0089] in, is the code AST node state vector of different functional modules, D i , D j T is the UI design diagram variable corresponding to different functional modules. i , T j are the requirement document variables corresponding to different functional modules, and α, β, and γ are the corresponding weight distributions.
[0090] Through the above scheme, using cosine similarity to calculate similarity can accurately reflect the similarity between different functional modules. The two-dimensional similarity evaluation mechanism provides a more accurate judgment standard for component encapsulation, which takes into account both the formal similarity of the code structure and the semantic consistency of the code function implementation, thereby more accurately identifying code modules suitable for encapsulation as components.
[0091] Step 130: If the module similarity satisfies the first set condition, determining the node similarity between the first node in the first functional module and the second node in the second functional module.
[0092] Here, if it is determined that the first functional module and the second functional module have duplicated code, it is further determined whether the duplicated code nodes in the first functional module and the second functional module can be packaged into a component.
[0093] Optionally, step 130 includes determining node similarity between a first node in the first functional module and a second node in the second functional module if the structural similarity satisfies a first threshold and the semantic similarity satisfies a second threshold. That is, if the first functional module and the second functional module have both high structural similarity and high semantic similarity, it indicates that the two functional modules likely implement the same or similar functions. Subsequently, determining node similarity further determines whether the two functional modules can be encapsulated into a reusable component.
[0094] Through the above scheme, the structural similarity and semantic similarity between functional modules are taken into consideration at the same time, and the corresponding threshold is set as the judgment standard, which significantly improves the accuracy and practicality of component encapsulation. Among them, structural similarity focuses on the formal characteristics of the code, such as document structure, tag nesting and style rules, while semantic similarity focuses on the functional implementation, business logic and interactive behavior of the code. This dual-dimensional similarity evaluation method avoids the incorrect encapsulation caused by only being based on a single dimension. In addition, the present application can also identify code fragments that are "different in structure but similar in function" or "similar in structure but different in function", providing developers with more encapsulation options, making component encapsulation more flexible and adaptable to the needs of different scenarios.
[0095] Alternatively, node similarity can be the attention weight between any two nodes. Here, the attention weight can be obtained from the AST. The attention weight can reflect the importance of the node in the code structure. The higher the weight, the more likely the node is the core function node of the code.
[0096] By quantifying node similarity as attention weights, this approach captures the strength of associations between code snippets. This allows similarity calculations to move beyond simple syntactic comparisons and instead reflect the actual importance and relevance of nodes in functional implementation. This attention weighting mechanism automatically strengthens connections between key nodes while simultaneously deemphasizing the influence of less important nodes, thereby improving the accuracy of identifying functional boundaries during component encapsulation.
[0097] Step 140: If the node similarity is higher than the module similarity, determine that the first node and the second node are encapsulable nodes with the same functionality.
[0098] Here, when the similarity between two nodes is higher than the overall similarity of the modules they belong to, it means that the two nodes have high functional similarity, and therefore can be identified as nodes with the same functional characteristics and are suitable for being encapsulated into a reusable component.
[0099] Step 150: Perform component encapsulation on any encapsulable node.
[0100] Here, the nodes determined as encapsulable nodes in step 140 are encapsulated as components. In subsequent code development, the encapsulated components can be called multiple times in different parts of the project, avoiding code redundancy. In subsequent code maintenance, when a certain function needs to be maintained, only the corresponding encapsulated components need to be maintained, without having to search for and modify all similar codes in the entire project.
[0101] Optionally, step 150 includes: encapsulating the same code portion of the encapsulable node in at least two functional modules as component constants, and extracting component parameters from the different code portions of the encapsulable node in at least two functional modules, thereby obtaining an encapsulated component.
[0102] In other words, the same code parts will be encapsulated as component constants of the component - that is, the unchanging and fixed parts of the component, and the different parts will be extracted as the "variable parameters" of the component. By combining the constant part (fixed structure) and the parameter part (variable configuration), a reusable component is finally generated.
[0103] Through the above solution, the same code parts in multiple functional modules are encapsulated as component constants, which reduces the amount of duplicate code and improves the overall quality and maintainability of the project code. When a common function needs to be modified, only one encapsulated component needs to be modified to achieve a unified update of all places where the component is used, which significantly improves the efficiency of code maintenance; the mechanism of extracting component parameters from different code parts ensures the flexibility and adaptability of the components. By parameterizing the changing parts (such as data sources, configuration options, style themes, etc.), the same component can exhibit different behaviors and appearances through different parameters in different scenarios. This parameterized design not only increases the scope of component reuse, but also retains the unique requirements of each functional module, realizing the "one-time development, multiple use" model.
[0104] The above solution automatically extracts component functional boundaries and interaction logic. Component functional boundary extraction refers to the process of breaking down a complex functional module into multiple independent components with clear functional boundaries during front-end development. This process helps improve code maintainability, reusability, and scalability. Interaction logic extraction refers to the process of extracting the internal functionality and implementation logic of a single component.
[0105] Further, optionally, component parameters are extracted from different code parts of the encapsulated node in at least two functional modules, including: extracting local semantic feature information from the upper and lower information of the encapsulated node; extracting global semantic feature information from at least two functional modules; and determining component parameters based on the local semantic feature information and the global semantic feature information.
[0106] Furthermore, optionally, the formula for encapsulating the same code portion of the encapsulable node in at least two functional modules as a component constant is as follows:
[0107]
[0108] KeyNodes = {v|α v >θ}
[0109] Among them, α v is the average attention weight of v in AST node, θ is the threshold (α v Nodes larger than θ are encapsulated component nodes, H is the number of nodes in AST, and KeyNodes are encapsulated nodes.
[0110] Through the above scheme, local semantic feature information can capture the unique needs and usage scenarios of encapsulated nodes based on the contextual environment around the encapsulated nodes; while global semantic feature information identifies the parameter configurations required in common in different scenarios from the overall perspective of multiple functional modules. This local + global dual semantic analysis method enables the extracted component parameters to meet the needs of specific scenarios and have universality across scenarios. Secondly, the application significantly improves the intelligence level of component encapsulation. Through semantic feature analysis rather than simple code structure analysis, by understanding the actual function and business meaning of the code, it can more accurately identify which variables should be used as component parameters. This semantic-based parameter extraction method avoids the errors that may be caused by traditional methods that rely solely on surface features such as variable names or positions for judgment, making the encapsulated components more in line with actual business needs.
[0111] Furthermore, optionally, the formula for extracting component parameters from different code parts of the encapsulable node in at least two functional modules is as follows:
[0112] p{prop k |C,D,T}=softmax(W p [h fused ;h ctx ])
[0113] Among them, h ctx is the current context encoding vector, h fused is the global semantic feature information (also called fusion vector), prop k is the candidate component parameter, C, D, T are different development documents (such as code snippet C, requirement document D, UI design diagram T), W p is a learnable weight matrix and softmax is a normalization function.
[0114] Optionally, after the above step 140, the following step may also be included: determining encapsulated nodes based on a knowledge graph formed by multiple development documents.
[0115] Through the above scheme, the use of abstract syntax can capture the grammatical structure and logical relationship of the code. Compared with simple text matching or keyword search, the present invention can more accurately identify code fragments with the same functional characteristics, reduce the possibility of misjudgment, and improve the accuracy of component encapsulation. Secondly, the introduction of knowledge graph as an optional scheme further enhances the ability to identify complex nested relationships. Due to the limitations of relying solely on abstract syntax tree analysis, the abstract syntax tree method cannot cover complex component nested relationships, while the knowledge graph can capture implicit nested relationships and cross-modal scenarios that are difficult for the abstract syntax tree to identify; therefore, this dual recognition mechanism can cope with code structures of various complexities, that is, for simple nested relationships, only the abstract syntax tree can be used for rapid recognition; and for complex nested relationships, a more comprehensive analysis can be performed in combination with the knowledge graph, thus expanding the scope of application of component encapsulation and improving the coverage and accuracy of component encapsulation.
[0116] Furthermore, optionally, the present invention may further include the following steps ae:
[0117] Step a: forming entities and relationships between entities in a knowledge graph based on multiple development documents;
[0118] Optionally, define the knowledge graph G = (V, E) as follows:
[0119] V={v i}
[0120] E={v i ,r k ,v j}
[0121] Among them, V represents entity nodes such as components, UI design styles, and requirement documents, and E represents entity elements v i and v j The concrete implementation relationship between them, such as the requirements of code modules, etc. k Is the relationship type.
[0122] Step b, determining the entities corresponding to the inclusion relationship among the relationships between entities;
[0123] Step c, for each entity in the entities having the containment relationship: determining the number of child nodes of the entity and the maximum path length from the entity to the deepest child node;
[0124] Step d, for each entity in the entities having the containment relationship: determining a first score of the entity according to the number of child nodes and the maximum path length;
[0125] Step e, for each of the entities with the containing relationship: if the first score is greater than the third threshold, determining that the node corresponding to the entity is an encapsulatable node.
[0126] Through the above scheme, by extracting entities and their relationships from multiple development documents to form a knowledge graph, the complexity and nesting depth of the code can be evaluated, and the nesting relationship of the code can be identified, providing a scientific basis for component encapsulation decision-making.
[0127] The present application converts code modules with high repeatability into reusable components, automatically encapsulates code nodes with the same functionality, significantly reduces code redundancy, and improves code reuse rate and maintenance efficiency. Compared with the traditional manual identification and encapsulation method, the automatic method based on multi-source semantic analysis of the present application not only improves the efficiency of component encapsulation, but also ensures the encapsulation quality through the similarity judgment standard, providing strong support for the componentization reform of large front-end projects.
[0128] The present application extracts semantic feature information from different types of development documents, analyzes the development documents recorded by programming languages (such as code libraries) and the development documents recorded by natural languages (such as requirement documents or UI design documents) at the same time, establishes a semantic bridge between code implementation and business requirements, and thus more comprehensively understands the purpose and implementation method of the functional module, significantly improving the accuracy and applicability of component encapsulation.
[0129] The above describes a component encapsulation method, and the following will illustrate the above method with a specific example.
[0130] Please refer to Figure 2 , which shows a schematic diagram of a component encapsulation device provided by the present application, which includes a multi-modal semantic analysis module 210, a component semantic analysis and automatic generation module 220, a knowledge graph-based component relationship dynamic optimization module 230, and a large model 240.
[0131] The multi-modal semantic analysis module 210 is used to scan the front-end code library, UI diagram and requirement document to realize intelligent analysis of the project, the component semantic analysis and automatic generation module 220 is used to identify the nodes that can be encapsulated into components, and encapsulate the encapsulatable nodes into components (i.e. the above steps 120-150); the knowledge graph-based component relationship dynamic optimization module 230 is used to determine the encapsulatable nodes according to the knowledge graph; finally, the large model 240 determines the encapsulatable components from the front-end business code based on the above multi-modal semantic analysis module 210, component semantic analysis and automatic generation module 220, knowledge graph-based component relationship dynamic optimization module 230, and outputs the component encapsulation result.
[0132] Based on the above component packaging device, the present invention provides a component packaging method, which can be divided into four steps. It can be understood that the following four steps can refer to the above steps 110 to 150. Specifically, step one can refer to the above step 110, step two can refer to the above steps 120 to 140, and step three can refer to the above steps a to e.
[0133] Step 1: For any functional module, extract semantic feature information from the code base (first development document), UI design drawing (second development document), and requirement document (second development document).
[0134] Here, different development documents can be parsed through the large model to parse different development documents into corresponding semantic data structures.
[0135] The following introduces the extraction of semantic feature information for different development documents.
[0136] (1) Visual feature vector a1, structured metadata a2, and interaction event matrix a3 are extracted from the UI design drawing, and the visual feature vector a1, structured metadata a2, and interaction event matrix a3 are fused to obtain a semantic variable X based on the UI design drawing.
[0137] Among them, the visual feature vector a1 is used to capture element image feature information such as component position, hierarchical relationship, and color space distribution; the structured metadata a2 is used to capture element layout feature information such as spacing constraints, grid system, and responsive breakpoints; and the interaction event matrix a3 is used to capture interaction feature information such as click areas and sliding paths.
[0138] Optionally, the visual feature vector a1 can be extracted using a CNN (such as ResNet50) to generate a vector of a set dimension (such as 512 dimensions).
[0139] In a specific example, the visual feature vector a1 is:
[0140] [0.1234, -0.5678, 0.9012, ... 0.4321]
[0141] Optionally, the structured metadata a2 can be recognized by OpenCV.
[0142] In a specific example, the structured metadata a2 is:
[0143]
[0144] Optionally, the interaction event matrix a3 can be obtained through prototyping tool metadata (such as Figma / Sketch plug-ins, etc.).
[0145] In a specific example, the interaction event matrix a3 is:
[0146]
[0147]
[0148] Finally, the visual feature vector a1, structured metadata a2, and interaction event matrix a3 are fused to obtain the semantic variable X based on the UI design diagram:
[0149]
[0150]
[0151] (2) Extract the AST structure variable b1, data flow graph variable b2, and component interface variable b3 from the code base, and fuse the AST structure variable b1, data flow graph variable b2, and component interface variable b3 to obtain the semantic variable Y based on the front-end code base.
[0152] Optionally, the AST structure variable b1 can use Babel to generate an abstract syntax tree, extract variable dependency paths, and generate component lifecycle hooks (lifecycle hooks refer to functions that are automatically executed at specific points in the process from component creation to destruction in the framework).
[0153] Optionally, the data flow graph variable b2 can be generated by analyzing the variable propagation path through CodeQL and constructing a cross-file data dependency graph (DFG).
[0154] Optionally, the component interface variable b3 can be generated by extracting the component interface specification based on JSDoc / TypeScript type annotations.
[0155] In a specific example, the semantic variable Y of the code base is:
[0156] {
[0157] "ast_hash":"a3f5c2",
[0158] "dependencies":[
[0159] {"source":"Button","target":"Form","type":"import"},{"source":"Table","target":"fetchData","type":"call"}
[0160] ],
[0161] "props_schema":[
[0162] {"name":"items","type":"Array","required":true},
[0163] {"name":"theme","default":"light"} ]
[0165] }
[0166] (3) Extract functional feature information c1, business process feature information c2, and demand priority feature information c3 from the requirement document, and fuse the functional feature information c1, business process feature information c2, and demand priority feature information c3 to obtain a semantic variable Z based on the requirement document.
[0167] Optionally, the functional feature information c1 can be obtained by extracting functional keywords (such as paging and topic switching) using a BERT model (such as CodeBERT-NL encoder).
[0168] Optionally, the business process characteristic information c2 can be obtained by constructing a demand logic diagram through dependency syntax analysis and identifying conditional branches.
[0169] Optionally, the demand priority feature information c3 can be obtained by marking the demand urgency (such as P0-P3 level) based on the TF-IDF weighted algorithm.
[0170] In a specific example, the semantic variable Z based on the requirements document is:
[0171]
[0172] Step 2: Determine module similarity and node similarity based on semantic feature information, and identify code modules with repeated functions as encapsulated nodes.
[0173] This step can optimize the algorithm training of the CodeBERT model to improve the semantic understanding of the front-end code by the large model, and then dynamically analyze the code modules in the code base that can be encapsulated into components.
[0174] This step can be specifically divided into the following three steps: (1) fusing multimodal semantic feature information; (2) calculating the module similarity of different functional modules; and (3) determining the nodes and parameters that can be encapsulated into components. These three steps are explained below.
[0175] (1) Fuse multimodal semantic feature information.
[0176] The three semantic feature information X, Y, and Z obtained in step 1 are multimodally fused. The fusion formula is as follows:
[0177] h fused =Layernorm(W c h c +W d D+W t T)
[0178] Among them, Wc, Wd, Wt are learnable weight matrices, h c is the AST node state vector, D is the UI element layout vector (a set of semantic variables X based on the UI design diagram, D∈R d , R d represents a d-dimensional real vector space), T is the natural language semantic embedding vector (a set of semantic variables Z based on the requirements document, T∈R t ), h fused It is the fusion vector of all functional modules.
[0179] h c It is calculated through the following steps: obtain the AST node sequence (i.e., the combination of semantic variables Y based on the front-end code base) c = {c1, c2, ..., cn}, input the AST node sequence into the pre-trained CodeBERT model to generate the AST node state vector hc = {hc1, hc2, ..., hcn}.
[0180] The above formula can map different modal features (semantic variable X based on UI design drawings, semantic variable Y based on front-end code base, and semantic variable Z based on requirement documents) into the same space, achieving trimodal semantic alignment of code, design, and requirements, and maximizing the semantic similarity among code, design, and requirements.
[0181] (2) Calculate the module similarity of different functional modules. Here, module similarity can include structural similarity and functional similarity.
[0182] Here, we first calculate the fusion vectors of different functional modules. The calculation formula is as follows:
[0183]
[0184] Here, i represents different functional modules, and Pool represents the aggregation of code AST node sequences.
[0185] Optionally, the formula for calculating structural similarity is as follows:
[0186]
[0187] Among them, C i, C j Code snippets for different functional modules, is the fusion vector of different functional modules.
[0188] Through the above formula, we can get the code fragment C i and C j The structural similarity of Sim1(C i , C j ) is a real number in the interval [0, 1]. The higher the value, the higher the code similarity. Here, we assume that the code structure similarity Sim1(C i , C j )=0.92.
[0189] Optionally, the semantic similarity is calculated as follows:
[0190]
[0191] in, is the code AST node state vector of different functional modules, D i , D j T is the UI design diagram variable corresponding to different functional modules. i , T j are the requirement document variables corresponding to different functional modules, and α, β, and γ are the corresponding weight distributions.
[0192] Through the above formula, we can get code fragment C i and C j Semantic similarity, Sim2(C i , C j ) is a real number in the interval [0, 1]. The higher the value, the higher the code similarity. Here, we assume that the code semantic similarity Sim2(C i , C j )=0.72.
[0193] (3) Determine the nodes and parameters that can be encapsulated into components.
[0194] Here, we first determine whether the code fragments in the functional modules with similar functions can be encapsulated into components, and then convert the Sim1(C i , C j ) and Sim2(C i , C j ) as the judgment standard, and Sim1(C i , C j ) and Sim2(C i , C j) is used as the node selection criterion to identify the functional boundaries of components (component functional boundary extraction refers to the process of decomposing a complex functional module into multiple independent components with clear functional boundaries during front-end development) to improve the accuracy of component encapsulation.
[0195] Here we assume that Sim1(C i , C j )>0.6 and Sim2(C i , C j )>0.6, then the two functional modules are judged to have similar functionality.
[0196] The formula for determining encapsulated component nodes is as follows:
[0197]
[0198] KeyNodes = {v|α v >θ}
[0199] Among them, α v is the average attention weight of v in AST node, θ is the threshold (here Sim1(C i , C j ) and Sim2(C i , C j ) average value), and α v Nodes larger than θ are encapsulated component nodes, H is the number of nodes in AST, and KeyNodes are encapsulated nodes.
[0200] The above formula can extract code nodes with higher attention weights and use these nodes as node elements of components to determine the functional boundaries of components.
[0201] The formula for determining component parameters is as follows:
[0202] p{prop k |C,D,T}=softmax(W p [h fused ;h ctx ])
[0203] Among them, h ctx Encoding vectors for the current context can focus on the local semantics of the current generation scenario (such as the real-time context of the code / requirements around a component when encapsulating it), prop k For candidate interface parameters (such as items, page, etc., which may be exposed by components in the corresponding code), p{prop k |C,D,T} is the conditional probability, which indicates that under the project context (code snippet C, requirement document D, UI design diagram T and other information), the attribute propk As the probability of the interface parameter, W p is a learnable weight matrix, a parameter optimized during model training, used to fused ;h ctx ] performs linear transformation to adapt the probability calculation, and softmax is the normalization function.
[0204] The above formula is used to generate parameterized interfaces (such as component props definitions) by calculating "a candidate property prop k As the probability of interface parameters, the big model can filter out reasonable interface parameters based on the project's multimodal information (code base, requirement documents, UI design drawings, etc.) to achieve parametric design during component packaging. In other words, the above formula can prioritize the extraction of frequently referenced variables as the interface of the packaged component.
[0205] Through the above scheme, the nodes and parameters of the encapsulated components can be determined.
[0206] In a specific example, the e-commerce project code base (including different modules for product lists, etc.), UI design drawings (including list layouts, etc.), and requirements documents (including requirements such as support for paging and sorting) are input into the big model 240. The code of the general product list component output by the big model 240 is as follows:
[0207]
[0208]
[0209] In the above code, the variables in props represent component parameters, that is, parameterized interfaces, and the functional boundaries represent the main functional tags that constitute the component, including the product list tag div and the data pagination tag Pagination.
[0210] Through component semantic analysis and automatic generation algorithms, this application can automatically perform semantic fusion on multimodal data (code libraries, UI design drawings, requirements documents, etc.), as well as component similarity analysis and encapsulation, enabling automated component encapsulation of code snippets with relatively simple nesting relationships. However, for code snippets with more complex nesting relationships, component encapsulation is less effective.
[0211] Specifically, although the component nesting recognition method based on AST syntax tree can recognize simple code nesting relationships and regular component tags, it has limitations for complex component nesting relationships. For example, when there are cross-modal scenarios and implicit nesting relationships (implicit nesting usually occurs at the logical level rather than the syntax tag level). For example, see Figure 3A schematic diagram shows a front-end code structure when the Form component imports the Input component through the import statement but is dynamically rendered in the template through the render function. For another example, in the high-order component scenario (such as const EnhancedForm = withAuth(Form)), although the AST can parse the basic grammatical structure, it is difficult to identify the impact of the withAuth function on the nested logic of Form (because the relevant nested logic is within the withAuth function and requires semantic analysis), making it difficult to understand the complete nested logical relationship.
[0212] Based on this, in order to further optimize the quality of large model encapsulation components, the present invention introduces a component relationship dynamic optimization module 230 based on the knowledge graph. For details, see step three below.
[0213] Step 3: Based on the knowledge graph formed by multiple development documents, determine the encapsulated nodes.
[0214] Here, we first construct a dynamic knowledge graph based on code features, then use the knowledge graph features to analyze the more complex component nesting relationships in the code through algorithms, output the component nesting relationships as structured data, and finally automatically complete the front-end code encapsulation based on the analysis results. Specifically, step three can include the following steps (1) to (7).
[0215] Step (1) defines the knowledge graph G = (V, E), and defines the entities in the knowledge graph and the relationships between entities as follows:
[0216] V={v i}
[0217] E={v i ,r k ,v j}
[0218] Among them, V represents entity nodes such as components, UI design styles, and requirement documents, and E represents entity elements v i and v j The concrete implementation relationship between them, such as the requirements of code modules, etc. k Is the relationship type.
[0219] Step (2), get the triple v i ,r k ,v j The fusion variable set T1 is as follows:
[0220]
[0221] Among them, F triplet is a triplet (vi ,r k ,v j ) function.
[0222] Step (3) determines the entities corresponding to the inclusion relationship among the relationships between entities. The formula is as follows:
[0223] T nest ={((v i ,contains,v j )|(v j ,contains,v i ))∈T1}
[0224] According to the above formula, if the element v i Contains element v j , and this containment relationship already exists in the original set T1, then this containment relationship constitutes a nested relationship and will be added to the nested edge set T nest middle.
[0225] Step (4), construct the nested subgraph G nest , the formula is as follows:
[0226] G nest =(V nest ,T nest )
[0227] Among them, the nested subgraph G nest It is a subgraph extracted from the original graph that contains only nested relationships. It consists of a node set V nest and edge set T nest composition.
[0228] Through the above solution, clear nested relationships can be extracted from complex UI structure diagrams or other hierarchical structures to help large models understand and analyze the hierarchical structure between components.
[0229] Please continue reading Figure 3 , assuming the front-end code structure is as follows Figure 3 As shown, after the above transformation, the structural subgraph G can be generated nest for:
[0230] (App, contains, Header)
[0231] (App, contains, Main)
[0232] (Main, contains, Form)
[0233] (Form, contains, Input)
[0234] (Form, contains, Button)
[0235] Step (5), for each entity in the entities with the containment relationship: determine the number of child nodes of the entity and the maximum path length from the entity to the deepest child node.
[0236] The formula for determining the child node tree is as follows:
[0237] FanOut(n)=|{n j |(n i ,contains,n j )∈G nest}|
[0238] The formula for determining the maximum path length from an entity to its deepest child node is as follows:
[0239]
[0240] Step (6) determines the first score of the entity based on the number of child nodes and the maximum path length. The formula is as follows:
[0241] Score wrap (C)=α·Depth(C)+β·FanOut(C)
[0242] The weights of α and β can be set according to experience. In a specific example, α is set to 0.5 and β is set to 0.5.
[0243] Step (7): If the first score is greater than a third threshold, the node corresponding to the entity is determined to be an encapsulated node.
[0244] The formula is as follows:
[0245] Score wrap (C)>θ
[0246] Through the above scheme, when the first score Score wrap (C) When the value is greater than the set component encapsulation threshold, the node corresponding to the entity is determined to be an encapsulated node. The threshold setting can be flexibly configured according to the business scenario. In a specific example, θ is set to 2.
[0247] In a specific example, the generated data structure is as follows:
[0248]
[0249] The above data structure is sent to the large model 240, and the large model 240 encapsulates the encapsulated nodes into components and outputs the encapsulation results.
[0250] Step 4: Based on the encapsulated nodes obtained in step 2 and step 3, encapsulate the front-end business code into components and output the component encapsulation results.
[0251] After that, whether it is code development or code maintenance, you can only modify the components without modifying all the code, which improves code reuse and daily development efficiency.
[0252] Based on the same concept, an embodiment of the present application also provides a component packaging device, which can execute the component packaging method introduced in the above content.
[0253] See also Figure 4 , a structural diagram of a component packaging device provided by an embodiment of the present invention is given, such as Figure 4 As shown, the component packaging device is applicable to multiple development documents with different functional modules, and the component packaging device includes:
[0254] Extraction module 401 is configured to extract, for any functional module, first semantic feature information corresponding to the functional module from a first development document, and second semantic feature information corresponding to the functional module from a second development document; the first development document is a development document recorded in a programming language; the second development document is a development document recorded in a natural language;
[0255] Determination module 402 is configured to determine module similarity between the first functional module and the second functional module based on first semantic feature information and second semantic feature information corresponding to the first functional module and the second functional module, respectively; if the module similarity satisfies a first set condition, determine node similarity between a first node in the first functional module and a second node in the second functional module; if the node similarity is higher than the module similarity, determine that the first node and the second node are encapsulated nodes having the same functionality; wherein the first functional module and the second functional module are any two different functional modules;
[0256] The encapsulation module 403 is used to perform component encapsulation on any encapsulable node.
[0257] In a possible implementation, the determination module 401 is specifically configured to determine the structural similarity and semantic similarity between the first functional module and the second functional module; the determination module 401 is specifically configured to determine whether the structural similarity satisfies a first threshold and the semantic similarity satisfies a second threshold.
[0258] In one possible implementation, node similarity is the attention weight between any two nodes; module similarity is the comprehensive similarity determined based on structural similarity and semantic similarity.
[0259] In one possible implementation, structural similarity and semantic similarity are determined as follows: for any functional module, global semantic feature information is determined based on the first semantic feature information and the second semantic feature information; structural similarity is determined based on the cosine similarity of the global semantic feature information of the first functional module and the global semantic feature information of the second functional module; semantic similarity is determined based on the cosine similarity of the first semantic feature information of the first functional module and the first semantic feature information of the second functional module and the weighted value of the cosine similarity of the second semantic feature information of the first functional module and the second semantic feature information of the second functional module.
[0260] In one possible implementation, the determination module 402 is also used to: for any functional module, determine the mixed feature information based on the first semantic feature information and the second semantic feature information; the structural similarity includes the cosine similarity of the mixed information of the first functional module and the mixed information of the second functional module; the semantic similarity includes the weighted value of the cosine similarity of the first semantic feature information of the first functional module and the first semantic feature information of the second functional module and the cosine similarity of the second semantic feature information of the first functional module and the second semantic feature information of the second functional module.
[0261] In one possible implementation, the encapsulation module 403 is specifically used to: encapsulate the same code parts of the encapsulable node in at least two functional modules as component constants; extract component parameters from the different code parts of the encapsulable node in at least two functional modules, thereby obtaining an encapsulated component.
[0262] In one possible implementation, the encapsulation module 403 is specifically configured to: extract local semantic feature information from context information of the encapsulable node; extract global semantic feature information from at least two functional modules; and determine component parameters based on the local semantic feature information and the global semantic feature information.
[0263] In one possible implementation, the determination module 402 is specifically used to: determine the encapsulated nodes based on the abstract syntax tree of the first development document; or determine the encapsulated nodes based on the abstract syntax tree of the first development document and a knowledge graph formed based on multiple development documents.
[0264] In one possible implementation, the determination module 402 is specifically used to: form entities and relationships between entities in a knowledge graph based on multiple development documents; determine entities corresponding to inclusion relationships among the relationships between entities; for each entity among the entities with inclusion relationships: determine the number of child nodes of the entity and the maximum path length from the entity to the deepest child node; determine a first score of the entity based on the number of child nodes and the maximum path length; if the first score is greater than a third threshold, determine that the node corresponding to the entity is an encapsulated node.
[0265] In one possible implementation, the first development document is a code library; the extraction module is specifically used to extract abstract syntax tree feature information corresponding to the functional module, data dependency feature information based on the data flow diagram, and parameter feature information based on the interface description from the first development document; the second development document is a UI design drawing; the extraction module is specifically used to extract element image feature information, element layout feature information, and interaction feature information corresponding to the functional module from the second development document; the second development document is a requirement document; the extraction module is specifically used to extract functional feature information, business process feature information, and requirement priority feature information corresponding to the functional module from the second development document.
[0266] See also Figure 5 , showing a structural diagram of another component packaging device provided by an embodiment of the present application, such as Figure 5 As shown, the component packaging device includes: a memory 501 and a processor 502, and the processor 502 is coupled to the memory 501. The memory 501 is used to store program instructions, and the processor 502 is used to call the program instructions stored in the memory 501 and execute the above component packaging method according to the obtained program.
[0267] Optionally, the component packaging apparatus may further include an interface circuit 503, which may be a transceiver or an input / output interface. The input / output interface is used to input and / or output information, where output can be understood as sending and input can be understood as receiving. The processor 502 may communicate with other devices in the test apparatus or other devices outside the component packaging apparatus via the interface circuit 503 to obtain information required for executing the above component packaging method.
[0268] When the data processing device 500 is used to implement Figure 1 In the method shown, the processor 502 is used to implement the functions of the above-mentioned extraction module 401, determination module 402, and packaging module 403.
[0269] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0270] The memory in the embodiments of the present application can be a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, a register, a hard disk, a mobile hard disk, a compact disc read-only memory (CD-ROM), or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. The storage medium can also be an integral part of the processor.
[0271] Based on the same technical concept, an embodiment of the present invention also provides a computer-readable storage medium, on which a computer program or instruction is stored. When the computer program or instruction is executed by a processor, the computer executes the above-mentioned component packaging method.
[0272] Based on the same technical concept, an embodiment of the present invention further provides a computer-readable program product, which, when executed, enables a computer to execute the above-mentioned component packaging method.
[0273] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0274] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0275] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0276] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0277] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A component packaging method, characterized in that: Applicable to multiple development documents with different functional modules, the method includes: For any functional module, extracting first semantic feature information corresponding to the functional module from a first development document, and extracting second semantic feature information corresponding to the functional module from a second development document; the first development document is a development document recorded in a programming language; the second development document is a development document recorded in a natural language; Determine, based on first semantic feature information and second semantic feature information corresponding to the first functional module and the second functional module, module similarity between the first functional module and the second functional module; if the module similarity satisfies a first set condition, determine node similarity between a first node in the first functional module and a second node in the second functional module; if the node similarity is higher than the module similarity, determine that the first node and the second node are encapsulated nodes having the same functionality; the first functional module and the second functional module are any two of the different functional modules; Perform component encapsulation on any encapsulable node.
2. The method according to claim 1, wherein The determining the module similarity between the first functional module and the second functional module includes: determining structural similarity and semantic similarity between the first functional module and the second functional module; The module similarity meeting the first set condition includes: The structural similarity satisfies a first threshold and the semantic similarity satisfies a second threshold.
3. The method according to claim 2, wherein The node similarity is the attention weight between any two nodes; the module similarity is the comprehensive similarity determined based on the structural similarity and the semantic similarity.
4. The method according to claim 2, wherein The structural similarity and the semantic similarity are determined as follows: For any functional module, determining global semantic feature information based on the first semantic feature information and the second semantic feature information; determining the structural similarity based on the cosine similarity between the global semantic feature information of the first functional module and the global semantic feature information of the second functional module; The semantic similarity is determined based on the cosine similarity of the first semantic feature information of the first functional module and the first semantic feature information of the second functional module and the weighted value of the cosine similarity of the second semantic feature information of the first functional module and the second semantic feature information of the second functional module.
5. The method according to claim 1, wherein The component encapsulation of any encapsulable node includes: The same code parts of the encapsulable node in at least two functional modules are encapsulated as component constants; and component parameters are extracted from the different code parts of the encapsulable node in at least two functional modules, thereby obtaining an encapsulated component.
6. The method according to claim 5, wherein Extracting component parameters from different code parts of the encapsulable node in at least two functional modules includes: Extracting local semantic feature information from the context information of the encapsulated node; Extracting global semantic feature information from the at least two functional modules; Component parameters are determined based on the local semantic feature information and the global semantic feature information.
7. The method according to any one of claims 1 to 6, wherein The determining that the first node and the second node are encapsulated nodes having the same functionality includes: Determine encapsulated nodes based on the abstract syntax tree of the first development document; or Based on the abstract syntax tree of the first development document and the knowledge graph formed based on the multiple development documents, encapsulated nodes are determined.
8. The method according to claim 7, wherein The step of determining encapsulable nodes based on the knowledge graph formed by the plurality of development documents includes: Forming entities and relationships between entities in the knowledge graph based on the multiple development documents; Determining, among the relationships between the entities, entities corresponding to a containment relationship; For each entity in the entities with the containment relationship: determining the number of child nodes of the entity and the maximum path length from the entity to the deepest child node; determining a first score of the entity based on the number of child nodes and the maximum path length; if the first score is greater than a third threshold, determining that the node corresponding to the entity is an encapsulated node.
9. The method according to any one of claims 1 to 6, wherein The first development document is a code library; Extracting first semantic feature information corresponding to the functional module from the first development document includes: Extracting abstract syntax tree feature information corresponding to the functional module, data dependency feature information based on a data flow graph, and parameter feature information based on an interface description from the first development document; The second development document is a UI design drawing; Extracting second semantic feature information corresponding to the functional module from the second development document includes: Extracting element image feature information, element layout feature information, and interaction feature information corresponding to the functional module from the second development document; The second development document is a requirements document; Extracting second semantic feature information corresponding to the functional module from the second development document includes: Functional feature information, business process feature information, and demand priority feature information corresponding to the functional module are extracted from the second development document.
10. A component packaging device, characterized in that: include: A processor, wherein the processor is coupled to a memory, the memory is used to store a computer program or instruction, and the processor is used to execute the computer program or instruction to implement the method according to any one of claims 1 to 9.
11. A computer-readable storage medium, characterized in that The computer stores a computer program executable by a computer device. When the program is run on the computer device, the computer device is caused to execute the steps of the method according to any one of claims 1 to 9.
12. A computer program product, characterized in that When the method is executed on a computer, the computer is enabled to execute the method according to any one of claims 1 to 9.