Method for AI intelligent document capable of team cooperation

By performing syntax parsing and cross-language semantic fusion on the source code modules of heterogeneous programming languages, an interface function call relationship graph and semantic topology graph are constructed. Semantic conflicts are identified and optimized, solving the interface definition problem of heterogeneous programming languages ​​in multi-team collaborative development. This achieves automated unification and continuous optimization of interface definitions, improving software development efficiency and quality.

CN121387399APending Publication Date: 2026-01-23NANJING TONGDAHAI INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511242473.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

In collaborative development involving multiple people and teams, the interface definitions of heterogeneous programming languages ​​suffer from semantic conflicts and integration difficulties, resulting in low software development efficiency, poor quality, and poor maintainability. Existing methods that rely on manually defining interface specifications are inefficient, error-prone, and lack a unified standard.

Method used

By parsing the source code modules of heterogeneous programming languages, constructing an interface function call relationship graph, using a cross-language semantic fusion model to fuse semantic features, constructing an interface semantic topology graph, identifying semantic conflicts, and generating standardized interface definition documents through dynamic iterative optimization of interface definitions driven by document semantic entropy.

Benefits of technology

It significantly improves the accuracy and consistency of interface definitions, reduces code integration failures and maintenance difficulties, improves conflict identification efficiency and overall software development efficiency and quality, and realizes a closed loop of continuous optimization of interface definitions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121387399A_ABST
    Figure CN121387399A_ABST
Patent Text Reader

Abstract

The invention provides an AI intelligent document method capable of team cooperation, which comprises the following steps of: performing grammar analysis on a heterogeneous programming language source code module, and constructing an interface function calling relation graph; fusing the interface semantic features by using a cross-language semantic fusion model to obtain unified interface semantics; constructing an interface semantic topological graph and automatically identifying semantic conflicts; automatically adjusting the interface semantic granularity based on the standard function calling definition, and generating a unified standard interface definition; and defining iterative optimization by using a document semantic entropy driving interface. According to the method, automatic unification of interface definition and accurate semantic conflict recognition are effectively realized, and team development efficiency and software quality are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of software technology, specifically to a method for creating AI-powered intelligent documents that enable team collaboration. Background Technology

[0002] As software development teams grow larger and projects become more complex, multi-person, multi-team collaborative development has become the mainstream development model in the field of software engineering. In actual engineering development scenarios, different developers or teams usually develop functional modules based on their own familiar programming languages, such as Java, Python, and C++, resulting in a large number of source code modules implemented in heterogeneous programming languages. Even if a team has uniformly adopted a certain programming language, it still faces problems such as inconsistent interface definitions, inconsistent parameter naming conventions, and ambiguous data structure definitions. These problems often lead to semantic conflicts and integration difficulties in interface definitions between different modules, seriously affecting software development efficiency, code quality, and subsequent maintainability.

[0003] Existing solutions typically rely on manually developing interface specifications or maintaining interface definition documents. However, this approach is not only inefficient and error-prone, but also makes it difficult to achieve rapid iteration and continuous integration. At the same time, due to the lack of unified standards, the definitions of interface documents are often ambiguous, uncertain, or even conflicting, causing repeated integration failures during the software integration phase due to interface call failures or parameter passing errors. Furthermore, the ambiguity and uncertainty of interface definitions further increase the complexity and difficulty of automated interface testing and continuous delivery.

[0004] Therefore, how to automatically unify the semantics of cross-language interfaces, accurately identify semantic conflicts in interface definitions, automatically generate standardized interface definition documents, and on this basis, achieve continuous iterative optimization of interface definitions are the technical problems that urgently need to be solved. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes a method for AI intelligent documents that enable team collaboration.

[0006] To achieve the above objectives, the present invention provides a method for creating AI-powered intelligent documents for team collaboration, comprising: S101: Perform syntax parsing on the source code modules of heterogeneous programming languages ​​to obtain the input parameters, output parameters, and execution operations of the interface, and construct the interface function call relationship graph; S102: A cross-language semantic fusion model is used to perform cross-language semantic feature fusion processing on the nodes of the function call relationship graph to obtain unified interface semantic features; S103: Based on the unified interface semantic features and the actual calling relationship between interfaces, construct an interface semantic topology graph, and determine the set of interfaces with semantic conflicts based on the topology graph; S104: Adjust the semantic granularity of the set of semantically conflicting interfaces to obtain a unified standard interface definition; S105: Calculate the semantic entropy of the document defined by the standard interface. If the rate of change of the semantic entropy of the document is less than the preset convergence threshold twice in a row, output the standard interface definition. Otherwise, adjust the node connection relationship of the interface semantic topology graph according to the semantic entropy of the document and repeat steps S103 to S104 until the rate of change of the semantic entropy is less than the preset convergence threshold.

[0007] Compared with the prior art, the beneficial effects of the present invention are: This invention significantly improves the accuracy and consistency of interface definitions through automated syntax parsing and cross-language semantic feature fusion methods. It avoids problems such as inconsistent parameter naming and ambiguous interface semantics that are prone to occur when defining interfaces manually in the traditional way, thereby effectively reducing code integration failures and maintenance difficulties caused by unclear interface definitions.

[0008] This invention uses interface semantic topology graph construction and automatic semantic conflict identification technology to accurately discover implicit semantic differences and potential conflicts in interface definitions, replacing the traditional method of relying on manual investigation. This significantly improves the efficiency of conflict identification and reduces the workload of repeated iterations and manual debugging caused by semantic conflicts during software integration and testing.

[0009] This invention automatically achieves a continuous optimization loop for interface definitions through a document semantic entropy-driven dynamic iterative optimization mechanism. This avoids the high cost and error-proneness of manually modifying interface documents repeatedly, and realizes continuous integration and automated iteration of interface definitions, effectively ensuring the overall efficiency and delivery quality of collaborative software development.

[0010] This invention achieves automatic unification of interface definition semantics and accurate conflict identification, effectively improving the collaborative efficiency and quality of team software development. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of the structure of the method of the present invention. Detailed Implementation

[0013] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0014] Example 1 Please see Figure 1 This invention provides a method for creating AI-powered intelligent documents that enable team collaboration, comprising: S101: Perform syntax parsing on the source code modules of heterogeneous programming languages ​​to obtain the input parameters, output parameters, and execution operations of the interface, and construct the interface function call relationship graph; It should be noted that the heterogeneous programming languages ​​mentioned in this invention refer to source code modules written in at least two different programming languages, which differ in syntax, parameter definition, and calling methods, such as but not limited to Java, Python, and C++. Specifically, the method for constructing the interface function call relationship graph includes: Syntax analysis tools are used to parse the source code modules of heterogeneous programming languages, and to extract the input parameter names, data types, and output parameter names and data types for each interface. It should be noted that syntax analysis tools are used to parse the source code modules of heterogeneous programming languages. These syntax analysis tools include, but are not limited to, open-source tools such as ANTLR and Flex. During the parsing process, specific syntax rules are set for each programming language to ensure that the input parameter names and data types, as well as the output parameter names and data types of the interface definition are obtained accurately. Specifically, an interface definition can be represented by a set, and the set of interface input parameters is as follows: , in, Indicates the first The parameter name of each input parameter. This represents the corresponding parameter data type; similarly, the set of interface output parameters is denoted as... Its representation method is consistent with the input parameter set, and will not be elaborated here; It should be understood that the specific syntax rules refer to the parsing rules established according to the function or method definition syntax of each programming language. For example, for the Java language, the specific syntax rules are as follows: First, identify function definition statements in Java source code that begin with access modifiers (such as public, private, protected, etc.) and contain a return type, method name, and parameter list; then, further parse the comma-separated list of input parameters within parentheses, extract the type and name of each parameter, and finally extract the return type specified before the method definition as the data type of the output parameter (when the method is void, the output parameter is empty). For another example, for the C++ language, the specific syntax rules are as follows: First, identify function definition statements that begin with a return type, followed by a function name and a parameter list within parentheses; second, parse the list of input parameters within parentheses, parsing each comma-separated parameter item as a combination of data type and parameter name, and the return type before the function definition is the data type of the output parameter. Specifically, the above-mentioned specific syntax rules for Java and C++ languages ​​are merely illustrative examples of the embodiments of the present invention. For those skilled in the art, based on mastering the above methods, they can formulate corresponding specific syntax rules according to the standard syntax definition of the specific language to be parsed, so as to accurately parse the interface input parameters and output parameters of the language, thereby completely constructing the interface function call relationship diagram. Based on the call statements of the interface methods in the source code module, extract the name of the specific functional operation performed by each interface; In practice, based on the aforementioned syntax analysis tools, the function or method call statements of the source code module are parsed line by line to obtain the function or method names explicitly expressed in the call statements. The specific parsing process includes: setting syntax parsing rules for call statements for each programming language, that is, parsing specific keywords, method call formats, and syntax structures in the call statements; for example, in Java, the call statement is usually "object name.method name(parameter list)", and the "method name" is extracted by setting parsing rules; as another example, in Python, the call statement is usually "function name(parameter list)", and the parsing rules extract the "function name"; Furthermore, this invention establishes a preset functional vocabulary library to perform semantic functional matching on the parsed function or method names. Specifically, the preset functional vocabulary library collects publicly available function or method name data from public code repositories (such as GitHub, StackOverflow, and enterprise internal knowledge bases), and performs semantic vectorization processing on the function or method names using word embedding models in natural language processing (such as Word2Vec or GloVe models). Subsequently, based on clustering analysis (such as the K-Means algorithm), function or method names with semantic similarity greater than or equal to a preset similarity threshold are clustered into several functional clusters. Then, those skilled in the art manually define natural language phrases representing the functions implemented by each cluster as functional operation names, ultimately forming a functional vocabulary library with semantic matching capabilities. In the specific process of extracting function operation names, the present invention inputs the function or method name obtained by the lexical analysis into the above-mentioned function vocabulary library. By calculating the cosine similarity between the function name to be identified and the semantic vector of the center of each function cluster in the function vocabulary library, the functional phrase corresponding to the cluster with the highest similarity is selected as the specific function operation name of the interface, thereby realizing the clear definition and unified expression of the interface function. Treat each interface as a graph node, and determine the direct call relationships between interfaces as graph edges based on the function call statements in the source code module; In practical implementation, each interface is used as a node in the function call relationship graph. Based on the function call statements existing in the source code module, the direct call relationships between each interface are determined. That is, when there is a function call relationship between interface nodes, a connection edge is constructed between the two nodes. Specifically, the interface function call relationship graph is defined as follows: , where the node set Represents all interface nodes and the set of connecting edges. interface Calling the interface This indicates the direct calling relationship between interface nodes; For example, suppose there are three interface nodes, interface A, interface B, and interface C, in the source code module. Analyze the source code module and find the following function call statements: interface A explicitly calls interface B during execution, and interface B explicitly calls interface C during execution. Based on this call relationship, the function call relationship diagram described in this invention can be represented as follows: take interface A, interface B, and interface C as nodes, and construct connecting edges between the interface nodes according to the above explicit call relationship, based on the direction of function calls between the interfaces. Specifically, construct a directed connecting edge from node A to node B between node A and node B, and construct a directed connecting edge from node B to node C between node B and node C. The final function call relationship diagram can intuitively and accurately describe the call relationship between the interface nodes, that is: A→B→C. For interface node pairs that are not directly called but whose input and output parameter names and data types are exactly the same, corresponding semantic association edges are added. Specifically, for each pair of interface nodes that do not have a direct calling relationship, their corresponding input parameter sets and output parameter sets are compared one by one; if the input parameter names and corresponding data types of the two interface nodes are completely identical and the output parameter names and corresponding data types are also completely identical, then it is determined that there is a semantic functional relationship between the two interface nodes, and a semantic relationship edge is established. It should be noted that the semantic association edges are represented by undirected dashed lines in the function call relationship graph. This means that there is similarity or substitutability between interface nodes in terms of data definition structure and potential function, but it does not indicate the direction of the call. For example, if the input and output parameters between interface nodes D and E are exactly the same, a semantic association edge represented by an undirected dashed line is added to the interface function call relationship graph to reflect the similarity between the two interface nodes in terms of function definition and interface granularity. The output includes an interface function call relationship graph containing interface nodes, direct call relationship edges, and semantic association edges; It is understandable that by extracting all interface nodes, direct call relationship edges, and semantic association edges through the above steps, a complete interface function call relationship graph is ultimately formed and output. Specifically, the interface function call relationship graph can be represented as a graph structure. ,in, For a set of interface nodes, This represents the set of edges indicating direct call relationships between interface nodes. This represents the set of semantic association edges established between interface nodes whose input and output parameter names and data types are exactly the same but have no calling relationship. It should be noted that the interface function call relationship graph constructed in this invention only focuses on the direct call relationship between interfaces and the semantic association relationship with the same interface signature, and does not consider the indirect call relationship between interface nodes through intermediate nodes; since indirect call relationships are unclear in cross-team collaborative development and are very easy to change in the process of business logic changes, they are not included as an explicit component of the graph structure in this invention, so as to ensure the stability and clarity of the interface call relationship definition. S102: A cross-language semantic fusion model is used to perform cross-language semantic feature fusion processing on the nodes of the function call relationship graph to obtain unified interface semantic features; In implementation, the construction of the cross-language semantic fusion model specifically includes: Collect historical code interface fragments that have the same behavior and function but are implemented in different programming languages; It should be noted that, through data collection tools, historical code interface fragments with completely consistent interface definitions and behaviors but implemented in different programming languages ​​(such as Java, Python, C++, etc.) are automatically collected from public code repositories such as GitHub, GitLab, and SourceForge, and used as the basic data source for training the cross-language semantic fusion model; It should also be noted that the code interface fragment includes input parameters and output parameters, that is, it provides the input parameters required when calling the function or method (such as user ID, file path, encryption key, etc.) and the output parameters returned after the interface is executed (such as boolean value, JSON data, status code, etc.); and the code interface fragment can implement specific functions that are marked or described by the developers, such as "user authentication", "data encryption and decryption", "network data transmission", "database query" and other functions. For example, historical code interface fragments obtained from public code repositories may appear as follows: a "user login authentication" interface fragment implemented in Java defines input parameters (username, password) and output parameters (authentication success indicator), and has a code comment stating "implements user login authentication function"; while in an interface fragment with the same function implemented in Python, the input parameters are also username and password, the output parameter is also an authentication success indicator, and there is also a code comment describing it as "implements user authentication function". Because these code fragments have clear functions, well-defined interfaces, and have been publicly released in different programming languages, they can serve as typical representatives of historical code interface fragments in this invention, used to train the cross-language semantic fusion model proposed in this invention. For each code interface segment, extract the interface's input parameters, output parameters, and corresponding functional actions; It should be noted that the extraction method is based on a syntax analysis tool, which will not be elaborated on here; The initial semantic feature encoding is obtained by using a pre-trained cross-language semantic mapping model to encode the interface input parameters, output parameters, and functional actions. The pre-training process of the cross-language semantic mapping model specifically includes: Collect code interface fragments in different languages ​​and their corresponding functional description texts from publicly available code repositories; A bidirectional attention encoder is used to extract features from the function call sequence and the corresponding function description text of the code snippet to obtain an initial feature embedding vector; It should be noted that the specific implementation of the bidirectional attention encoder can adopt the self-attention mechanism of the existing Transformer model to extract features from the function call sequence and the function description text sequence respectively, and generate corresponding initial feature embedding vectors. For example, for an input sequence X = {x1, x2, …, x} of length n n The formula for calculating the feature embedding vector of the bidirectional attention mechanism is as follows: , In the formula, Q represents the query matrix, i.e., the feature query of the input sequence itself; K represents the key matrix, i.e., the feature keys of the input sequence; and V represents the value matrix, i.e., the feature values ​​of the input sequence. The dimension of each vector in the key matrix, used for normalization; Calculate the cosine distance between the initial feature embedding vector of the code snippet and the initial feature embedding of the corresponding functional description text, and perform comparative training based on this cosine distance; For example, the initial feature embedding vector of the function call sequence is calculated. Initial feature embedding vectors of the corresponding functional description text The cosine distance between them is used for comparative training. The formula for calculating the cosine distance is as follows: , This represents the dot product of the code feature vector and the text feature vector. These represent the magnitudes of the code feature vector and the text feature vector, respectively. This represents the cosine distance between code feature embeddings and text feature embeddings. When the training iterations reach a point where the cosine distance is less than a preset convergence threshold, the trained cross-lingual semantic mapping model is output. It should be noted that when the cosine distance The smaller the value, the higher the semantic similarity between the two feature embedding vectors. The goal of training is to minimize the cosine distance until the cosine distance value is stably less than the preset convergence threshold (e.g., 0.05), at which point training ends and the trained cross-language semantic mapping model is output. It should be understood that the convergence threshold is obtained by statistically analyzing the range of cosine distance changes during the training process through repeated training experiments. When the cosine distance stabilizes to a certain range and the model performance no longer improves significantly (e.g., the improvement in classification accuracy or semantic matching accuracy does not exceed 1%), the present invention takes the average cosine distance within this stable range as the preset convergence threshold in the formal training process. The specific value can be adjusted according to the accuracy requirements of different application scenarios. A self-supervised contrastive learning algorithm is adopted to optimize the initial encoding of semantic features based on the same or different relationships between the functional action execution sequences of code interface fragments, so as to form a unified cross-language semantic feature and output the trained cross-language semantic fusion model. In practice, based on the initial semantic feature encoding vector of each code interface segment, pairs of code interface segments with the same functional action execution sequence (positive sample pairs) and pairs of code interface segments with different functional action execution sequences (negative sample pairs) are constructed. A contrastive learning loss function is then used to train and optimize these positive and negative sample pairs. The mathematical expression for the contrastive learning loss function is as follows: , In the formula, This represents the total number of sample pairs used for training. Indicates the first The label value of a sample pair, when the sample pair has the same functional action sequence (i.e., a positive sample pair), When the functional action sequences of sample pairs are different (i.e., negative sample pairs), then , Indicates the first The similarity between feature vectors in a sample pair; It should be noted that the similarity is calculated by... The cosine similarity of the initial semantic feature encoding vectors of two code segments in a sample pair is obtained; During training, this invention optimizes the contrastive learning loss function LLL to gradually increase the similarity of feature encoding vectors between code segments executing the same functional action sequence, while decreasing the similarity of feature encoding vectors between code segments executing different functional actions. After multiple rounds of iterative optimization, when the contrastive learning loss function is optimized in two consecutive iterations... When the decrease is less than the preset stability threshold, the training process is considered to have reached a stable convergence state, and the cross-language unified semantic features after stable convergence are output, thus obtaining the trained cross-language semantic fusion model. It should be understood that the preset stability threshold was obtained through statistical analysis of multiple experiments, which will not be elaborated on here. S103: Based on the unified interface semantic features and the actual calling relationship between interfaces, construct an interface semantic topology graph, and determine the set of interfaces with semantic conflicts based on the topology graph; In specific implementation, the interface semantic topology graph is a graph structure built based on unified interface semantic features and interface call relationships. It is used to intuitively reflect the semantic connections and differences between interfaces, thereby effectively identifying interface node pairs with semantic conflicts. Specifically, the method for constructing the interface semantic topology graph includes: The interfaces corresponding to the obtained unified interface semantic features are respectively used as nodes in the topology graph; It should be noted that each node in the interface semantic topology graph is the interface corresponding to the unified interface semantic feature. In other words, the interfaces represented by the nodes in the topology graph have eliminated the semantic differences between different programming languages ​​and have a unified semantic feature representation. Traverse each pair of interfaces with a calling relationship recorded in the function call relationship graph, and treat the calling direction as a directed connection edge between nodes; Understandably, to reflect the actual functional call relationships between interface nodes, this invention traverses the constructed functional call relationship graph and, for each pair of interface nodes with a call relationship, constructs directed connection edges between the nodes in the direction of the interface call; specifically, when there are interface nodes... Calling the interface node At that time, construct a slave node in the topology graph. Pointing to node Directed edges are used to visually represent the actual call order and functional dependencies between interfaces; Extract the complete set of data fields for the input and output parameters corresponding to each interface from the unified interface semantic features; It should be understood that, in order to accurately represent the functional definition details of each interface node, this invention extracts a complete set of data fields for the input and output parameters corresponding to each interface node from the unified interface semantic features; specifically, the complete set of data fields is defined as the set of names of the interface input and output parameters and their corresponding data types; for example, if the input parameter corresponding to a certain interface node is... Integer ,String The output parameters are Boolean Then the complete set of data fields corresponding to this interface node is the union of the sets of input parameters and output parameters mentioned above, that is... Integer ,String Boolean ; Compare each pair of interface nodes that do not have a direct calling relationship but have semantic association edges. If their data field sets of input and output parameters are completely identical, inherit and continue the semantic association edges already established in the above function calling relationship graph, and add corresponding undirected connection edges in the interface semantic topology graph. It should be noted that the undirected connection edges in the interface semantic topology graph directly correspond to the existing undirected dashed semantic association edges in the function call relationship graph, emphasizing the structural similarity between interfaces at the unified semantic abstraction level, and further demonstrating the semantic granularity consistency between interface nodes in the semantic topology graph through undirected solid lines. For example, in the interface function call relationship graph, interface node D and interface node E have established a semantic association edge through an undirected dashed line. When constructing the interface semantic topology graph, this invention inherits this relationship as an undirected solid line connection edge to reflect that the two have a stable structural similarity at the unified semantic level, which facilitates the subsequent identification of interface semantic conflicts. It should be understood that the interface semantic topology graph inherits not only the direct call relationships (represented by directed solid lines) from the function call relationship graph during its construction, but also the semantic association relationships (i.e., undirected dashed lines established by identical input and output parameters) from the function call relationship graph. This inheritance mechanism ensures the logical continuity and information integrity between the interface function call relationship graph and the interface semantic topology graph, preventing semantic association information from being ignored or lost during the construction of the semantic topology graph. Specifically, the interface function call relationship graph is defined as: a set of nodes. Directly call the relation edge set Call Semantic related edge set The interface semantic topology graph inherits and redefines itself as a set of nodes that inherits from the interface function call relationship graph, and a set of directed edges that inherits from the set of direct call relationship edges. The set of undirected connected edges inherits from the set of semantically related edges. ; Output the interface semantic topology graph, which consists of nodes, directed edges, and undirected edges. It should be understood that the final output interface semantic topology graph is a combination of nodes (representing unified semantic interfaces), directed edges (representing interface call relationships), and undirected edges (representing potential functional similarity relationships between interfaces). This topology graph can comprehensively and intuitively reflect the specific relationships between interface nodes in terms of functional calls and semantic features, laying the foundation for subsequent identification of interface semantic conflicts; Specifically, the method for determining the set of interface semantic conflicts includes: Extract the unified interface semantic features corresponding to each interface node in the interface semantic topology graph; Initial function tags are generated using unified interface semantic features, where the initial function tags are natural language phrases that describe the specific functions of the interface. It should be noted that generating initial function tags using unified interface semantic features involves generating natural language phrases describing the functional behavior of each interface node based on the unified semantic features of each interface node using a natural language generation model (such as BERT or GPT model). These phrases are called initial function tags. Specifically, each initial function tag consists of several natural language words, such as "user login authentication" and "account information query," to represent the functions implemented by the interface. The initial function tags are compared pairwise to determine the overlap of function tags between any two interface nodes. The overlap represents the proportion of the number of shared words in function tags to the total number of tags. In practice, the label overlap degree is defined to evaluate the semantic similarity of functional labels of any two interface nodes. The formula for calculating the label overlap degree is as follows: , In the formula, Indicates interface node With interface node The degree of overlap in functional tags between them Indicates interface node With interface node The number of common words contained in the initial functional tags. Indicates interface node With interface node The total number of all unique words in the initial feature tags; If the overlap of the functional labels is less than the preset conflict threshold, then the interface node pair is determined to have a semantic conflict. It should be noted that the preset conflict threshold is used to determine whether there is a semantic conflict between interface node pairs; when the tag overlap... When the value is less than the preset conflict threshold, it is considered that the functional semantic difference between the two interface nodes is significant, and a semantic conflict is determined to exist. It should be understood that the preset conflict threshold is determined automatically by collecting interface function tag data from the public code library and using a cluster analysis method to select the most obvious "inter-cluster boundary" between the interface function tag overlap clusters as the preset conflict threshold based on the boundary value between the clusters. For example, if the cluster analysis finds that the function tag overlap sample data is obviously divided into two clusters, with the overlap value of one cluster mainly concentrated between 0 and 0.28, and the overlap value of the other cluster mainly concentrated between 0.32 and 1.0, then the preset conflict threshold is determined to be the boundary value between the two clusters, for example, 0.30. All output pairs of interface nodes that are determined to have semantic conflicts are used to form a set of interface semantic conflicts for subsequent adjustment of interface semantic granularity. For example, the present invention outputs all interface node pairs that are determined to have semantic conflicts as a set, forming an interface semantic conflict set, specifically defined as: , in, This represents the set of interface semantic conflicts, that is, the set of all interface node pairs that are determined to have semantic conflicts. This represents any two distinct interface nodes. S104: Adjust the semantic granularity of the set of semantically conflicting interfaces to obtain a unified standard interface definition; In practice, the specific implementation methods for adjusting the interface semantic granularity include: Extract the specific execution operation and corresponding data fields from each semantic conflict interface; For example, based on the set of semantic conflicts in the interfaces, the specific operation performed by each semantic conflict interface node and the complete set of data fields involved in the operation are extracted. The specific operation refers to the specific behavior of each interface at the functional implementation level, such as "user login verification", "data file upload", "information encryption processing", etc.; while the corresponding data fields specifically refer to the names of the input parameters and output parameters of the above interface operation and their corresponding data type sets, such as the input parameters being "user ID (String)" and "password (String)", and the output parameters being "verification status (Boolean)", etc. For interfaces with different execution operations, the execution operations in each interface are broken down one by one, and each execution operation is independently verified to fully realize the corresponding original function call; It should be noted that the method for generating the original function call includes: obtaining an initial set of call logs by collecting function or method call log data in the actual operating environment; using sequence mining algorithms (such as the PrefixSpan algorithm) and frequent pattern mining algorithms (such as the FP-Growth algorithm) to extract call patterns with high call frequency and relatively stable parameter definitions from the initial set of call logs, forming a candidate set of function call patterns; for the candidate set of function call patterns, this invention further introduces a functional semantic verification step to eliminate data defects and adapt to changes in business logic; for the call patterns that have undergone the above automated semantic verification and manual confirmation, a natural language generation model (such as the GPT model) is used to generate a corresponding structured "original function call", including function identifier, function description text, input parameter set and output parameter set, thereby ensuring that the generated "original function call" can accurately reflect the current real business logic and has a certain degree of adaptability to cope with future business changes; The functional semantic verification step includes: taking the code snippets and related comments and documentation corresponding to the candidate function call patterns as auxiliary inputs, and automatically evaluating the semantic validity and consistency of the call patterns using a pre-trained code-text joint semantic model (such as the CodeBERT model) to ensure that the mined call patterns match the actual business semantic description; ranking the call patterns according to the semantic consistency score automatically evaluated by the model, and automatically filtering out call patterns with semantic consistency below a preset threshold; further, manually confirming the call patterns that have passed the automatic filtering to see if the business functions corresponding to the call patterns have changed or iterated in the business requirements; if there are changes or iterations, updating the call pattern set with the new business logic through the manual verification results, replacing the old call patterns, and avoiding the system from hindering the normal evolution of business logic; For example, a raw function call can be represented by the following structured data definition: Function ID: A string identifier that uniquely identifies a specific business function, such as "validateUserLogin"; Function Description: A brief natural language description that explains the specific operation or business purpose performed by the function, such as "user login verification"; Input Parameter Set: , in, For the name of the input parameter (e.g., "userld", "password"), The data type of the corresponding parameter (such as String, Integer, Boolean, etc.); Output Parameter Set: , in, The name of the output parameter (e.g., "authStatus"). The data type of the corresponding parameter (such as Boolean, JSON object, Binary data, etc.); It should be understood that the present invention performs fine-grained decomposition of the execution operations of the interface node. That is, when an interface node defines multiple execution operations, the present invention decomposes each independent execution operation one by one, and performs integrity verification of the original function call for each of the decomposed execution operations independently. The verification method for whether the execution operation can completely realize the corresponding original function call specifically includes: Extract the corresponding input and output parameters for each operation. For example, each execution operation is broken down. If a certain execution operation after being broken down is "file download", then its corresponding set of input parameters (e.g., "file ID (Integer)" and "target path (String)") and corresponding set of output parameters (e.g., "download success status (Boolean)" and "file data (Binary)") are extracted. Analyze the complete input conditions required for the original function call corresponding to the operation; It is understood that this invention analyzes the complete input conditions necessary for each execution operation to perform the corresponding original function call, specifically defined as the set of all input parameters required to implement the function call; the complete input conditions are determined based on business logic; for example, the complete input conditions for the "file download" function call may be that it must have three input parameters: "file ID (Integer)", "target path (String)" and "authorization verification token (String)"; Compare whether the input parameters for the operation cover all parameters required for a complete set of input conditions; It should be noted that if the set of input parameters for the operation completely covers all the required parameters, it means that the operation can fully meet the requirements of the original function call at the input end; otherwise, it is considered not to meet the requirements. Compare whether the output parameters of the executed operation cover the complete output results required by the original function call; It should be noted that the complete output of the original function call is also based on the set of data fields defined by the business logic. For example, the complete output of the "file download" function call can be defined as having two output parameters: "download success status (Boolean)" and "file data (Binary)". It should also be noted that the actual output parameters of the operation are compared item by item with the complete output result defined above. If the actual output parameter set completely covers all the parameters of the complete output result, then the output meets the integrity requirement; otherwise, it does not. An execution operation whose input and output parameters both satisfy the requirements of the original function call is considered an execution operation that can fully realize the original function call. The execution operation that can fully realize the original function call is defined as a new interface; It should be understood that for an execution operation whose output and input / output parameters both meet the integrity requirements of the original function call, it is defined as a separate new interface to eliminate conflicts caused by inconsistent semantic granularity in the interface semantic conflict set. Specifically, for each execution operation that can independently and completely realize the original function call, the present invention assigns a new unique interface identifier and defines the corresponding set of input and output parameters for that interface to form a new interface definition. For example, assuming that the "file download" operation has been verified to meet the complete input and output requirements, the present invention defines the "file download" operation as a new standard interface. This standard interface definition includes complete input parameters (e.g., "file ID (Integer)", "target path (String)", "authorization verification token (String)") and complete output parameters (e.g., "download success status (Boolean)", "file data (Binary)"), and assigns a new interface identifier (e.g., "downloadFile"). For execution operations that cannot achieve the original function call on their own, merge other execution operations related to the function to generate a new combined interface definition; Understandably, for execution operations that cannot fully realize the original function call requirements on their own after verification, the functional correlation between the execution operation and other execution operations in the interface semantic conflict set is analyzed. Specifically, multiple execution operations that are closely related and jointly realize the specific original function call requirements are merged to generate a new combined interface definition. In the specific implementation process, each execution operation that cannot realize the original function call on its own is compared with the correlation of data fields and the similarity of functional semantics with other execution operations in the function call process to determine the combination of execution operations that can be merged. For example, if the execution operation "permission verification" of a certain interface node cannot fully realize the original function call on its own, but can be combined with another execution operation "file download" to jointly realize the original function call requirement, then the present invention defines the above two execution operations as a new interface after combination. The new interface definition includes all input parameters and output parameters of the two sets of execution operations, and then assigns a new interface identifier (e.g., "validateAndDownloadFile"). Output all functionally verified new interface definitions as standard interface definitions; It should be understood that all the new interface definitions that have undergone functional verification are used as unified standard interface definitions. The output standard interface definitions fully include interface identifiers, specific functional descriptions, complete sets of input parameters, and complete sets of output parameters, thereby ensuring clear and unified guidance for interface implementation and invocation in cross-language and cross-team collaborative development environments. For example, in a specific implementation, if the "file download" operation satisfies the complete input and output requirements, then the new interface is defined as follows: Interface ID: downloadFile, Function Description: File download function interface, Input parameter set: File ID (Integer), Target path (String), Permission verification Token (String), Output parameter set: Download success status (Boolean), File data (Binary); If the execution operation "Permission verification" of a certain interface node cannot fully realize the original function call on its own, but can be combined with another execution operation "file download" to jointly realize the original function call requirements, then the new interface is defined as follows: Interface ID: validateAndDownloadFile, Function Description: Permission verification and file download combined function interface, Input parameter set: User ID (String), Permission verification Token (String), File ID (Integer), Target path (String), Output parameter set: Permission verification status (Boolean), Download success status (Boolean), File data (Binary); S105: Calculate the semantic entropy of the document defined by the standard interface. If the change rate of the semantic entropy of the document is less than the preset convergence threshold twice in a row, output the standard interface definition. Otherwise, adjust the node connection relationship of the interface semantic topology graph according to the semantic entropy of the document and repeat steps S103 to S104 until the change rate of the semantic entropy is less than the preset convergence threshold. It should be noted that document semantic entropy reflects the degree of uncertainty in the semantics of the document in the interface definition. The smaller the document semantic entropy, the clearer and more unified the interface definition is, while the larger the semantic entropy, the higher the ambiguity or uncertainty in the interface definition. In implementation, the method for adjusting the node connection relationships of the interface semantic topology graph using document semantic entropy specifically includes: The interface node pairs that need to be adjusted are determined based on the document semantic entropy value; The method for determining the interface node pairs that need adjustment specifically includes: For each interface node, count the number of parameters that do not satisfy the original function call in the standard interface definition, and calculate the proportion of the number of such parameters to the total number of parameters of that interface node. It should be understood that, for each interface node in the interface semantic topology diagram, the number of parameters that do not meet the original functional call requirements in the current standard interface definition is counted, and the proportion of these unmet parameters to the total number of parameters of that interface node is further calculated. The specific formula is as follows: , In the formula, Indicates the first The proportion of interface node parameters that do not meet the original function call requirements. Indicates the first The number of parameters in each interface node that do not meet the original function call requirements. Indicates the first The total number of parameters for each interface node; Based on the above ratio, the document semantic entropy of each interface node is calculated using the information entropy formula; Specifically, based on the above ratio, the document semantic entropy of each interface node is calculated using the information entropy formula, which is as follows: , In the formula, This represents the document semantic entropy of the i-th interface node. Represents the logarithmic function with base 2; It should be noted that when the above formula contains a parameter ratio... or At that time, due to Undefined, this invention defines the corresponding document semantic entropy value as 0, that is, when all interface node parameters meet or do not meet the original function call requirements, its document semantic entropy is defined as 0, indicating that the semantic determinism reaches its maximum at this time; All interface nodes are sorted according to the calculated document semantic entropy, and the interface nodes sorted from largest to smallest entropy value are assigned adjustment priority from high to low. It is understandable that this invention applies the document semantic entropy of all interface nodes. Sort the interfaces from largest to smallest. The higher the entropy value, the greater the semantic uncertainty of the interface node. Therefore, the higher the priority in the subsequent adjustment process. Based on the sorted adjustment priority, select the interface nodes that are greater than the preset priority threshold, and combine the selected interface nodes in pairs. Output the combined interface node pairs as the interface node pairs that need to be adjusted. It should be noted that the preset priority threshold is determined by preliminary statistical analysis based on the sorted node list, using the distribution of node document semantic entropy values ​​to determine the proportion of nodes with high semantic entropy. For example, by analyzing the entropy value statistical distribution of multiple actual interface documents, it was found that the set of interface nodes with significantly higher document semantic entropy values ​​(e.g., the entropy values ​​of the top 20% of nodes are higher than those of the other nodes) contributes the most to the reduction of overall document semantic entropy and has the most significant adjustment effect. It should be understood that this percentage (e.g., 20%) is an empirical value determined based on the distribution of semantic entropy of multiple interface nodes in the actual project. In practice, those skilled in the art can adjust this value appropriately based on the complexity of the interface definition, the characteristics of the document semantic entropy distribution, and the actual requirements for the stability of the interface definition in the specific project, without departing from the concept of this invention. For the interface nodes that need adjustment, re-extract the actual call relationships and shared data fields between the interfaces; It should be noted that the process of re-executing the actual call relationship between interfaces and extracting shared data fields involves re-using the syntax analysis tool and data field extraction method to re-analyze whether there is an actual call relationship between interface node pairs from the current source code module or standard interface definition, and further extract the shared data field information between interface node pairs. The new connection relationships between interface nodes are determined based on the latest actual call relationships and data field sharing. In the specific implementation process, the rules for defining new connection relationships between interface nodes are as follows: If there is an actual calling relationship between two interface nodes (e.g., interface nodes...) Calling the interface node Then, a slave node is established in the interface semantic topology graph. Pointing to node Directed connecting edges; If there is no calling relationship between two interface nodes, but the set of data fields shared between the two interface nodes is completely identical (i.e. the input and output parameter names and data types are exactly the same), then an undirected connection edge is established between the interface node pair to reflect the potential functional correlation between the interface nodes. If there is neither an actual calling relationship nor a data field sharing relationship between the interface node pairs, then no connection edge will be established to reflect the independence of the interface nodes in terms of functional implementation; Remove the existing connections between interface nodes in the topology graph, update them with new connections, and output the updated interface semantic topology graph. It should be understood that the original connection relationships (including directed and undirected connections) between corresponding interface nodes in the original interface semantic topology graph are removed, and then the new connection relationships between interface nodes are added to the interface semantic topology graph to update the interface semantic topology graph structure so as to reflect the semantic relationship of the interface definition after the latest adjustment. It should be noted that the updated interface semantic topology graph structure specifically includes the latest combined relationship graph of nodes (representing interfaces), directed connections (call relationships), and undirected connections (shared data field relationships); the output updated interface semantic topology graph will serve as the input basic data for the new round of interface semantic conflict set determination process, and will be continuously iterated and adjusted until the interface semantic topology graph structure is stable, the document semantic entropy change rate is less than the preset convergence threshold for two consecutive times, and the standard interface definition is output. It should also be noted that the aforementioned consecutive document semantic entropy refers to the weighted average of the document semantic entropy of all interface nodes to obtain a global document semantic entropy. The rate of change of the global document semantic entropy during the two consecutive iterations is calculated using the following formula: , in, This represents the global document semantic entropy obtained in the Kth iteration. If the rate of change of the document semantic entropy in two consecutive iterations is greater than or equal to the preset convergence threshold, the iteration adjustment continues. If the rate of change of the document semantic entropy in two consecutive iterations is less than the preset convergence threshold, the standard interface definition is considered to have reached a stable convergence state, and the standard interface definition is directly output as the final result.

[0015] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired or wireless network. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0016] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only one method, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0017] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0018] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0019] Some of the data in the above formula are calculated by removing dimensions and taking their numerical values. The formula is the closest to the real situation obtained by software simulation of a large amount of collected data. The preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data.

[0020] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.

Claims

1. A method for creating AI-powered intelligent documents for team collaboration, characterized in that: include: S101: Perform syntax parsing on the source code modules of heterogeneous programming languages ​​to obtain the input parameters, output parameters, and execution operations of the interface, and construct the interface function call relationship graph; S102: A cross-language semantic fusion model is used to perform cross-language semantic feature fusion processing on the nodes of the function call relationship graph to obtain unified interface semantic features; S103: Based on the unified interface semantic features and the actual calling relationship between interfaces, construct an interface semantic topology graph, and determine the set of interfaces with semantic conflicts based on the topology graph; S104: Adjust the semantic granularity of the set of semantically conflicting interfaces to obtain a unified standard interface definition; S105: Calculate the semantic entropy of the document defined by the standard interface. If the rate of change of the semantic entropy of the document is less than the preset convergence threshold twice in a row, output the standard interface definition. Otherwise, adjust the node connection relationship of the interface semantic topology graph according to the semantic entropy of the document and repeat steps S103 to S104 until the rate of change of the semantic entropy is less than the preset convergence threshold.

2. The method for a team-collaborative AI-powered intelligent document according to claim 1, characterized in that, The method for constructing the interface function call relationship graph includes: Syntax analysis tools are used to parse the source code modules of heterogeneous programming languages, and to extract the input parameter names, data types, and output parameter names and data types for each interface. Based on the call statements of the interface methods in the source code module, extract the name of the specific functional operation performed by each interface; Treat each interface as a graph node, and determine the direct call relationships between interfaces as graph edges based on the function call statements in the source code module; For interface node pairs that are not directly called but whose input and output parameter names and data types are exactly the same, corresponding semantic association edges are added. The output is a graph of interface function call relationships, including interface nodes, direct call relationship edges, and semantic association edges.

3. The method for a team-collaborative AI-powered intelligent document according to claim 1, characterized in that, The construction of the cross-language semantic fusion model includes: Collect historical code interface fragments that have the same behavior and function but are implemented in different programming languages; For each code interface snippet, extract the interface's input parameters, output parameters, and corresponding functional actions; The initial semantic feature encoding is obtained by using a pre-trained cross-language semantic mapping model to encode the interface input parameters, output parameters, and functional actions. A self-supervised contrastive learning algorithm is adopted to optimize the initial encoding of semantic features based on the same or different relationships between the functional action execution sequences of code interface fragments, so as to form a unified cross-language semantic feature and output the trained cross-language semantic fusion model.

4. The method for a team-collaborative AI-powered intelligent document according to claim 3, characterized in that, The pre-training process of the cross-language semantic mapping model includes: Collect code interface fragments in different languages ​​and their corresponding functional description texts from publicly available code repositories; A bidirectional attention encoder is used to extract features from the function call sequence and the corresponding function description text of the code snippet to obtain an initial feature embedding vector; Calculate the cosine distance between the initial feature embedding vector of the code snippet and the initial feature embedding of the corresponding functional description text, and perform comparative training based on this cosine distance; When the training iterations reach a point where the cosine distance is less than a preset convergence threshold, the trained cross-lingual semantic mapping model is output.

5. The method for a team-collaborative AI-powered intelligent document according to claim 1, characterized in that, The method for constructing the interface semantic topology graph includes: The interfaces corresponding to the obtained unified interface semantic features are respectively used as nodes in the topology graph; Traverse each pair of interfaces with a calling relationship recorded in the function call relationship graph, and treat the calling direction as a directed connection edge between nodes; Extract the complete set of data fields for the input and output parameters corresponding to each interface from the unified interface semantic features; Compare each pair of interface nodes that do not have a direct calling relationship but have semantic association edges. If their data field sets of input and output parameters are completely identical, inherit and continue the semantic association edges already established in the above function calling relationship graph, and add corresponding undirected connection edges in the interface semantic topology graph. Output the interface semantic topology graph, which consists of nodes, directed edges, and undirected edges.

6. The method for a team-collaborative AI-powered intelligent document according to claim 1, characterized in that, The method for determining the set of interface semantic conflicts specifically includes: Extract the unified interface semantic features corresponding to each interface node in the interface semantic topology graph; Initial function tags are generated using unified interface semantic features, where the initial function tags are natural language phrases that describe the specific functions of the interface. The initial function tags are compared pairwise to determine the overlap of function tags between any two interface nodes. The overlap represents the proportion of the number of shared words in function tags to the total number of tags. If the overlap of the functional labels is less than the preset conflict threshold, then the interface node pair is determined to have a semantic conflict. Output all interface node pairs that are determined to have semantic conflicts, forming an interface semantic conflict set.

7. The method for AI-powered intelligent documents enabling team collaboration according to claim 1, characterized in that, The specific implementation method for adjusting the semantic granularity of the interface includes: Extract the specific execution operation and corresponding data fields from each semantic conflict interface; For interfaces with different execution operations, the execution operations in each interface are broken down one by one, and each execution operation is independently verified to fully realize the corresponding original function call; The execution operation that can fully realize the original function call is defined as a new interface; For execution operations that cannot achieve the original function call on their own, merge other execution operations related to the function to generate a new combined interface definition; Output all new interface definitions that have been functionally verified, and use them as standard interface definitions.

8. The method for a team-collaborative AI-powered intelligent document according to claim 7, characterized in that, The verification method for whether the execution operation can completely realize the corresponding original function call includes: Extract the corresponding input and output parameters for each operation. Analyze the complete input conditions required for the original function call corresponding to the operation; Compare whether the input parameters for the operation cover all parameters required for a complete set of input conditions; Compare whether the output parameters of the executed operation cover the complete output results required by the original function call; An execution operation whose input and output parameters both satisfy the requirements of the original function call is considered an execution operation that can fully realize the original function call.

9. The method for a team-collaborative AI-powered intelligent document according to claim 1, characterized in that, The method for adjusting the node connection relationships of the interface semantic topology graph using document semantic entropy includes: The interface node pairs that need to be adjusted are determined based on the document semantic entropy value; For the interface nodes that need adjustment, re-extract the actual call relationships and shared data fields between the interfaces; The new connection relationships between interface nodes are determined based on the latest actual call relationships and data field sharing. Remove the existing connections between interface nodes in the topology graph, update them with new connections, and output the updated interface semantic topology graph.

10. The method for a team-collaborative AI-powered intelligent document according to claim 9, characterized in that, The method for determining the interface node pairs that need adjustment includes: For each interface node, count the number of parameters that do not satisfy the original function call in the standard interface definition, and calculate the proportion of the number of such parameters to the total number of parameters of that interface node. Based on the above ratio, the document semantic entropy of each interface node is calculated using the information entropy formula; All interface nodes are sorted according to the calculated document semantic entropy, and the interface nodes sorted from largest to smallest entropy value are assigned adjustment priority from high to low. Based on the sorted adjustment priority, select interface nodes that are greater than the preset priority threshold, and combine the selected interface nodes in pairs. Output the combined interface node pairs as the interface node pairs that need to be adjusted.