Code and non-code resource bidirectional mapping and consistency maintenance method and system

By introducing a two-way mapping between code and non-code resources and a consistency maintenance method into the software development process, the problems of excessive coupling between task management and code development and semantic gaps are solved. It realizes the automatic conversion of task constraints to code verification logic and the generation of structured knowledge objects, thereby improving the collaborative efficiency and knowledge reuse capability of software development.

CN121349508AActive Publication Date: 2026-01-16189CSP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511913570.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-01-16
Estimated Expiration
2045-12-18

AI Technical Summary

Technical Problem

In the existing software development process, task management and code development are too closely coupled, there is a semantic gap between requirement documents and code implementation, and there is a lack of automated experience accumulation and knowledge reuse mechanisms, resulting in high collaboration thresholds, difficulty in auditing task closure loops, difficulty in reusing knowledge, and documentation lagging behind implementation.

Method used

It adopts a bidirectional mapping and consistency maintenance method between code and non-code resources. By configuring independent storage space and associated identifiers, it uses an intelligent semantic processing module to realize the automatic conversion of task constraints into code verification logic, parses code change characteristics to generate structured knowledge objects, and introduces a multi-agent collaborative mechanism for real-time cross-analysis and synchronization.

Benefits of technology

It achieves deep decoupling and dynamic synchronization between tasks and code, ensures that the requirements document provides mandatory guidance for code development, makes implicit development experience explicit, improves the traceability and knowledge reuse capability of the entire software lifecycle, and reduces the probability of code merging conflicts and rework costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121349508A_ABST
    Figure CN121349508A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer data processing, discloses a code and non-code resource bidirectional mapping and consistency maintenance method and system, and aims to effectively decouple tasks and codes and keep strong semantic association and automatic synchronization between the tasks and the codes. The method comprises the following steps: configuring a first storage space and a second storage space which are physically independent and logically associated to store source codes and task descriptions respectively; generating a global unique association identifier and realizing cross-space entity mapping through metadata marking; and configuring an automatic cooperation engine to monitor changes in real time, and triggering the intelligent semantic processing module to perform cross analysis: analyzing task constraint conditions to dynamically generate verification logic in the first storage space, or analyzing code change characteristics to automatically generate structured knowledge objects in the second storage space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer data processing technology, and in particular to a technology for performing full lifecycle data association, state synchronization and consistency verification for heterogeneous data storage sources in a distributed software development environment. Background Technology

[0002] As software development grows in scale and complexity, modern software engineering places higher demands on collaboration efficiency and quality management. Existing software development processes typically involve two core management dimensions: first, version control and building of source code, usually hosted on code hosting platforms (such as GitHub, GitLab, etc.); and second, management of task requirements, design documents, and project schedules, usually relying on project management tools or used as document file storage.

[0003] Currently, most mainstream development models adopt a "single-repository strategy" or a loosely coupled architecture between code and documentation. In the single-repository model, source code, build scripts, requirements documents, and various evaluation data are often stored together in the same version control repository. While this approach centralizes resources to some extent, it has many drawbacks in practical applications.

[0004] First, task management and code development are too tightly coupled. Because non-technical requirements documents are mixed with highly technical source code, it's difficult for non-technical members such as product managers and testers to directly participate in the task flow and confirmation, increasing the collaboration barrier. At the same time, the massive amount of code files often obscures crucial task context, making task closure difficult to audit.

[0005] Secondly, there is a "semantic gap" between requirements documents and code implementation. In existing technologies, requirements documents serve only as static references for human readers. For example, the document might specify that "the application programming interface (API) response time must be less than 200ms," but this constraint cannot be directly understood and executed by the build system. This leads to a "document decay" phenomenon: as the code iterates, the document often lags behind the implementation and fails to serve a true constraint.

[0006] Secondly, there is a lack of automated mechanisms for accumulating experience and reusing knowledge. While existing development processes widely employ continuous integration / continuous deployment to automate builds and releases, these processes primarily focus on the compilation and deployment status of code. High-value knowledge generated during development, such as successful experiences, troubleshooting paths, and architectural decisions, is often scattered in commit logs or fragmented comments in merge requests, lacking systematic collection and structured storage. This results in subsequent projects being unable to quickly reuse existing experience, leading to frequent recurrence of the same pitfalls. Summary of the Invention

[0007] This application provides a method and system for bidirectional mapping and consistency maintenance of code and non-code resources, in order to solve the technical problem of how to establish a mechanism that can effectively decouple tasks and code while maintaining a strong semantic relationship and automatic synchronization between the two.

[0008] This application discloses a method for bidirectional mapping and consistency maintenance of code and non-code resources, including the following steps: Configure a first storage space and a second storage space that are physically independent but logically related; wherein, the first storage space is used to store the source code and build configuration of the software project, and the second storage space is used to store the task description data, constraints and structured knowledge objects of the software project; Generate a globally unique association identifier and synchronously store the association identifier in the relevant task description in the second storage space, and associate it with the corresponding code entity in the first storage space through metadata tagging; Configure an automated collaboration engine to monitor data change events related to the associated identifier in the first storage space or the second storage space in real time. When the data change event is detected, the intelligent semantic processing module is triggered to extract context data from the two storage spaces based on the association identifier, perform cross-analysis, and perform the following operations: parse the task constraints in the second storage space and dynamically generate the corresponding verification logic or test cases in the first storage space; parse the code change features in the first storage space and automatically generate or update the structured knowledge object in the second storage space.

[0009] In a preferred embodiment, the step of parsing the task constraints in the second storage space and dynamically generating corresponding verification logic or test cases in the first storage space further includes: From a task description data in the second storage space, identify and extract acceptance metrics in natural language or pseudocode form; The intelligent semantic processing module is used to compile the acceptance indicators into an automated test script that can be executed in the first storage space; The automated test script is injected into the construction process of the first storage space and executed; if the execution fails, a blocking signal is generated to intercept the code release process, and the failure log is fed back to the second storage space.

[0010] In a preferred embodiment, the step of parsing the code change features in the first storage space and automatically generating or updating the structured knowledge object in the second storage space further includes: Monitor the code commit records of the first storage space; When a detached code change without the associated identifier is detected, or when the complexity of the code change exceeds a preset threshold: the intelligent semantic processing module is used to analyze the syntactic structure differences and runtime metric impact of the code change; a task draft containing a summary of the change intent, a list of affected modules, and suggested acceptance criteria is automatically generated in the second storage space; and in response to a confirmation instruction for the task draft, the task draft is converted into the structured knowledge object, while the associated mapping is completed.

[0011] In a preferred embodiment, the step of associating the corresponding code entity to the first storage space via metadata tagging further includes: During the code editing or submission phase in the first storage space, the source code structure is parsed using abstract syntax tree analysis technology. The associated identifier is injected as metadata into the comments, annotations, or extended attributes of the function, class, or module nodes related to the task to form a code fingerprint; When the source code is renamed, files are moved, or logic is refactored, the code fingerprint is tracked using static analysis tools to maintain a persistent binding between the associated identifier and the corresponding code logic segment.

[0012] In a preferred embodiment, the method further includes: When creating or modifying task description data in the second storage space, the intelligent semantic processing module is used to analyze the semantic intent of the task description and predict the target code files that may be involved in the first storage space. Construct a virtual sandbox to simulate potential code changes that may occur after the current task is implemented, and perform collision analysis with code changes involved in other ongoing tasks in the first storage space; Calculate the probability of potential logical conflicts and output warning information in the second storage space before actual code writing.

[0013] In a preferred embodiment, the structured knowledge object is configured as a dynamic entity with lifecycle attributes, and the method further includes: Record the frequency of citations, creation time, and activity level of associated code for the structured knowledge objects; The vitality index of the structured knowledge object is calculated according to a preset algorithm; When the vitality index is lower than a preset threshold, an archiving or deweighting operation is automatically performed; When the vitality index is higher than a preset threshold and the content complexity increases, the intelligent semantic processing module automatically proposes to split the structured knowledge object into multiple sub-objects or extract a common pattern to the organizational knowledge base.

[0014] In a preferred embodiment, the intelligent semantic processing module employs a multi-agent collaborative mechanism: For the same associated identifier, multiple virtual agent roles with different functions are instantiated in the system background. The roles include at least an agent responsible for architecture dependency analysis, an agent responsible for code implementation strategy, and an agent responsible for metric measurement. Multiple virtual agent roles independently evaluate the same data change event based on their respective preset prompts and concerns, and generate comprehensive decision recommendations through preset negotiation protocols.

[0015] In a preferred embodiment, the method further introduces a third storage space as an intent verification layer: Using the intelligent semantic processing module, the expected code skeleton is generated in reverse based solely on the task description data in the second storage space and stored in the third storage space; Perform a semantic comparison between the actual submitted code in the first storage space and the expected code skeleton in the third storage space; The deviation between the intended and the actual result is calculated, and the deviation data is written back to the second storage space as a quality metric.

[0016] In a preferred embodiment, the association identifier includes a hierarchical coding structure, which includes at least an item pedigree code, a generational evolution code, and a mutation point code, to support cross-item task gene retrieval and success rate prediction.

[0017] This application also discloses a bidirectional mapping and consistency maintenance system for code and non-code resources, including: The storage module is configured to maintain physical isolation between the first and second storage spaces. The identifier management module is configured to generate unique associated identifiers and maintain their mapping relationship between the two storage spaces; The automated collaboration engine is configured to listen for data change events in any storage space; The intelligent semantic processing module includes a processor and a memory storing computer instructions, which, when executed, are used to implement the method described above.

[0018] In the implementation of this application, by configuring an independent dual-warehouse architecture for tasks and code, establishing a cross-warehouse metadata mapping mechanism based on globally unique associated identifiers, and introducing an intelligent semantic processing module to respond to and cross-analyze bidirectional change events in real time, it can effectively achieve deep decoupling and dynamic synchronization between business requirements and technical implementation during software development. Specifically, this solution breaks down the barriers of non-technical personnel's difficulty in participating and the difficulty in accumulating knowledge assets in the traditional single-warehouse model, and uses an automated collaboration engine to build a bridge connecting "document constraints" and "code entities." On the one hand, by parsing task constraints to dynamically generate verification logic, the mandatory guiding role of requirement documents for code development is established, preventing the disconnect between requirements and implementation; on the other hand, by parsing code changes to automatically feed back to knowledge objects, implicit development experience is made explicit and structurally accumulated, thereby significantly improving the traceability of the entire software lifecycle, closed-loop management efficiency, and organizational-level knowledge reuse capabilities.

[0019] Furthermore, by intelligently extracting acceptance metrics in natural language form from task descriptions and compiling them into executable automated test scripts for the build process, static document constraints can be transformed into dynamic "executable contracts." This allows for the immediate interception of changes that do not meet business expectations during the code build phase, ensuring that the final deliverables strictly adhere to established quality and functional standards and fundamentally eliminating the phenomenon of "documents being used only as a reference."

[0020] Furthermore, by monitoring changes to detached code or highly complex changes that do not carry identifiers, and using AI to reverse analyze their syntax structure and runtime impact to generate task drafts and completion maps, a reverse closed-loop management mechanism of "deriving causes from results" can be established. This effectively captures process data in emergency repairs or temporary changes, fills the management blind spots caused by the "code first, task later" approach in traditional processes, and ensures the integrity of the audit chain and the absence of knowledge assets.

[0021] Furthermore, by utilizing Abstract Syntax Tree (AST) analysis technology to inject associated identifiers as metadata into the microstructure of code entities (such as functions, class annotations, or extended attributes), and combining this with static analysis tools to trace code fingerprints, it is possible to maintain a persistent and precise binding between task IDs and code logic fragments even in complex scenarios such as code renaming, file moving, or large-scale refactoring. This solves the fragility problem of traditional file path or commit record-based associations and enables fine-grained code tracing.

[0022] Furthermore, by predicting the scope of code involved based on semantic intent during the task creation phase and simulating code changes in a virtual sandbox to conduct collision analysis between multiple tasks, potential logical conflicts and integration risks can be predicted before the actual code is written. This gives developers a "God's-eye view" of foresight, significantly reducing the probability of code merging conflicts and rework costs in the later stages, and achieving a "leftward shift" in development risk management.

[0023] Furthermore, by endowing structured knowledge objects with lifecycle attributes such as vitality index and self-mutation record, and automatically performing archiving, splitting, or pattern extraction operations based on activity and content complexity, a dynamic knowledge base with "metabolism" capability can be constructed. This ensures that the accumulated experience cards always maintain timeliness and high value, solving the industry pain point of traditional knowledge bases becoming rigid and outdated due to lack of maintenance.

[0024] Furthermore, by introducing a multi-agent collaboration mechanism, virtual agent roles that focus on architecture, implementation, and metrics are instantiated in the system backend and negotiation agreements are established. This can upgrade the general assistance of a single model to simulate the decision-making of a "virtual committee" of an expert team. Through multi-dimensional adversarial and consensus mechanisms, the professionalism, comprehensiveness, and feasibility of AI-generated suggestions can be significantly improved.

[0025] Furthermore, by introducing a third storage space as an intent verification layer, the expected code skeleton is generated in reverse based on the task description and semantically compared with the actual submitted code. This allows for the quantitative calculation of the "intent-implementation deviation" metric, thereby effectively identifying code implementations that, although passing the test, deviate from the original design intent. This provides a new semantic consistency dimension for code quality assessment.

[0026] Furthermore, by adopting a hierarchical coding structure of "task DNA" that includes information on project lineage, generational evolution, and mutation points, it is possible to achieve gene-level tracking of task evolution paths, thereby supporting success pattern recognition and risk prediction across projects and cycles. This enables organizations to extract patterns from the "genetic information" of historical tasks to guide future project decisions and resource allocation.

[0027] The various technical features disclosed in the above-described invention, the various technical features disclosed in the following embodiments and examples, and the various technical features disclosed in the accompanying drawings can be freely combined to form various new technical solutions (all of which should be considered as having been recorded in this specification), unless such a combination of technical features is technically infeasible. For example, in one example, feature A+B+C is disclosed, and in another example, feature A+B+D+E is disclosed. Features C and D are equivalent technical means that serve the same function, and technically only one needs to be used; it is impossible to use both simultaneously. Feature E can be technically combined with feature C. Therefore, the solution A+B+C+D should not be considered as having been recorded because it is technically infeasible, while the solution A+B+C+E should be considered as having been recorded. Attached Figure Description

[0028] Figure 1 This is an overall architecture diagram of a bidirectional mapping and consistency maintenance system for code and non-code resources according to an embodiment of this application; Figure 2 This is a schematic diagram of a multi-agent collaborative process according to an embodiment of this application; Figure 3 This is a schematic diagram of the ET card browser plugin interface according to an embodiment of this application. Detailed Implementation

[0029] In the following description, many technical details are presented to help the reader better understand this application. However, those skilled in the art will understand that the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments.

[0030] Example 1: Basic Implementation of Dual-Card Architecture and Identifier System This embodiment provides a basic implementation scheme for a bidirectional mapping and consistency maintenance method between code and non-code resources. In this embodiment, both the first and second storage spaces are configured based on the GitHub platform, each hosting different types of project resources. Specifically, the first storage space is configured as a development repository named EdgeTeam, with a directory structure including the source code directory src / , the build artifact directory dist / , and the automated workflow configuration directory .github / workflows / . The second storage space is configured as a task repository named EdgeTeam-TaR, with a directory structure including the task definition directory tasks / , the structured knowledge object storage directory cards / , and the context directory context / . The two repositories are completely independent at the physical level and are managed uniformly through GitHub organizational accounts, but are closely linked at the logical level through association identifiers. Figure 1It demonstrates a typical system architecture, including a dual-compartment (optional triple-compartment) structure and the connection relationships between the modules.

[0031] The association identifier is designed using a hierarchical coding structure to support cross-project task gene retrieval and success rate prediction. In this embodiment, the encoding format of the association identifier is defined as ET-{project pedigree code}-{generational evolution code}-{mutation point code}-{phenotypic code}. The project pedigree code uses Greek letters, such as Alpha, Beta, Gamma, etc., to identify the project product line to which the task belongs; the generational evolution code uses a format of G plus a number, such as G1, G2, G3, indicating the number of evolutionary layers the task has undergone from its original requirements; the mutation point code uses a format of M plus a version number, such as M1.0, M2.1, recording major changes and sub-changes that occur during the task's execution; the phenotypic code describes the current state of the task, with possible values ​​including Draft, Active, Success, Archive, etc. For example, ET-Alpha-G3-M2.1-Success indicates that this identifier represents a third-generation derivative task belonging to the Alpha product line, achieving a successful state after the first sub-change of the second major change.

[0032] To achieve persistent binding between associated identifiers and code entities, this embodiment employs abstract syntax tree (AST) analysis technology to implant metadata tags at the code level. Specifically, during code editing or submission, the system performs AST parsing on the source code by calling the TypeScriptCompiler API or Babel Parser, identifying function declaration nodes, class declaration nodes, and module export nodes. Subsequently, the system injects the associated identifier as a JSDoc comment at the beginning of the relevant nodes, forming a code fingerprint. Taking JavaScript code as an example, the injected code structure is as follows: the system adds a JSDoc block containing the @taskId tag before the function definition, with the content ET-Alpha-G3-M2.1, thereby establishing a mapping relationship between the function and the corresponding task.

[0033] When source code is renamed, files are moved, or logic is refactored, the system traces the migration path of the code fingerprint using static analysis tools. This embodiment employs a node matching algorithm based on semantic similarity. This algorithm first extracts the abstract syntax tree of the code before and after refactoring, then calculates the feature vector of each node, including dimensions such as node type, parameter signature, and function body hash value. Finally, it uses cosine similarity matching to find the corresponding nodes before and after refactoring. When the similarity exceeds a preset threshold (0.85 in this embodiment), the system automatically migrates the associated identifier of the original node to the new node, thereby maintaining a persistent binding relationship between the code fingerprint and the task.

[0034] The automated collaboration engine is configured with an event-driven architecture based on GitHub Actions. The engine deploys listeners in two repositories to monitor data change events related to associated identifiers in real time. In the first repository, the listened-to events include push (code push), pull request (pull_request), and release (version release); in the second repository, the listened-to events include task file change, status label change, and card creation. When any event is triggered, the engine uses a webhook mechanism to pass the event payload to the intelligent semantic processing module for further analysis.

[0035] Example 2: Automatic Conversion of Constraints to Executable Tests This embodiment, based on Embodiment 1, describes in detail the technical implementation of automatically generating executable verification logic for the first storage space from task constraints in the second storage space. This embodiment upgrades the metric cards in the structured knowledge object into executable contract tests, enabling document constraints to directly control the code construction process.

[0036] In the Markdown file of the metric card in the second storage space, acceptance metrics are embedded using a specific pseudocode format. This embodiment defines a set of acceptance metric description languages, with the basic syntax being Expect: {metric name} {comparison operator} {threshold} {unit}. For example, for API response time constraints, it is recorded in the metric card as Expect: API_Response_Time < 200 ms; for test coverage constraints, it is recorded as Expect: Code_Coverage >= 80%; and for concurrent user support constraints, it is recorded as Expect: Concurrent_Users >= 1000 users. These constraints are embedded in the Markdown document in the form of code blocks, facilitating recognition and extraction by the intelligent semantic processing module.

[0037] The intelligent semantic processing module uses a large language model to compile and convert acceptance metrics into automated test scripts. In this embodiment, the system first extracts all constraints in the metric card that conform to the Expect syntax using regular expression matching. Then, it assembles the extracted results along with the project's technology stack information (such as using Jest as the testing framework and k6 as the performance testing tool) into prompt words and calls the DeepSeek inference model to generate the corresponding test script. Taking API response time constraints as an example, the k6 performance test script generated by the model includes the following logic: defining a virtual user scenario, sending an HTTP request to the target API endpoint, recording the response time, and verifying whether the response time is less than 200 milliseconds using a check function. The generated test script is stored in the test / contracts / directory of the first storage space, and the file name follows the format {association identifier}_contract.test.js.

[0038] Automated test scripts are injected into the build process of the first storage space via GitHub Actions workflows. The workflow configuration file defines the contract testing phase, which is executed after unit tests pass and before deployment. Specifically, the workflow first pulls the metric card file corresponding to the identifier associated with the current branch from the second storage space, then triggers the build script to generate the latest test cases, and finally executes the tests and collects the results. If the test execution fails, i.e., any acceptance metric fails to meet the standard, the workflow will generate a blocking signal, terminating subsequent build and deployment steps, and simultaneously adding a failure tag to the corresponding pull request via the GitHub API.

[0039] The failure log feedback mechanism ensures a closed information loop between the two storage units. When a contract test fails, the system automatically generates a failure report file in the corresponding task directory of the second storage space. The report includes the name of the failed acceptance metric, the expected value, the actual value, the failure timestamp, and the hash value of the code commit that triggered the failure. In addition, the system updates the status field of the metrics card from "Monitoring" to "Violated" and displays it in red on the task dashboard. Project managers and developers can directly view the failure details through the ET Card browser plugin and quickly jump to the relevant code for repair.

[0040] Example 3: Reverse Task Generation Based on Code Changes This embodiment describes a technical solution for generating task descriptions for a second storage space from code changes in a first storage space, addressing the problem of code creation preceding task creation in actual development. Traditional development processes assume tasks always precede code creation; however, in practice, numerous temporary fixes, technical research, and emergency changes occur, often lacking corresponding task records, leading to blind spots in project management.

[0041] The automated collaboration engine continuously monitors code commit records in the first storage space, identifying detached code changes that lack associated identifiers. The identification algorithm is based on the following rules: First, it parses the commit message of each code commit, checking for associated identifiers conforming to the ET-xxx format; second, it scans the code fingerprint comments in the change files to confirm the presence of the @taskId tag; finally, it checks whether the naming of the branch to which the commit belongs conforms to the naming convention feat / ET-xxx or fix / ET-xxx. When all three checks result in a negative outcome, the commit is marked as a detached code change.

[0042] In addition to detecting missing identifiers, the system also assesses whether the complexity of the code changes exceeds a preset threshold. This embodiment defines three dimensions for the complexity threshold: more than 50 lines of code modified, more than 5 files involved, or a cyclomatic complexity increment exceeding 10. Complexity calculation is performed using static analysis tools. The system automatically analyzes the changed code after each commit, extracts the aforementioned metrics, and compares them with the thresholds. When any dimension exceeds the threshold, even if the commit includes an associated identifier, the system will mark it as a high-complexity change requiring review. These thresholds can be configured and adjusted according to project size and complexity.

[0043] When detached code changes or highly complex changes are detected, the intelligent semantic processing module is triggered to execute the reverse task generation process. The module first collects complete contextual information about the change, including a list of changed files, the content of the differences, the commit author, the commit time, and relevant runtime metrics (such as changes in build time and test coverage). Then, the module calls a large language model to perform semantic analysis on the context, generating a task draft document. The draft includes a summary of the change intent (inferred by the model based on code differences and commit information), a list of affected modules (derived from static dependency analysis), and suggested acceptance criteria (recommended by the model based on the change type).

[0044] The generated task draft is pushed to the `drafts / ` directory in the second storage space, and a notification is sent to the relevant developers via the ET browser plugin. The notification states: "A code change without a task record has been detected. The change involves {number of files} and {number of lines modified} lines of code. Should we accept it as a formal task ET-XXX?" Developers can preview the task draft through the plugin interface and choose to accept, modify, or reject. If they choose to accept, the system automatically assigns a formal association identifier to the task, moves the draft to the formal task directory, writes back the code commit record in the first storage space, and appends the association identifier comment to the end of the commit message. If they choose to modify, the system opens the draft editing interface for manual adjustments. If they choose to reject, the system marks the change as a reviewed, task-free change and will not send a repeat notification.

[0045] Example 4: Virtual Sandbox and Conflict Prediction Mechanism This embodiment describes a virtual sandbox technology solution for predicting potential conflicts before code writing. In traditional development processes, code conflicts are usually discovered only during the merging phase, at which point the cost of resolving conflicts is already very high. This embodiment achieves proactive risk management in the development process by predicting conflicts during the task drafting phase.

[0046] When a user creates or modifies task description data in the second storage space, the intelligent semantic processing module first performs semantic intent analysis on the task description. The analysis process uses natural language understanding technology to extract key entities and action intents from the task description. For example, for a task described as optimizing the response speed of the user login module, reducing the verification time from the current 500ms to less than 200ms, the module identifies key entities including the user login module, verification time, and the action intent as performance optimization. Subsequently, the module matches the identification results with the code repository index in the first storage space to predict the target code files that the task may involve, such as auth / login.js, services / validator.js, etc.

[0047] Based on the predicted list of target code files, the system constructs a virtual merge sandbox for collision analysis. The sandbox environment is a temporary clone of the main branch in the first storage space, where the system simulates potential code changes that might occur after the current task is implemented. The generation of simulated changes also relies on a large language model, which infers the possible scope and manner of modification based on the task description and the existing code in the target files. The generated simulated changes are stored as patch files but are not actually applied to any branch.

[0048] The collision analysis algorithm cross-compares the simulated changes of the current task with the code changes involved in other ongoing tasks in the first storage space. The algorithm calculates the conflict probability from the following dimensions: file overlap (the ratio of the intersection to the union of the target file sets involved in the two tasks); function overlap (the proportion of the intersection of the function sets that the two tasks may modify); and logical dependency (the degree of overlap of the indirect influence scope determined by call chain analysis). The final conflict probability is calculated by a weighted average of the three dimensions, with weights set to 0.4, 0.4, and 0.2 in this embodiment.

[0049] The system adopts different response strategies based on different conflict probability ranges. When the conflict probability is below 60%, the system only displays a low-risk warning on the task details page and does not take mandatory intervention. When the conflict probability is between 60% and 80%, the system outputs an orange warning message in the task description in the second storage space, with the following format: Note: Your task requirements may have a logical conflict of {probability value} with task number {associated identifier}. It is recommended to communicate with {responsible person's name} before development. When the conflict probability exceeds 80%, the system outputs a red mandatory warning and sends a notification to the responsible persons of both parties via instant messaging, requiring coordination to be completed before code writing. The warning message also includes a list of conflicting files and a list of conflicting functions, making it easier for developers to quickly locate potential problem areas.

[0050] Example 5: Lifecycle Management of Structured Knowledge Objects This embodiment describes a technical solution for dynamic lifecycle management of structured knowledge objects. In this application, structured knowledge objects include three types: success cards, component cards, and metric cards. They are configured as dynamic entities with lifecycle attributes, capable of automatically evolving, splitting, or archiving according to usage, thereby solving the industry pain point of long-term maintenance of knowledge bases.

[0051] Each structured knowledge object's metadata includes a lifecycle tracking field. The system continuously records the following metrics: citation frequency (the number of times the object has been referenced by other tasks, code comments, or documentation in the past 90 days); creation time (the timestamp of the object's first creation); last modification time (the timestamp of the object's most recent update); and associated code activity (the commit frequency of the code files associated with the object via an association identifier in the past 30 days). These metrics are updated daily via a background scheduled task.

[0052] The vitality index is calculated using the formula V = α × R × T × A, where R represents the standardized citation frequency, T represents the time-related decay factor, A represents the associated code activity factor, and α is the normalization coefficient. The time-related decay factor uses an exponential decay model, calculated as T = e (-λ×d)Where d represents the number of days since creation, and λ is the decay constant, set to 0.002 in this embodiment, corresponding to a lifespan decay of approximately 50% over about one year. The associated code activity factor is calculated based on the commit frequency: A is 1.0 when the associated code has been committed in the past 30 days, 0.7 when there have been no commits but there have been commits in the past 90 days, and 0.3 when there have been no commits for more than 90 days. The final lifespan index V is normalized to the range of 0 to 1.

[0053] The system triggers different automated operations based on a vitality index threshold. When the vitality index falls below 0.1, the system performs archiving or deweighting operations: moving the structured knowledge object from the active directory cards / to the archive directory archive / , reducing its weight in the organization's knowledge base search results, and adding an archived tag and a reason for archiving to the object's metadata. Before archiving, the system extracts knowledge from the object's content using a large language model, generating a final summary that records patterns and ideas with transferable value for reference in subsequent projects.

[0054] When the vitality index is higher than 0.8 and the content complexity exceeds a preset threshold, the system triggers an automatic splitting proposal. The complexity threshold is calculated based on the number of characters, chapters, and embedded elements of the object. For highly active and complex objects that meet the criteria, the intelligent semantic processing module analyzes their content structure, identifies knowledge units that can stand alone, and generates splitting suggestions. For example, a success card containing three parts—architectural design, implementation details, and performance optimization—may be suggested to be split into three independent objects: an architecture success card, an implementation component card, and a performance measurement card. The splitting suggestion is presented to the object maintainer in a pop-up window for confirmation and execution. In addition, when the system detects multiple similar structured knowledge objects (similarity is calculated using the cosine distance of semantic vectors, with a threshold set to 0.85), it will also propose merging them into one object while retaining their respective best practice content.

[0055] Example 6: Multi-Agent Collaborative Decision-Making Mechanism This embodiment describes the technical implementation of the multi-agent collaboration mechanism in the intelligent semantic processing module. Unlike traditional single large-model solutions, this embodiment breaks down the intelligent semantic processing capability into multiple virtual agent roles with different functions, forming more comprehensive and professional decision-making suggestions through negotiation between the agents.

[0056] For a task corresponding to the same associated identifier, the system instantiates four virtual agent roles in the background. The first is the architect agent, whose function is to analyze the impact of the task on the overall system architecture, focusing on module dependencies, interface compatibility, and technical debt risks. The second is the implementation agent, whose function is to evaluate the specific implementation strategy of the task, focusing on code complexity, implementation time estimation, and testability. The third is the metrics agent, whose function is to monitor task-related metrics, focusing on performance benchmarks, test coverage, and acceptance criteria achievement. The fourth is the retrospective agent, whose function is to extract valuable lessons from the task, focusing on success factors, lessons learned, and reusable patterns.

[0057] Each agent role has its own unique prompt template and output format specification. Taking the architect agent as an example, its prompt includes the following core instructions: As a system architect, please analyze the impact of this task on the existing architecture, focusing on evaluating the following aspects: whether it introduces new external dependencies, whether it changes the responsibility boundaries of existing modules, whether there is a risk of circular dependencies, and whether it conforms to the current technology selection strategy. Please output the evaluation results in JSON format, including three fields: impact_level (high / medium / low), risk_points (a list of risk points), and recommendations (a list of recommendations). Other agents have similar prompt structures, but their focus dimensions and output fields differ.

[0058] At key decision-making junctures (such as task initiation approval, code merge review, and version release review), the system triggers a proxy negotiation process. The process consists of three phases (see...). Figure 2 In the independent evaluation phase, each agent generates its own evaluation result based on the current data change event and the dual-warehouse context, according to its own prompts. In the cross-review phase, the evaluation results from each agent are aggregated and redistributed to all agents, and each agent can provide supplementary opinions or objections to the evaluation results of other agents. In the consensus formation phase, the system uses a weighted voting protocol to synthesize the opinions of all agents and generate the final decision recommendation. In the weighted voting, the weight of each agent is dynamically adjusted according to the decision type. For example, in technical architecture-related decisions, the architect agent has a weight of 0.4, and in schedule management-related decisions, the indicator officer agent has a weight of 0.4.

[0059] For situations where there are significant disagreements between agents, the system employs a conflict resolution mechanism. When the evaluation conclusions of any two agents are logically contradictory (e.g., the architect agent believes the solution is feasible while the implementation agent believes it is not), the system extracts the points of disagreement and presents them to the human decision-maker, along with the arguments of each agent. After the human decision-maker makes a final judgment, the system writes the decision result and its rationale into the decision log of the corresponding task, serving as material for the agent model's subsequent learning and optimization. Through this human-machine collaborative approach, the agent's judgment capabilities can continuously evolve as the project progresses.

[0060] Example 7: Third Storage Space and Intent Verification Layer This embodiment describes a technical solution that introduces a third storage space as an intent verification layer. This solution is used to address the implicit problem that code meets the requirements but deviates from the original intention. By reverse-engineering the expected code skeleton and comparing it with the actual code, the degree of deviation between the intent and the implementation is quantified.

[0061] The third storage space is configured as a shadow repository named EdgeTeam-Shadow. Its directory structure is consistent with the first storage space, but it only stores the code skeleton files reverse-engineered by the intelligent semantic processing module, rather than the actual executable source code. The creation and updating of the shadow repository are fully automated and do not require direct intervention from developers. Whenever the task description in the second storage space changes, the system triggers a synchronization update process for the shadow repository.

[0062] The reverse engineering process for generating a code skeleton is as follows: First, the intelligent semantic processing module reads task description data from the second storage space, including functional requirements, acceptance criteria, and technical constraints. Second, the module combines the existing code structure in the first storage space (such as directory layout, naming conventions, and dependency graphs) and calls the large language model to generate the code skeleton that the task should theoretically produce. Finally, the generated code skeleton is stored in the corresponding location in the third storage space. A code skeleton is an abstract representation that includes the signature definitions of classes and functions, key data structures, and pseudocode descriptions of the core algorithm flow, but does not contain specific implementation details.

[0063] When a code submission related to a task is generated in the first storage space, the system automatically performs intent-implementation comparison analysis. The comparison process employs a multi-level semantic matching algorithm: the structural layer comparison checks the correspondence between the actual code and the expected skeleton in terms of class, function, and module organization; the interface layer comparison verifies the consistency of function signatures, parameter types, and return types; and the logic layer comparison evaluates the similarity of the algorithm flow through abstract syntax tree analysis. Each level outputs a matching score between 0 and 1. The final deviation is calculated as 1 minus the weighted average of the matching scores from the three levels. In this embodiment, the weights are set to 0.3, 0.3, and 0.4, respectively.

[0064] Deviation data is written back to the secondary storage space as a quality metric. The system adopts different processing strategies based on different deviation ranges: when the deviation is below 30%, it is considered normal deviation and recorded only as statistical data; when the deviation is between 30% and 50%, the system adds a review suggestion label to the task's metric card, prompting code reviewers to focus on the consistency between the implementation and the requirements; when the deviation exceeds 50%, the system triggers an early warning mechanism, sending a notification to the task owner, requesting confirmation of whether the actual implementation matches the original intent, or whether the task description needs to be updated to reflect design changes during the implementation process. In addition, the system will track the average deviation trend of each project over a long period, providing data support for continuous improvement of requirements writing quality.

[0065] Example 8: System Device Implementation This embodiment provides a bidirectional mapping and consistency maintenance system for code and non-code resources, the structure of which is as follows: Figure 1 As shown, this system is used to implement the methods described in Embodiments 1 to 7 above. The system comprises four core components: a storage module, an identifier management module, an automated collaboration engine, and an intelligent semantic processing module.

[0066] The storage module is configured to maintain physically isolated first storage space, second storage space, and optional third storage space. In this embodiment, the storage module is implemented based on the repository service of the GitHub platform, with the three storage spaces corresponding to three independent Git repositories. The storage module also includes a distributed caching layer deployed using a Redis cluster to cache frequently accessed task metadata, code index information, and associated mapping tables to improve system response performance. Data synchronization between the caching layer and the GitHub repository is achieved via Webhooks. When the repository content changes, the relevant cached items are marked as invalid and reloaded on the next access.

[0067] The identifier management module is configured to generate unique associated identifiers and maintain their mapping relationships across various storage spaces. This module includes two sub-components: an identifier generator and a mapping registry. The identifier generator generates new associated identifiers according to the hierarchical encoding rules described in Example 1, ensuring global uniqueness of the identifiers through a distributed lock mechanism during the generation process. The mapping registry is implemented based on a relational database and stores the many-to-many mapping relationships between associated identifiers and code entities in the first storage space, task files in the second storage space, and skeleton files in the third storage space. The registry also records the timestamp and operator information for each mapping change, supporting historical tracing of mapping relationships.

[0068] The automated collaboration engine is configured to listen for data change events in any storage space and trigger corresponding processing flows. The engine adopts an event-driven architecture, with core components including an event collector, an event router, and a task scheduler. The event collector subscribes to change notifications from various repositories via the GitHub Webhook interface, standardizes the received event payload, and pushes it to a message queue. The event router consumes events from the message queue and distributes them to the corresponding processing pipelines based on the event type and associated identifier. The task scheduler manages the execution of each processing pipeline, supporting task priority settings, concurrency control, and failure retry mechanisms. In this embodiment, the message queue is deployed using Apache Kafka, supporting a processing throughput of thousands of events per second.

[0069] The intelligent semantic processing module includes a processor and a memory storing computer instructions. The processor is deployed using a GPU server cluster; in this embodiment, it is configured with four computing nodes equipped with high-end GPUs to run large language model inference tasks. The computer instructions stored in the memory include model loaders, prompt word templates, multi-agent negotiation protocols, and various semantic analysis algorithms. The module provides a RESTful API interface, supporting functions such as acceptance metric compilation, reverse task generation, conflict prediction, vitality index calculation, agent negotiation, code skeleton generation, and deviation analysis. For computationally intensive tasks, the module adopts an asynchronous processing mode; the caller obtains the task ID after submitting a request, and subsequently retrieves the processing results through polling or callback methods.

[0070] The system also includes a unified entry module, implemented as an ET card in the form of a browser plugin. The ET card, by calling the API interfaces of the aforementioned modules, displays the version release status of the first storage space, the task progress dashboard of the second storage space, the mapping relationship diagram of associated identifiers, and the analysis results of the intelligent semantic processing module in a unified interface. The ET card supports deep integration with GitHub pages; when users browse code files, they can directly view the associated task descriptions, and when users browse task files, they can directly jump to the associated code location. Furthermore, the ET card provides quick operation entry points, such as one-click task creation, one-click test case generation, and one-click proxy negotiation triggering, lowering the operational threshold for users to use the system. Figure 3 This shows a schematic diagram of the interface of an ET card browser plugin.

[0071] Accordingly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the various method embodiments of this application. Computer-readable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device. As defined herein, computer-readable storage media do not include transient computer-readable media, such as modulated data signals and carrier waves.

[0072] Furthermore, embodiments of this application also provide a computer program product, including computer-executable instructions that, when executed by a processor, implement the steps in the above-described method embodiments.

[0073] It should be noted that in this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. In this application, if it refers to performing an action according to an element, it means performing the action at least according to that element, including two cases: performing the action only according to that element, and performing the action according to that element and other elements. Expressions such as "multiple," "repeatedly," and "various" include two, two times, two kinds, and more than two, more than two times, and more than two kinds.

[0074] This specification includes combinations of various embodiments described herein. Individual references to embodiments are made (e.g., "one embodiment," "some embodiments," or "preferred embodiments"); however, these embodiments are not mutually exclusive unless indicated to be mutually exclusive or are readily apparent to those skilled in the art. It should be noted that the word "or" is used in a non-exclusive sense throughout this specification unless the context explicitly indicates or requires it.

[0075] All references to this specification are considered to be incorporated integrally into the disclosure of this application so that they can serve as the basis for modifications if necessary. Furthermore, it should be understood that the above descriptions are merely preferred embodiments of this specification and are not intended to limit the scope of protection of this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of one or more embodiments of this specification should be included within the scope of protection of one or more embodiments of this specification.

[0076] In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A code and non-code resource bidirectional mapping and consistency maintenance method, characterized in that, The method comprises the following steps: configuring a first storage space and a second storage space which are physically independent and logically associated; wherein the first storage space is used for storing source code and build configuration of a software project, and the second storage space is used for storing task description data, constraint conditions and structured knowledge objects of the software project; generating a globally unique association identifier, and synchronously storing the association identifier in the related task description of the second storage space, and associating the association identifier to the corresponding code entity of the first storage space by means of metadata marking; configuring an automated collaboration engine, and monitoring data change events related to the association identifier in the first storage space or the second storage space in real time; when the data change event is monitored, triggering an intelligent semantic processing module, extracting context data in the two storage spaces based on the association identifier for cross analysis, and performing the following operations: analyzing task constraint conditions in the second storage space, and dynamically generating corresponding verification logic or test cases in the first storage space; analyzing code change characteristics in the first storage space, and automatically generating or updating the structured knowledge objects in the second storage space.

2. The method of claim 1, wherein, The operation of analyzing task constraint conditions in the second storage space and dynamically generating corresponding verification logic or test cases in the first storage space further comprises: identifying and extracting acceptance indicators in the form of natural language or pseudo code from a piece of task description data in the second storage space; compiling the acceptance indicators into an automated test script executable in the first storage space by using the intelligent semantic processing module; injecting the automated test script into the build process of the first storage space and executing it; if the execution fails, generating a blocking signal to intercept the code release process, and feeding back the failure log to the second storage space.

3. The method of claim 1, wherein, The operation of analyzing code change characteristics in the first storage space and automatically generating or updating the structured knowledge objects in the second storage space further comprises: monitoring code submission records of the first storage space; when detecting free code changes not carrying the association identifier, or detecting that the complexity of code changes exceeds a preset threshold: analyzing the syntax structure difference and runtime indicator influence of the code changes by using the intelligent semantic processing module; automatically generating a task draft containing change intention abstract, affected module list and recommended acceptance standard in the second storage space, and in response to a confirmation instruction for the task draft, converting the task draft into the structured knowledge object while completing the association mapping.

4. The method of claim 1, wherein, The operation of associating to the corresponding code entity of the first storage space by means of metadata marking further comprises: analyzing the source code structure by using abstract syntax tree analysis technology during the code editing or submission stage of the first storage space; injecting the association identifier as metadata into the annotation, note or extension attribute of the function, class or module node related to the task, to form a code fingerprint; The association identifier is kept persistently bound to the corresponding code logic segment by tracking the code fingerprint via static analysis tools when the source code undergoes renaming, file moving or logic refactoring.

5. The method of claim 1, wherein, The method further comprises: When the task description data in the second storage space is created or modified, the semantic intention of the task description is analyzed by the intelligent semantic processing module, and the target code file that may be involved in the first storage space is predicted; A virtual sandbox is constructed to simulate the code changes that may be generated after the current task is implemented, and a collision analysis is performed with the code changes involved in other ongoing tasks in the first storage space; The potential logic conflict probability is calculated, and a warning information is output in the second storage space before the actual code is written.

6. The method of claim 1, wherein, The structured knowledge object is configured as a dynamic entity with a life cycle attribute, and the method further comprises: The reference frequency, creation time and activity of the associated code of the structured knowledge object are recorded; The vitality index of the structured knowledge object is calculated according to a preset algorithm; When the vitality index is lower than a preset threshold, an archiving or de-authorization operation is automatically performed; When the vitality index is higher than a preset threshold and the content complexity increases, the structured knowledge object is automatically proposed to be split into multiple sub-objects or the general pattern is extracted to the organizational level knowledge base by using the intelligent semantic processing module.

7. The method of claim 1, wherein, The intelligent semantic processing module adopts a multi-agent collaborative mechanism: for the same association identifier, multiple virtual agent roles with different functions are instantiated in the background of the system, including at least an agent responsible for architecture dependency analysis, an agent responsible for code implementation strategy and an agent responsible for index measurement; the multiple virtual agent roles with different functions independently evaluate the same data change event based on their respective preset prompt words and focus points, and generate a comprehensive decision suggestion through a preset negotiation protocol.

8. The method of claim 1, wherein, The method further introduces a third storage space as an intention verification layer: Using the intelligent semantic processing module, only based on the task description data in the second storage space, the expected code skeleton is reversely generated and stored in the third storage space; The actual submitted code in the first storage space is compared with the expected code skeleton in the third storage space; The deviation degree of intention and implementation is calculated, and the deviation degree data is written back to the second storage space as a quality measurement index.

9. The method of claim 1, wherein, The association identifier contains a hierarchical coding structure, which at least includes a project lineage code, a generational evolution code and a mutation point code, for supporting cross-project task gene retrieval and success rate prediction.

10. A code and non-code resource bidirectional mapping and consistency maintenance system, characterized in that, The method comprises: A storage module configured to maintain a physically isolated first storage space and a second storage space; An identifier management module configured to generate a unique association identifier and maintain its mapping relationship in the two storage spaces; An automated collaboration engine configured to listen to data change events in any storage space; An intelligent semantic processing module including a processor and a memory storing computer instructions, when the instructions are executed, for implementing the method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Data change response method and device

    CN113918427A

  • Specific project software code knowledge management platform and construction method thereof

    CN113986340A

  • Business data docking method and device, electronic equipment and storage medium

    CN117009363A

  • Code and development document automatic synchronization method based on large model

    CN120560715A

  • Research and development test bidirectional synchronization method based on artificial intelligence

    CN120994571A