Zero code visualization and automatic routing method for large model servitization orchestration

By building an entity intent graph to generate state snapshots and combining them with the global session ID, the problems of session state continuity and routing decision deviation in large-model service-oriented orchestration are solved, cross-node state continuity and dynamic adaptation of routing strategies are achieved, service matching is improved, and development difficulty is reduced.

CN120745846AActive Publication Date: 2025-10-03KARAMAY HONGYOU SOFTWARE

Patent Information

Application Number
CN202511247772.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-10-03
Estimated Expiration
2045-09-03

AI Technical Summary

Technical Problem

Existing large-model service-oriented orchestration technology has problems such as poor session state continuity and routing decision deviation, resulting in weak adaptability in complex session scenarios, affecting user experience and restricting large-scale applications.

Method used

By building an entity intent graph to generate a state snapshot, combining it with a globally unique session ID to bind it to distributed storage, cross-node state continuity is established, and by fusing the state snapshot with the new request semantics to calculate the routing weight offset and correct the decision table, zero-code visual configuration is achieved.

Benefits of technology

It solves the problems of session state loss and intent drift, improves the matching degree between service response and actual needs, lowers the development threshold, and is suitable for the agile construction of large-model collaborative applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120745846A_ABST
    Figure CN120745846A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of routing context connection, in particular to a zero-code visualization and automatic routing method for large-model servitization orchestration, which comprises the following steps of: analyzing a user session text, extracting a core entity and an intention dependency relationship, constructing an entity intention graph, compressing to generate a state snapshot, and distributing a global session ID (Identity); the method comprises the steps that a snapshot is bound to a distributed routing decision tree node and stored, cross-node state continuity is established, when a new request is received, the snapshot and new session semantics are fused, a routing weight offset correction decision table is calculated, and a routing decision logic flow is generated by dragging an interface configuration rule and strategy. The method improves the service matching degree, reduces the development threshold, and is suitable for agile construction of large-model collaborative application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of routing context connection technology, and in particular to a zero-code visualization and automatic routing method for large-model service-oriented orchestration. Background Art

[0002] Routing context connection is a crucial technology. As large-scale model applications rapidly gain popularity, enabling multi-model collaboration to respond to complex demands through service-oriented orchestration is crucial for lowering development barriers, improving service deployment efficiency, and enhancing the applicability of large-scale models. This technology transcends the traditional reliance on specialized expertise for code development, providing an intuitive and efficient solution for enterprises to rapidly build large-scale model applications. Traditional model integration methods that rely on manual coding are no longer able to meet the agile development needs of diverse scenarios.

[0003] However, existing large-model service-oriented orchestration technologies suffer from core problems such as poor session state continuity and routing decision deviations. Traditional solutions do not perform structured storage and dynamic tracking of the entity intent of user sessions. When sessions interact across nodes, state loss can easily lead to context breaks, making it impossible to accurately understand user continuity needs. At the same time, there is a lack of a dynamic adjustment mechanism for routing weights based on real-time session semantics. Fixed routing strategies are unable to cope with the drift of user intent, resulting in a decrease in the match between service responses and actual needs. The combination of these problems makes large-model services less adaptable in complex session scenarios, which not only affects user experience but also restricts the large-scale application of large-model service-oriented services, making it difficult to meet enterprises' needs for efficient and accurate model orchestration. To address this technical problem, we provide a zero-code visualization and automatic routing method for large-model service-oriented orchestration. Summary of the Invention

[0004] The purpose of the present invention is to provide a zero-code visualization and automatic routing method for large-scale model service orchestration to solve the problems raised in the above background technology.

[0005] 1. Because traditional solutions do not structure the storage of session entity intents, cross-node interactions are prone to state loss, resulting in context fragmentation. Therefore, this case builds an entity intent graph and generates state snapshots, which are then bound to session IDs and stored in a distributed storage pool. This establishes cross-node state continuity and accurately understands coherence requirements.

[0006] 2. Because traditional solutions lack a real-time semantic routing adjustment mechanism, fixed strategies are difficult to cope with intent drift, resulting in response deviations. Therefore, this case integrates state snapshots with new request semantics, calculates routing weight offsets, and corrects the decision table, eliminating intent drift and improving service matching.

[0007] To achieve the above objectives, one of the objectives of the present invention is to provide a zero-code visualization and automatic routing method for large-scale model service orchestration, comprising the following steps: S1. Parse the text content of the user's current conversation, extract the core entities and the intent dependencies between them, and construct an entity intent graph based on the core entities and the intent dependencies between them. Perform binary compression on the entity intent graph to generate a state snapshot and resolve state loss. S2. Assign a globally unique session ID to the current session, bind the state snapshot to the corresponding node of the distributed routing decision tree through the session ID, and store the bound state snapshot in the distributed session storage pool to establish cross-node state continuity; S3. When a new user session request with the same session ID is received, the corresponding state snapshot is loaded from the distributed session storage pool, and the loaded state snapshot is fused and analyzed with the semantic content of the new user session request to obtain a fused context state. A routing weight offset is calculated based on the fused context state, and the routing decision table is modified according to the routing weight offset to eliminate intent drift. S4. Configure entity extraction rules and state snapshot generation strategies through a drag-and-drop interface to generate routing decision logic flows that integrate contextual states, achieving zero-code development.

[0008] Compared with the prior art, the present invention has the following beneficial effects: 1. By extracting core entities and intent dependencies to build a graph, and combining binary compression to generate a state snapshot containing the topological structure, this approach accurately preserves key session information while significantly reducing storage volume. A globally unique session ID is then bound to the nodes of the distributed routing decision tree, and a copy-on-write mechanism is used to implement multi-node redundant storage, ensuring that state is not lost during cross-node interactions. This completely resolves the contextual disconnection issue inherent in traditional solutions, allowing the system to consistently understand user needs throughout complex conversations and improve interaction fluidity.

[0009] 2. When receiving a new session request, the system loads a historical state snapshot and performs a fourth-order fusion with the new request semantics to generate a complete context state. Based on this, the routing weight offset is calculated to dynamically modify the decision table, enabling the routing strategy to adapt to user intent drift in real time, avoiding response deviations caused by fixed routes, and significantly improving the matching degree between service responses and actual needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is the overall workflow diagram of the present invention. DETAILED DESCRIPTION

[0011] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0012] See also Figure 1 As shown, this embodiment provides a zero-code visualization and automatic routing method for large-scale model service orchestration, including the following steps: S1. Parse the text content of the user's current conversation, extract the core entities and the intent dependencies between them, and build an entity intent graph based on the core entities and the intent dependencies between them. Perform binary compression on the entity intent graph and generate a state snapshot to resolve state loss. Parse the text content of the user's current conversation and extract the core entities and the intent dependencies between entities, including: In the process of large model service orchestration, parsing the user's current session text to extract the core entities and intent dependency relationships is the basis for constructing subsequent routing decisions. This step provides a basis for accurately understanding the user's needs by structuring the session content. The domain key entities are extracted by using the named entity recognition module of the pre-trained large model. Specifically, after splitting the user session text by sentence, it is input into the named entity recognition module loaded with domain-specific dictionaries (such as "product", "order" in the e-commerce domain, "account", "transaction" in the financial domain). The module uses the fine-tuning parameters of the pre-trained model to identify noun phrases that conform to the domain characteristics in the text (such as "smartphone", "transfer amount"), and annotates the entity types (such as "product", "amount"), while filtering out虚词 or general vocabulary without practical significance (such as "of", "then") to ensure that the extracted entities are all core elements strongly related to the current session topic. On this basis, a dynamic dependency network is constructed based on the co-occurrence frequency and semantic similarity between entities. The co-occurrence frequency refers to the number of times two entities appear simultaneously in the same session context window (such as within 3 consecutive sentences). The higher the frequency, the closer the entity association; the semantic similarity is obtained by calculating the cosine distance of the entities in the pre-trained language model word vector space. The smaller the distance, the stronger the semantic association between the entities. Both are used as the quantitative basis for the dependency relationship between entities. The co-occurrence frequency reflects the co-occurrence pattern of entities in the session, and the semantic similarity verifies the rationality of the association from the semantic level. When constructing, first take the extracted entities as nodes, and then assign initial weights to the connections between nodes according to the weighted values of the co-occurrence frequency and semantic similarity to form a preliminary dependency network. Subsequently, the weights are dynamically adjusted according to the real-time session flow (such as when a pair of entities appears again in a new session, the weight is increased by 5%) so that the network can reflect the changes in entity associations in real time. To further optimize the network structure, strong association entity pairs are screened using the session context window. Specifically, taking the current session sentence as the center, a context window including the first 2 sentences and the last 2 sentences is set, and the co-occurrence times and average semantic similarity of entity pairs within the window are statistically calculated. Entity pairs with co-occurrence times ≥ 2 and semantic similarity ≥ 0.7 (full score 1) are determined as strong association entity pairs (such as "product" and "price" frequently co-occur and have a strong semantic association in a shopping session), while isolated noise entities refer to entities that co-occur with other entities 0 times or have semantic similarity < 0.3 in the context window (such as the "weather" accidentally mentioned in the session has no association with other entities in the e-commerce consultation scenario).During screening, strongly associated entity pairs and their connection weights are retained, isolated noise entities are eliminated, and finally a weighted entity-intent association set is output. The weight is calculated comprehensively based on the frequency of entity occurrence in the conversation (the more times it appears, the higher the weight) and position (the weight of entities at the beginning or core position of the sentence is increased by 20%). For example, if "product A" appears 5 times in the conversation and 3 times at the beginning of the sentence, its weight will be higher than "coupon" which only appears 2 times. This provides high-quality input data for the subsequent construction of the entity intent map, so that the entire conversation parsing process is not only in line with the domain characteristics, but also can dynamically adapt to changes in conversation content, laying a solid foundation for achieving accurate service routing.

[0013] Build an entity intent graph based on core entities and the intent dependencies between entities, including: After extracting the core entities and the intent dependencies between entities, the next step is to convert these associations into a structured entity intent graph. This process is both a visual presentation of the previous parsing results and the basis for the subsequent generation of state snapshots. Specifically, the entity-intent association set with weight labels is first mapped into a directed graph structure: each entity is an independent node in the graph, and the node attributes include the entity type and frequency of occurrence. The intent dependencies between entities are converted into weighted directed edges. The direction of the edge is determined by the directionality of the dependency (for example, "product" points to "price" to indicate that the "product" in the user's intent depends on the "price" information). The initial weight of the edge is temporarily replaced by the weight value in the entity-intent association set. In order to make the graph more in line with actual routing needs, the weight of the directed edge needs to be initialized based on the historical routing decision path. The system retrieves historical conversation data from the past 30 days that aligns with the current conversation domain, extracts the routing decision paths for all entity pairs, and counts the number of times each entity pair co-occurs in successful routing cases. If an entity pair (such as "order" and "logistics") is frequently used in routing decisions in historical decisions, with a success rate exceeding 80%, the corresponding directed edge is assigned a higher initial weight (e.g., 0.8). Conversely, if the entity pair is less relevant in historical routing or has a success rate below 30%, the initial weight is reduced (e.g., 0.3). To mitigate the timeliness of historical data, historical paths from the past seven days are weighted 1.2 times as much. This initialization ensures that the weights of directed edges reflect both overall patterns and recent trends, ensuring that they truly reflect the impact of entity dependencies on routing decisions. Based on this, the directed edge weights are dynamically adjusted using real-time conversation flow, and weakly connected subgraphs are pruned and merged. When new conversation content is input, the system will monitor the interaction frequency of entity pairs in the current conversation in real time: if an entity pair is repeatedly mentioned in three consecutive rounds of conversation (such as the user repeatedly asking about "the inventory and distribution of product A"), the weight of the corresponding directed edge will be increased by 10%; if the entity pair does not appear again within five rounds of conversation, the weight will be reduced by 5%. In the adjusted graph, if there is a weakly connected subgraph (i.e., the average weight of all directed edges within the subgraph is less than 0.2, and the weight of the edge connecting the subgraph to the main graph is less than 0.1), the graph is judged to be redundant. For example, if "product reviews" and "after-sales service" are only mentioned once in a shopping session and do not affect routing decisions, the resulting subgraph is weakly connected to the main graph and the subgraph is pruned. If the entity types of two weakly connected subgraphs are highly related (such as the "coupon" subgraph and the "discount" subgraph), they are merged into one subgraph and the edge weights are recalculated to reduce graph redundancy. This graph not only retains the core structure of user intent, but also achieves adaptability to real-time sessions through dynamic weight adjustment. It provides accurate structured data for subsequent compression and generation of state snapshots, enabling the state snapshots to truly reflect the evolution of session intent and lay the foundation for establishing cross-node state continuity.

[0014] Perform binary compression on the entity intent graph to generate a state snapshot, including: After the entity intent graph is constructed, it needs to be binary compressed to generate a state snapshot for easy storage and cross-node transmission. This process uses a layered compression strategy to achieve efficient processing of graph data while ensuring that key information is not lost. First, Huffman encoding is performed on the label data of the graph nodes: First, count the frequency of occurrence of all node labels, assign shorter binary codes to high-frequency labels (such as "product" and "order" entities that appear frequently in the conversation), and assign longer codes to low-frequency labels. For example, "product" appears 50 times and corresponds to the code "01", and "coupon" appears 5 times and corresponds to the code "1011"; all node labels are converted into binary streams according to this encoding rule to form compressed label data blocks. This step can significantly reduce the number of bytes required for label storage while retaining the unique identification information of the node. Differential pulse code modulation is performed on the matrix of directed edge weights. Specifically, the weight matrix is ​​first converted into a one-dimensional sequence in row order (such as the weight value of the 1st row and 3rd column in the matrix follows the 1st row and 2nd column), and then the two adjacent weights in the sequence are calculated. The difference between the values ​​(such as the difference between the current value 0.8 and the previous value 0.6 is 0.2), the difference is used to replace the original weight value to form a differential sequence, and then the differential sequence is pulse-coded, and the numerical range is mapped to the corresponding pulse signal length (such as the difference 0.2 corresponds to 2 pulse units, encoded as "110"). In this way, the continuous weight value is converted into a discrete pulse code, and the data volume is further compressed. It is worth noting that key topological structure identifiers are synchronously retained during the compression process, such as the position index corresponding to the edge that marks the core dependency relationship between entities (directed edge with weight ≥ 0.5), and the boundary information of the connected components of the graph. These identifiers ensure that the connection relationship and intention dependency path of the entity nodes can be accurately restored during decompression to avoid topological structure distortion. Finally, the compressed label data block, weight matrix encoding block, and topology structure identifier are integrated into a binary block, and the compression version number (used to distinguish different compression algorithm versions) and CRC32 checksum (used to verify data integrity) are attached to form a complete state snapshot. This provides a compact and reliable data carrier for subsequently binding the state snapshot to the distributed routing decision tree nodes and establishing cross-node state continuity, making the storage and transmission of state information more adaptable to the resource requirements of the distributed system.

[0015] S2. Assign a globally unique session ID to the current session, bind the state snapshot to the corresponding node of the distributed routing decision tree using the session ID, and store the bound state snapshot in the distributed session storage pool to establish cross-node state continuity. Assign a globally unique session ID to the current session, including: After generating a state snapshot of the entity intent graph, in order to achieve cross-node session state tracking, a globally unique session ID must be assigned to the current session. This identifier will serve as the core identifier for subsequent state binding and storage. Specifically, when generating a composite identifier, first perform a hash operation on the user identity information using the SHA-256 algorithm to obtain a 256-bit user identity hash value, then obtain the geographic coordinates of the service node that currently processes the session and convert it into a 32-bit binary code, while intercepting the milliseconds of the current timestamp and converting it into a 10-bit binary number. These three parts of data are concatenated in the order of "user identity hash value + service node geographic coordinate code + timestamp millisecond code" to form a long string, and then perform an MD5 operation on the string to generate a 128-bit fixed-length composite identifier, namely the session ID. After the session ID is generated, it is injected into the root node of the routing decision tree. Metadata: The system locates the root node of the distributed routing decision tree, writes the session ID in its metadata field, and associates the root node's creation time and expiration time parameters, making the root node the anchor point of the session state lifecycle. When a session starts, the root node associates the state snapshot through the session ID; while the session continues, the state updates of all child nodes are based on the root node's session ID. When the session terminates (such as timeout and inactivity), the session ID in the root node metadata triggers the expiration cleanup mechanism, ensuring that the entire lifecycle of the session state is manageable and controllable. This allows the binding and storage of subsequent state snapshots to be accurately associated with a specific session, laying the foundation for establishing cross-node state continuity.

[0016] Bind the state snapshot to the corresponding node of the distributed routing decision tree through the session ID, including: After assigning a globally unique session ID to the current session, the next step is to bind the state snapshot to the corresponding node of the distributed routing decision tree through the session ID to achieve an accurate association between the state and the routing decision. Specifically, when locating the partition node of the decision tree based on the hash value of the session ID, first perform a CRC32 hash operation on the session ID to obtain a hash value, and then perform a modulo operation on the hash value and the number of partitions of the distributed routing decision tree. The result is the index of the target partition node. If the number of partitions is 16 and the hash value modulo is 5, then the 5th partition node is located. When creating a snapshot index slot in the partition node memory, the system will allocate a fixed-size cache area in the node memory, and establish an index table in the order of the hash value of the session ID. Each index slot corresponds to an entry for a session ID. When storing the physical address pointer of the state snapshot, the physical storage address of the state snapshot in the distributed storage pool is converted into a 32-bit pointer variable and written into the corresponding index slot, so that the partition node can directly access the state snapshot data through the pointer. On this basis, a real-time relationship between the snapshot version number and the decision tree node status is established. Synchronization mechanism: the system assigns an auto-incrementing version number to each state snapshot (for example, the initial version is V1.0, which increases by 0.1 each time it is updated), and deploys a version monitoring process in the partition node. This process compares the snapshot version number with the version number currently recorded in the decision tree node every 100 milliseconds. If the snapshot version number is found to have increased, indicating that the state has been updated, the state synchronization of the decision tree node is immediately triggered. First, the read and write operations of the node are locked, the entity intent map in the latest state snapshot is loaded, the weight value of each routing path in the node is recalculated, and then the node is unlocked and the state update event is broadcast, so that the child nodes of the partition node synchronously adjust the routing logic to ensure that the state of the decision tree node is always consistent with the latest snapshot version, avoiding routing decision deviations due to state lag, and providing a reliable node-level association foundation for subsequently storing the state snapshot in the distributed session storage pool and establishing cross-node state continuity.

[0017] Storing the bound state snapshot in a distributed session storage pool to establish cross-node state continuity, including: After binding the state snapshot to the distributed routing decision tree node, in order to ensure the stable continuation of the state across nodes, the bound state snapshot needs to be stored in the distributed session storage pool. This process achieves state continuity through multi-node redundancy and real-time synchronization mechanism. Specifically, when using the write-time copy mechanism to persist the snapshot to at least three geographically dispersed storage nodes, the system first locks the memory copy of the current state snapshot, creates a writable copy, and then distributes the copy to geographically different storage nodes through asynchronous transmission, such as one node in East China, one in North China, and one in South China. After each node receives the copy, it verifies the data integrity by comparing the snapshot's checksum. After confirming that it is correct, it writes it to the local disk. Only when at least three nodes have completed the write and returned a successful response, is the persistence operation considered complete. This mechanism not only avoids interference with the original snapshot during the write process, but also reduces the risk of data loss caused by single point failures through multi-node storage. At the same time, a metadata chain containing the session ID, version number, and expiration time is created for each snapshot. The metadata chain is stored in a linked list structure. Each node records the update information of a snapshot. The session ID is used to associate the corresponding session, and the version number is used to associate with the corresponding session. The version of the state snapshot remains consistent, and the expiration time is dynamically set according to the session activity. For example, 24 hours after the last interaction, the head of the metadata chain points to the latest version, and the tail retains the historical version record to facilitate tracing the state evolution process. On this basis, state change events are pushed to relevant service nodes through the subscription and publication mode. The system deploys an event bus in the distributed session storage pool. All service nodes participating in session processing subscribe to the state change topic of the bus. When the state snapshot is updated, the storage pool automatically publishes an event containing the session ID, new version number and change time to the event bus. After receiving the event, the subscribing node immediately triggers a local cache update, pulls the latest state snapshot from the storage pool and replaces the old version, ensuring that the state information used by each node is consistent, providing a solid storage and synchronization foundation for establishing cross-node state continuity, so that when receiving new session requests subsequently, each node can quickly obtain a unified state snapshot to support coherent routing decisions.

[0018] S3. When a new user session request with the same session ID is received, the corresponding state snapshot is loaded from the distributed session storage pool. The loaded state snapshot is fused and analyzed with the semantic content of the new user session request to obtain the fused context state. The routing weight offset is calculated based on the fused context state. The routing decision table is modified according to the routing weight offset to eliminate intent drift. The loaded state snapshot is integrated with the semantic content of the new user session request for analysis, including: When a new user session request with the same session ID is received, the state snapshot loaded from the distributed session storage pool needs to be fused and analyzed with the semantic content of the new request. This process uses the fourth-order fusion logic to achieve seamless connection between historical and current session information, provide a complete context for accurate routing decisions, and align the semantic namespaces of the snapshot entity and the new request entity. The semantic namespace refers to the domain-specific semantic category set to avoid ambiguity in entity names. For example, in the e-commerce field, "Apple" refers to mobile phones under the "electronic equipment" namespace and to fruits under the "fresh food" namespace. Specifically, the entity synonym table in the domain knowledge graph is called to align the old entity in the snapshot with the new entity. The name "goods" and the new name "commodity" of the entity in the new request are mapped to a unified name, all of which are unified as "commodity". At the same time, the namespace label "commodity [e-commerce]" is attached to each entity to ensure the semantic consistency of cross-session entities, detect entity attribute conflicts and start confidence voting. Entity attribute conflicts refer to inconsistent attribute descriptions of the same entity in the snapshot and the new request. For example, the "price of commodity A" in the snapshot is 200 yuan, and it is 250 yuan in the new request. The system will extract the historical occurrence frequency of the conflicting attributes in the snapshot and the context weight in the new request, and generate the confidence of each attribute value (the old attribute value is set to 0) in combination with the domain common sense library (whether there is a fluctuation range in the commodity price). Confidence 0.3 for the original attribute value and 0.7 for the new attribute value), and then the attribute value with the highest confidence is selected as the fusion result through a voting mechanism. If the confidence difference is less than 0.2, the double attributes are retained and marked as "pending confirmation". The continuity feature in the historical entity chain is inherited. The continuity feature refers to the stable association relationship maintained by the entity in multiple rounds of conversations ("Product A-Order 123" is always bound in the historical conversation). The system will traverse the temporal relationship of the entity chain in the snapshot ("Product A→Add to Cart→Place Order"), and copy the entity association relationship that appears more than 3 times in a row to the fusion context, and mark the association strength ("strong association" "weak association") to ensure that the history will be The entity logic chain formed in the conversation is not interrupted. The intent label of the new request is corrected based on the intent migration trajectory. First, the evolution path of the historical intent is extracted from the snapshot (for example, from "inquire about price" → "understand inventory" → "inquire about delivery"), and then the sequence alignment algorithm is used to calculate the similarity between the new request intent label ("check logistics") and the most recent intent in the trajectory. If the similarity is ≥0.8, it is corrected to a label that better fits the trajectory (correcting "check logistics" to "inquire about delivery progress"), so that the intent label is more in line with the evolution of user needs. This not only preserves the traceability of information, but also forms a coherent and accurate conversation context, providing a reliable basis for the subsequent calculation of routing weight offsets.

[0019] Calculate the routing weight offset based on the fused context state, and modify the routing decision table based on the routing weight offset, including: After obtaining the fused context state, it is necessary to calculate the routing weight offset based on it and modify the routing decision table to dynamically adapt to changes in user intent and service node status. This process achieves accurate quantification of the weight offset through the joint calculation of three dynamic factors: the historical routing success rate factor is determined by the ratio of the number of successful routes to the total number of successful routes for the entity intent combination in the past 24 hours, the real-time service node load factor is calculated by monitoring the node CPU usage and memory occupancy, and the intent deviation coefficient is calculated based on the degree of deviation between the new intent and the historical intent trajectory. The three are weighted and summed at a ratio of 4:3:3 to obtain the final routing weight offset (such as 0.8×0.4+0.3×0.3+0.2×0.3=0.47), which is used when modifying the routing decision table. Using a hot-swap partitioning mechanism, the routing decision table is first partitioned into multiple logical partitions based on service nodes, with each partition corresponding to a group of node routing weights. The calculated weight offsets are then grouped by service node, with nodes within the same group sharing the same offset adjustment value (for example, a uniform 0.47 offset is added to the e-commerce service node group). The group size is dynamically adjusted based on network latency. If the average inter-node latency is less than 50ms, the offset is injected in batches of 10 nodes. If the latency is greater than 100ms, the offset is reduced to 3 nodes per group to avoid configuration synchronization delays caused by network congestion. During the injection process, the system first freezes the routing decision function of the target partition, applies the offset correction, and then verifies the validity of the decision table through health checks. Once verified, the partition is unfrozen and the new configuration is enabled. This process does not interrupt the normal operation of other partitions, ensuring routing service continuity. In this way, the routing decision table can respond to changes in the fusion context state in real time, ensuring that the routing strategy is both aligned with historical success and adapts to current service node load and user intent, significantly improving the accuracy and efficiency of service matching.

[0020] S4. Configure entity extraction rules and state snapshot generation strategies through a drag-and-drop interface to generate routing decision logic flows that integrate contextual states, achieving zero-code development.

[0021] Configure entity extraction rules and state snapshot generation strategies through a drag-and-drop interface to generate routing decision logic flows that integrate contextual states, including: A visual graph editor is provided. Users can drag entity type icons to the canvas to automatically generate extraction rules. The association rules and state snapshot compression parameters form a configuration template. When the configuration changes, the intermediate representation layer code is automatically compiled and generated. After sandbox verification, it is mapped to the node logic of the routing decision tree, and a three-dimensional topology diagram of the decision path is rendered in real time for users to adjust parameters and verify.

[0022] To achieve zero-code development, the system provides a drag-and-drop interface for users to configure entity extraction rules and state snapshot generation strategies, thereby generating a routing decision logic flow that integrates the context state. The specific process is as follows: First, the system provides a visual map editor. The left side of the interface displays entity type icons classified by field (such as "product" and "order" in the e-commerce field, and "account" and "transaction" in the financial field). Users can directly drag the required icons to the central canvas. The canvas will automatically generate corresponding entity extraction rules based on the icon type. After dragging the "product" icon, the editor loads the basic rules of "product name contains brand word + category word" and "price is number + currency unit" by default. Users can add custom conditions ("product inventory > 0") by double-clicking the rule entry. The logical relationship of the rule ("and" and "or") can be adjusted by dragging the connecting line to realize the visual configuration of the extraction rule. Then, the configured entity extraction rule is associated with the state snapshot compression parameter to form a configuration template. In the parameter panel on the right side of the editor, users can set the snapshot compression layer strategy, compression version number rules, and checksum algorithm. Clicking the "Associate Template" button packages the rules and parameters into an XML-formatted configuration template. The template contains the rule ID, parameter value, and association tag for easy reuse and modification. When users adjust the configuration, the system automatically triggers the compilation process, converting the configuration template into intermediate representation layer code. This code contains the judgment logic for entity extraction, the step instructions for snapshot generation, and the trigger conditions for context fusion. It is then sent to the sandbox environment for verification. The sandbox simulates real-world conversation scenarios. After entering test text, the intermediate code is run to check entity extraction accuracy, snapshot compression ratio, and decompression integrity. If there are rule conflicts or parameter anomalies, the specific error location is returned and corrections are prompted until verification passes. The verified intermediate representation layer code is mapped to the node logic of the distributed routing decision tree: entity extraction rules correspond to the decision tree's leaf nodes (responsible for determining whether the input text contains the target entity), state snapshot parameters correspond to intermediate nodes (control the timing and format of snapshot generation), and context fusion logic corresponds to the association strategy of the root node. At the same time, the editor will render a three-dimensional topological diagram of the decision path in real time, using nodes of different colors to indicate the priority of the rules (red indicates high priority), and lines with arrows to indicate the weight flow of the routing path (the wider the line, the higher the weight). Users can drag nodes to adjust the path weight, or modify the snapshot compression parameters using the slider. The topological diagram will be updated in real time and synchronously feedback the adjusted routing success rate prediction value for users to intuitively adjust and verify the parameters. Through this drag-and-drop configuration and visual verification mechanism, users can complete the full-process logic design from entity extraction and snapshot generation to routing decision without writing code. The generated routing decision logic flow can naturally integrate the context state, which is not only adapted to the personalized rules configured by the user, but also meets the technical requirements of cross-node state continuity and intent drift correction, greatly reducing the development threshold of large-model service-oriented orchestration.

[0023] The present invention parses user conversation text, extracts core entities and intent dependencies, constructs an entity intent graph, compresses and generates state snapshots, assigns a global conversation ID, binds the snapshots to distributed routing decision tree nodes and stores them, establishes cross-node state continuity, and when receiving a new request, fuses the snapshots with the new conversation semantics, calculates the routing weight offset to correct the decision table, configures rules and policies through a drag-and-drop interface, and generates a routing decision logic flow. This method solves the problems of conversation state loss and intent drift, improves service matching, lowers the development threshold, and is suitable for the agile construction of large-model collaborative applications.

[0024] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A zero-code visualization and automatic routing method for large-scale service-oriented orchestration, characterized by: The following steps are involved: S1. Parse the text content of the user's current conversation, extract the core entities and the intent dependencies between them, and construct an entity intent graph based on the core entities and the intent dependencies between them. Perform binary compression on the entity intent graph to generate a state snapshot and resolve state loss. S2. Assign a globally unique session ID to the current session, bind the state snapshot to the corresponding node of the distributed routing decision tree through the session ID, and store the bound state snapshot in the distributed session storage pool to establish cross-node state continuity; S3. When a new user session request with the same session ID is received, the corresponding state snapshot is loaded from the distributed session storage pool, and the loaded state snapshot is fused and analyzed with the semantic content of the new user session request to obtain a fused context state. A routing weight offset is calculated based on the fused context state, and the routing decision table is modified according to the routing weight offset to eliminate intent drift. S4. Configure entity extraction rules and state snapshot generation strategies through a drag-and-drop interface to generate routing decision logic flows that integrate contextual states, achieving zero-code development.

2. The zero-code visualization and automatic routing method for large-scale model service orchestration according to claim 1 is characterized by: Parsing the text content of the user's current conversation and extracting core entities and the intent dependency relationships between entities specifically includes: Key entities in the field are extracted through the named entity recognition module of the pre-trained large model, and a dynamic dependency network is constructed based on the co-occurrence frequency and semantic similarity between entities. The session context window is used to screen strongly associated entity pairs and filter out isolated noise entities. Finally, a weighted entity-intent association set is output, where the weight is determined by the frequency and position of the entity in the session.

3. The zero-code visualization and automatic routing method for large-scale model service orchestration according to claim 1 is characterized by: The entity intent graph is constructed based on the core entities and the intent dependency relationships between entities, including: Map the entity-intent association set into a directed graph structure, where entities are nodes and intent dependencies are weighted directed edges; The directed edge weights are initialized based on the historical routing decision path, and the directed edge weight values ​​are dynamically adjusted through real-time conversation flow. The weakly connected subgraphs in the directed graph structure are pruned and merged to generate an entity intent graph.

4. The zero-code visualization and automatic routing method for large-scale model service orchestration according to claim 1 is characterized by: Perform binary compression on the entity intent graph to generate a state snapshot, including: A layered compression strategy is adopted. Huffman coding is first performed on the label data of the graph nodes, and then differential pulse code modulation is performed on the matrix of directed edge weights. Key topological structure identifiers are retained during the compression process, so that the intention dependency path can be losslessly restored during decompression, and finally a binary block containing the compressed version number and check code is generated.

5. The zero-code visualization and automatic routing method for large-scale model service orchestration according to claim 1 is characterized by: The process of allocating a globally unique session ID to the current session includes: A composite identifier is generated by combining the user identity hash value, the service node geographic coordinates, and the timestamp milliseconds. The global uniqueness of the ID is ensured through a preset distributed consistency protocol. The session ID is injected into the root node metadata of the routing decision tree, making the root node the anchor point of the session state lifecycle.

6. The zero-code visualization and automatic routing method for large-scale model service orchestration according to claim 1 is characterized by: Binding the state snapshot to a corresponding node of the distributed routing decision tree through a session ID includes: The partition node of the decision tree is located according to the session ID hash value, and a snapshot index slot is created in the partition node memory to store the physical address pointer of the state snapshot. A real-time synchronization mechanism is also established between the snapshot version number and the decision tree node state to ensure that the decision tree reconstruction is automatically triggered when the state is updated.

7. The zero-code visualization and automatic routing method for large-scale model service orchestration according to claim 1 is characterized by: Storing the bound state snapshot in a distributed session storage pool to establish cross-node state continuity, including: A copy-on-write mechanism is used to persist snapshots in at least three geographically dispersed storage nodes. A metadata chain containing the session ID, version number, and expiration time is created for each snapshot. State change events are pushed to relevant service nodes through a subscription-publishing model to maintain cross-node state consistency.

8. The zero-code visualization and automatic routing method for large-scale model service orchestration according to claim 1 is characterized by: The fusing and analyzing the loaded state snapshot with the semantic content of the new user session request includes: Execute the fourth-order fusion logic, align the semantic namespaces of the snapshot entity and the newly requested entity, detect entity attribute conflicts and initiate confidence voting, inherit the continuity features in the historical entity chain, correct the intent label of the new request based on the intent migration trajectory, and generate a fusion context state with version traceability tags.

9. The zero-code visualization and automatic routing method for large-scale model service orchestration according to claim 1 is characterized by: The calculating of the routing weight offset based on the fused context state and modifying the routing decision table according to the routing weight offset includes: The weight offset is calculated jointly by three dynamic factors: the historical routing success rate factor, the real-time service node load factor, and the intention deviation coefficient. When correcting the decision table, a hot replacement partition mechanism is used to inject the offset in batches according to the service node group, and the group size is dynamically adjusted according to the network delay.

10. The zero-code visualization and automatic routing method for large-scale model service orchestration according to claim 1 is characterized by: The configuration of entity extraction rules and state snapshot generation strategies through the drag-and-drop interface to generate a routing decision logic flow that integrates the context state includes: A visual graph editor is provided. Users can drag entity type icons to the canvas to automatically generate extraction rules. The association rules and state snapshot compression parameters form a configuration template. When the configuration changes, the intermediate representation layer code is automatically compiled and generated. After sandbox verification, it is mapped to the node logic of the routing decision tree, and a three-dimensional topology diagram of the decision path is rendered in real time for users to adjust parameters and verify.

Citation Information

Patent Citations

  • Government affair service field multi-strategy fusion dialogue method based on knowledge graph

    CN116628172A

  • Replacement method for guaranteeing MEC service continuity

    CN117336808A

  • Large model deployment method and system based on multi-level cache mechanism

    CN119739809A

  • Large model prompt project optimization system and method fusing domain knowledge graph

    CN120196734A

  • Large model task routing method, device and medium

    CN120429307A

Cited By

  • Web3D zero code interactive design platform and method based on unified model

    CN121560308A

  • Toll station vehicle transaction state sharing and seamless continuing method and system

    CN121884473A

  • Model-driven code-free configuration development method and development system

    CN121934833A