Multi-modal system fusing dynamic semantic orchestration and cross-platform agent collaborative reasoning

By integrating dynamic semantic orchestration with cross-platform intelligent agent collaborative reasoning into a multimodal system, the challenges of diverse user intents and complex business logic processing in existing technologies are solved, enabling customer service robots to achieve flexibility and accuracy in complex business scenarios, and ensuring the robustness and fluency of the system.

CN121683858BActive Publication Date: 2026-05-15HANGZHOU XINGYU ARTIFICIAL INTELLIGENCE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU XINGYU ARTIFICIAL INTELLIGENCE CO LTD
Filing Date
2026-02-11
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing customer service robots or intelligent assistants struggle to translate users' unstructured, multi-intent natural language expressions into precise and ordered backend operation sequences when handling dynamic tasks involving real-time data interaction and complex business logic. This results in rigid interaction processes, an inability to handle abnormal interruptions, insufficient robustness, and an inability to meet the application needs of complex business scenarios.

Method used

A multimodal system integrating dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning is adopted. Through the user's original speech stream acquisition module, speech stream multi-intent parsing module, logical segment matching module, and compound prompt word construction module, a hybrid structure of compound prompt words is constructed to guide the large language model to perform deterministic reasoning and cross-platform collaborative operation.

Benefits of technology

It enables flexible dialogue and accurate completion of complex tasks such as queries and order placement in real-time interaction, solving the problems of illusion and insufficient robustness of traditional solutions, and ensuring the flexibility and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121683858B_ABST
    Figure CN121683858B_ABST
Patent Text Reader

Abstract

The application relates to the field of intelligent collaborative interaction, and specifically discloses a multi-modal system fusing dynamic semantic arrangement and cross-platform agent collaborative reasoning. First, the original continuous speech stream of a user is deeply analyzed to accurately identify multiple intents coexisting therein. Subsequently, the system matches and searches the intents with a pre-constructed service logic topology depicting a complete business context, and dynamically arranges the same into a logic segment set. Finally, the logic segments are constructed into a hybrid structure composite prompt word fusing natural language, function calling and code instruction. Submitting the highly structured prompt word to a large language model can guide the same to perform deterministic reasoning and seamless cross-platform collaborative operation, thereby ensuring that complex tasks such as query and ordering can be flexibly conversed, accurately and reliably completed in real-time interaction, and effectively solving the problems of illusion and insufficient robustness of traditional schemes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent collaborative interaction, and more specifically, to a multimodal system that integrates dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning. Background Technology

[0002] With the rapid development of Large Language Models (LLM) and multimodal technologies, automated customer service systems are undergoing profound changes. However, current mainstream customer service robots or intelligent assistants still struggle with dynamic tasks requiring real-time data interaction and complex business logic, such as simultaneously querying real-time inventory and member information and completing reservations during telephone conversations. The fundamental reason is that users' natural language expressions are often unstructured and contain multiple intentions, and traditional systems struggle to transform this dynamic and ambiguous input into precise and ordered sequences of backend operations. This leads to rigid interaction processes, an inability to handle abnormal interruptions, and the inability to simultaneously execute thoughts and actions like human customer service representatives, thus limiting the depth and breadth of their application in complex business scenarios.

[0003] To address these issues, existing technologies have explored various approaches. For instance, Retrieval Augmentation (RAG)-based solutions attempt to improve the accuracy of model responses through external knowledge bases. However, when rigorous business logic needs to be executed rather than simple question-and-answer sessions, the accuracy and logical coherence of the retrieved content are difficult to guarantee. Multi-agent collaborative solutions, on the other hand, often suffer from insufficient robustness when faced with frequently changing user intent or mid-dialogue backtracking due to the complexity and fragility of the collaborative links and contextual transmission, leading to process interruptions. Furthermore, relying solely on natural language prompts to drive large models for end-to-end operations cannot completely avoid the model illusion problem, failing to meet the stringent accuracy requirements of fields such as finance, e-commerce, and catering. Therefore, existing technologies lack a systematic solution capable of real-time parsing of dynamic, multi-intent-based user speech streams and efficiently and accurately orchestrating and scheduling them into complex task flows encompassing dialogue, querying, computation, and cross-platform execution.

[0004] To bridge the gap between the flexibility of natural language and the determinism of computer system execution, a completely new technological solution is urgently needed. Summary of the Invention

[0005] To address the aforementioned technical challenges, this application is proposed. According to this application, a multimodal system integrating dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning includes:

[0006] The user's original voice stream acquisition module is used to acquire the user's original voice stream.

[0007] The speech stream multi-intent parsing module is used to perform multi-intent parsing on the user's original speech stream to obtain an intent set;

[0008] The logical fragment matching module is used to perform matching and retrieval of the intent set based on the service logical topology to obtain the logical fragment set;

[0009] The compound prompt word construction module is used to construct compound prompt words with mixed structures based on a set of logical fragments;

[0010] The execution result generation module is used to submit compound prompt words with mixed structures to the large language model to obtain the execution result, which includes the response draft text and real-time data.

[0011] The audio conversion module is used to convert the execution result into a synthesized audio response after the response synthesis.

[0012] Compared to existing technologies, this application provides a multimodal system that integrates dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning. First, it performs deep analysis of the user's continuous raw speech stream to accurately identify multiple coexisting intents. Then, the system matches these intents with a pre-constructed service logic topology that depicts the complete business context, dynamically orchestrating them into a set of logical fragments. Finally, these logical fragments are constructed into a hybrid structured prompt word that integrates natural language, function calls, and code instructions. Submitting this highly structured prompt word to a large language model guides it to perform deterministic reasoning and seamless cross-platform collaborative operations, ensuring both flexible dialogue and accurate and reliable completion of complex tasks such as queries and order placement in real-time interactions. This effectively solves the problems of illusion and insufficient robustness inherent in traditional solutions. Attached Figure Description

[0013] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain the application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0014] Figure 1 This is a block diagram of a multimodal system that integrates dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning according to an embodiment of this application.

[0015] Figure 2 This is a schematic diagram of data flow in a multimodal system that integrates dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning according to an embodiment of this application.

[0016] Figure 3This is a block diagram of the speech stream multi-intent parsing module in a multimodal system that integrates dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning according to an embodiment of this application.

[0017] Figure 4 This is a schematic diagram of the data flow of the logical segment matching module in a multimodal system that integrates dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning according to an embodiment of this application. Detailed Implementation

[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0019] This application addresses the issues of process rigidity and operational inaccuracy in existing technologies when handling complex business interactions with dynamic and multi-intents. Figure 1 This is a block diagram of a multimodal system that integrates dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning according to an embodiment of this application. Figure 2 This is a schematic diagram of data flow in a multimodal system integrating dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning according to embodiments of this application. Specifically, as shown... Figure 1 and Figure 2 As shown, the multimodal system 100 integrating dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning according to an embodiment of this application includes: a user original speech stream acquisition module 110, used to acquire the user's original speech stream; a speech stream multi-intent parsing module 120, used to perform multi-intent parsing on the user's original speech stream to obtain an intent set; a logical segment matching module 130, used to perform matching and retrieval on the intent set based on service logical topology to obtain a logical segment set; a compound prompt word construction module 140, used to construct compound prompt words with a mixed structure based on the logical segment set; an execution result generation module 150, used to submit the compound prompt words with a mixed structure to a large language model to obtain an execution result, the execution result including a response draft text and real-time data; and an audio conversion module 160, used to convert the execution result into a synthesized audio response after the response synthesis.

[0020] Specifically, the user's original voice stream acquisition module 110 is used to acquire the user's original voice stream. It is understood that in automated customer service scenarios that heavily rely on voice interaction, such as telephone customer service robots, accurately understanding the user's verbal instructions and needs is a prerequisite for efficient service. Traditional customer service systems often struggle to capture the complex intentions and real-time dynamic information contained in user speech, especially when user expression is not singular and structured, but contains multiple intentions and incorporates real-time data interaction needs. Existing technical solutions exhibit rigidity in their processing flow, making it difficult to adapt to such dynamic changes, thus limiting the system's application depth in complex business scenarios. Therefore, to overcome the service limitations caused by the inability to accurately acquire and understand user voice information in existing technologies, and to provide reliable raw input for subsequent intelligent processing, ensuring the complete and real-time acquisition of the user's original voice stream is crucial. This lays the foundation for subsequent multi-intention parsing of the voice stream, thereby achieving accurate capture and recognition of user intentions, and ultimately supporting more refined and personalized service responses.

[0021] In a feasible technical solution, the user's original voice stream acquisition module 110 processes the data as follows: The user's original voice stream specifically refers to the continuous digital audio data sequence formed by the sound wave signal emitted by the user, such as through a telephone, smart speaker, or computer voice input device, after being digitized by the analog-to-digital converter inside the device. These uncompressed or unencoded original data sequences completely contain all the acoustic information of the user when speaking, such as the language content, the speaker's unique intonation, speaking speed, subtle changes in emotional state, and background noise that may exist in the call environment.

[0022] In practice, the user's original voice stream acquisition module is responsible for establishing and maintaining a stable and reliable data transmission channel with the external audio source. For example, in a typical telephone customer service application scenario, this module needs to be deeply integrated with the backend call center infrastructure. This may include traditional interface communication with the Public Switched Telephone Network (PSTN) or receiving data through an IP telephony system based on the Session Initiation Protocol (SIP). When a customer dials the customer service hotline and begins to speak, their sound waves are first captured by the microphone of the telephone terminal device and then converted into analog electrical signals. These analog signals are then converted into standard digital audio formats, such as the widely used Pulse Code Modulation (PCM) format, by the analog-to-digital converter (ADC) inside the telephone terminal. This PCM data is then transmitted in real time to the user's original voice stream acquisition module via the communication network in the form of continuous data packets or data blocks.

[0023] The user's raw voice stream acquisition module continuously and efficiently listens to, receives, and buffers these incoming digital audio data frames during this stage to ensure the continuity and integrity of the data stream, laying a solid foundation for subsequent voice processing. For example, when a user interacts with the system via telephone and says, "I'd like to order dinner for 6 PM tonight, and by the way, do you have parking spaces?", this module will receive the complete raw digital audio data corresponding to this entire sentence in real time from the established communication channel. Mathematically, this raw data can be abstractly represented as a series of continuous, discrete digital samples: ,in, To obtain the entire original user voice stream, This represents an audio sample value sampled and quantized at time j. These sample values ​​are precise digital representations of the original sound wave amplitude at a specific point in time. They are collected at a high density using a preset sampling frequency, such as the industry standard 8kHz or 16kHz, and each sample has a certain quantization precision, such as 8 bits or 16 bits, to ensure the detail of the acoustic information. To illustrate with a practical example, if we use the 8kHz sampling rate and 16-bit quantization precision commonly used in telephone voice communication, this means that the amount of raw voice stream data per second is approximately 8000 samples / second × 16 bits / sample = 128000 bits / second, or 16000B / second. These raw digital sample sequences constitute the most basic, unprocessed auditory input, which is stored sequentially and in real-time in the module's internal buffer. The internal implementation of this module depends on the correct configuration of the underlying operating system's audio interface, network driver, and communication protocol stack. For example, to optimize the real-time transmission performance of audio data, it may be necessary to finely configure the parameters of the Transmission Control Protocol (TCP / UDP), the packet size of the Real-Time Transport Protocol (RTP), and the buffer strategy. These preset values ​​are determined based on a comprehensive balance of the bandwidth conditions of the target communication network, the system's strict requirements for end-to-end latency, and the actual processing capabilities of the hardware used, in order to minimize packet loss, transmission delay, and jitter, thereby ensuring the fidelity, continuity, and real-time performance of the acquired user's original voice stream.

[0024] Specifically, the speech stream multi-intent parsing module 120 is used to perform multi-intent parsing on the user's original speech stream to obtain an intent set. Correspondingly, in complex human-computer interaction scenarios, the user's original speech stream is not only a carrier of information but also a multimodal signal rich in semantics, emotion, and environmental cues. However, the original digital audio data itself is unstructured and cannot be directly used for advanced logical decision-making and reasoning. Existing technologies have a significant gap in transforming this continuous, dynamic, and often multi-intent-containing fuzzy input into precise, ordered machine-executable instructions, resulting in rigid interaction processes and an inability to truly simulate the fluency of simultaneous thinking and operation in human dialogue. Therefore, deep multi-intent parsing of the user's original speech stream becomes a crucial and necessary preprocessing step. Its purpose is to transform this raw, unstructured audio data stream into a high-dimensional information set composed of clear, structured intents. This provides clear, rich, and operable input for subsequent dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning, and is the fundamental guarantee for achieving flexible and accurate responses from the entire system.

[0025] Figure 3 This is a block diagram of a speech stream multi-intent parsing module in a multimodal system integrating dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning, according to an embodiment of this application. Figure 3 As shown, in a feasible technical solution, the speech stream multi-intent parsing module 120 includes: a speech stream processing unit 121, used to perform speech stream preprocessing and multimodal feature extraction on the user's original speech stream to obtain standardized text and acoustic feature vectors; an additional context generation unit 122, used to perform dependency parsing and context information fusion on the standardized text, acoustic feature vectors, and dialogue history to obtain a set of clauses with additional context; and an intent analysis unit 123, used to perform multi-intent weighted scoring and structured intent generation on the set of clauses with additional context to obtain an intent set.

[0026] Specifically, the processing of the speech stream multi-intent parsing module 120 is as follows: In a feasible technical solution, the speech stream processing unit 121 includes: a speech segmentation subunit 1211, used to input the user's original speech stream into the speech activity detection module to obtain speech segments; and a text acoustic feature extraction subunit 1212, used to input the speech segments into the speech recognition engine and the acoustic feature extractor respectively to obtain standardized text and acoustic feature vectors.

[0027] First, the speech segmentation subunit 1211 is responsible for the initial processing of the input continuous user speech stream. Taking the user speech stream obtained in the above embodiment as an example, this speech stream not only contains the user's valid utterance "I want to order dinner at 6 pm tonight, and by the way, do you have parking spaces?", but may also contain brief silences before the user speaks, pauses between sentences, and environmental noise after the utterance. These non-speech parts are redundant information for semantic understanding and may even interfere with the accuracy of subsequent processing. Therefore, the primary responsibility of the speech segmentation subunit is to use its built-in speech activity detection module to accurately segment the valid speech segments from the entire data stream. In terms of technical implementation, the speech activity detection module adopts a deep learning-based model architecture, such as a recurrent neural network based on a gated recurrent unit (GRU). The input of this model is not the original audio sample points, but rather the acoustic features extracted from each frame after the original speech stream is segmented into continuous short frames (e.g., each frame is 25 milliseconds long with a frame shift of 10 milliseconds), such as Mel-frequency cepstral coefficients (MFCCs). These MFCCs effectively simulate the auditory characteristics of the human ear and are standard features in the field of speech signal processing. The GRU model consists of several stacked GRU layers to capture the temporal dependencies between audio frames, and a final sigmoid activation function, the sigmoid output layer. This output layer provides a probability value between 0 and 1 for each input frame, representing the confidence that the current frame belongs to speech. The model's weights and biases are obtained through supervised training on a massive audio dataset with manually labeled speech and non-speech segments. The training objective is to minimize the binary cross-entropy loss between the model's predictions and the manual annotations. During runtime, when the user's original speech stream is input, the speech segmentation subunit 1211 first converts it into an MFCC feature sequence, and then inputs this sequence into the pre-trained GRU model. The model outputs a speech activity probability curve synchronized with the audio frames. Subsequently, a decision logic unit processes this probability curve. By setting a predefined probability threshold, such as 0.6, which is determined by balancing recall and precision on the validation set, frames exceeding this threshold are initially identified as speech frames. To avoid excessive fragmentation of speech segments due to momentary noise or brief pauses, a smoothing mechanism is applied. For example, a speech segment only begins when the probability of speech exceeds the threshold for more than 5 consecutive frames, and ends when the probability of speech falls below the threshold for more than 15 consecutive frames. Finally, for the input raw speech stream containing "I'd like to order dinner at 6 PM tonight, and by the way, do you have parking?", the speech segmentation subunit removes the silences before and after it, outputting one or more compact speech segments containing only the valid spoken content.

[0028] After obtaining the speech segments, the text acoustic feature extraction subunit 1212 begins parallel processing. The speech segments are simultaneously fed into two independent components: a speech recognition engine and an acoustic feature extractor. The speech recognition engine's task is to convert the acoustic signals of the speech segments into text. This engine can employ an advanced end-to-end model architecture, such as an attention-based Transformer or Conformer model. Taking the Conformer model as an example, its architecture consists of an encoder and a decoder. The encoder, through stacked Conformer modules, combines the ability of convolutional neural networks to capture local features with the ability of self-attention mechanisms to capture global context, performing deep feature extraction on the log-Melogram of the input speech segments. The decoder, using an attention mechanism, generates corresponding text characters or sub-word units one by one, guided by the acoustic representations output by the encoder. The model's parameters are also obtained by jointly training on a large-scale corpus containing hundreds of thousands of hours of annotated speech, optimizing the Connectionized Temporal Classification (CTC) loss and cross-entropy loss. When the clean speech segments are input into the engine, it outputs the most probable text sequence, i.e., the original text. In our example, the original output text might be: "I'd like to order dinner for 6 PM tonight, and by the way, do you have parking?" This original text is then fed into a text normalization module. This module consists of a series of rule-based algorithms or finite-state transformers (FSTs) that format and normalize the original text. For example, it uses predefined regular expressions and a dictionary to convert the colloquial "six o'clock" into the standard 24-hour format "18:00" based on the context of "tonight," correcting potential recognition errors and adding appropriate punctuation. Finally, the module outputs the normalized text: "I'd like to order dinner for 18:00 tonight, and by the way, do you have parking?". Simultaneously, on a parallel processing path, an acoustic feature extractor analyzes the same speech segment to extract paralinguistic information beyond the literal meaning. This information is crucial for determining the user's emotion, attitude, and the emphasis of their speech. This extractor consists of a set of specialized digital signal processing algorithms used to calculate multiple acoustic metrics. For example, for pitch features, the extractor uses algorithms such as YIN or autocorrelation function to calculate the fundamental frequency (F0) of the speech frame by frame. Specifically, the autocorrelation function method calculates the similarity between the speech signal and itself at different time delays. When the similarity reaches a peak, the corresponding time delay is regarded as the pitch period, and its reciprocal is the fundamental frequency. The YIN algorithm improves upon the autocorrelation function method by introducing steps such as difference function and cumulative normalization, which can more accurately estimate the fundamental frequency and effectively reduce common octave or half-frequency errors.For the part "I'd like to order food for 6 PM tonight," the tone of voice is likely to be relatively steady, with a fundamental frequency average of approximately 110 Hz, likely belonging to an adult male. However, for the part "By the way, do you have parking spaces?", the tone at the end of the sentence might rise, causing a local fundamental frequency peak to reach 140 Hz. The extractor calculates statistical quantities such as the fundamental frequency mean, standard deviation, maximum, and minimum values ​​for the entire speech segment. For energy features, the volume variation can be quantified by calculating the root mean square energy of each frame. This calculation process involves squaring the amplitude values ​​of all audio sampling points within a frame, calculating the average of these squared values, and finally taking the square root of the average to obtain a value that objectively reflects the average amplitude within the speech frame. When stating the primary need of ordering food, the user's volume may be higher, while when stating the secondary issue of parking spaces, the volume may decrease slightly. Similarly, statistical indicators such as the energy mean, maximum, and dynamic range are calculated. Speech rate is calculated by dividing the number of syllables or words in the standardized text by the duration (in seconds) of the speech segment. This calculation relies on two key data points: the total number of characters or syllables in the standardized text obtained by the speech recognition engine, and the precise duration of the speech segment determined by speech activity detection. Dividing these two values ​​yields the speech rate value in "words per second" or "syllables per second". A steady speech rate may correspond to a declarative tone, while a sudden increase or decrease in speech rate may suggest changes in the user's emotions, such as anxiety or hesitation. Ultimately, all these calculated acoustic metrics, such as mean pitch, standard deviation of pitch, mean energy, dynamic range of energy, speech rate, and proportion of high-frequency energy, are combined into a multi-dimensional floating-point vector, namely the acoustic feature vector. For example, a specific acoustic feature vector might be represented as: [115.2,12.5,140.1,98.3,0.78,-15.4,4.5], where each dimension corresponds to a specific acoustic statistic, namely the fundamental frequency mean, standard deviation, maximum value, minimum value, energy mean, energy standard deviation, speech rate, etc.

[0029] In a feasible technical solution, the additional context generation unit 122 includes: a dependency parsing graph generation subunit 1221, used to input standardized text into a dependency parsing model to obtain a dependency parsing graph; a semantic core segmentation subunit 1222, used to perform semantic core segmentation on the dependency parsing graph to obtain a clause set; a current context information generation subunit 1223, used to perform dual-channel context feature vectorization on standardized text, acoustic feature vectors, and dialogue history to obtain current context information; and an encapsulation subunit 1224, used to perform entity recognition on the clause set, and then perform context fusion encapsulation with the current context information to obtain a clause set with additional context.

[0030] First, the dependency parsing graph generation subunit 1221 receives the standardized text from the previous module. The core of this subunit is a pre-trained dependency parsing model, whose function is to reveal the grammatical dependencies between words in a sentence and express these dependencies in the form of a directed acyclic graph. Technically, this model can employ a graph-based parsing method, with its internal architecture containing a deep neural network module, such as a structure combining a bidirectional long short-term memory network (Bi-LSTM) and a multilayer perceptron (MLP). The model processes the standardized text as follows: First, each word in the text (after word segmentation) is fed into an embedding layer. The core architecture of this embedding layer is a large trainable lookup matrix, with the number of rows equal to the size of the pre-built vocabulary, and the number of columns being the preset embedding dimension (a hyperparameter determining the density of the vectors, e.g., 300). During processing, each word is first mapped to a unique integer index by consulting the vocabulary. This index is then input to the embedding layer, which performs an efficient lookup operation, retrieving the specific row vector corresponding to that index from the matrix. This vector is the dense vector representation of the word, i.e., the word vector. The entire lookup matrix constitutes all the trainable parameters of this layer. Its initial values ​​can be set by loading word vectors pre-trained on massive unlabeled text corpora (e.g., generated by GloVe or Word2Vec algorithms) to inject rich prior semantic knowledge. During the overall supervised training of the dependency parsing model, these word vector parameters, along with all other network parameters, will undergo end-to-end gradient updates and fine-tuning through backpropagation, thereby making their representation ability specifically adapted to the current syntactic parsing task and ensuring that words with similar syntactic functions are located closer together in the vector space. This layer maps each word to a high-dimensional real-valued vector, i.e., the word vector. In the second step, the word vector containing the sequence order is input into a bidirectional long short-term memory network. Bi-LSTM traverses the entire word sequence in both forward and backward directions, thereby generating a context-aware hidden state vector for each word. This vector not only contains information about the word itself but also incorporates contextual information from its preceding and following words, making the representation richer and more accurate. Third, for any two word pairs in the sentence, a multilayer perceptron specifically designed for scoring uses their context-aware vectors as input and calculates a score. This score represents the probability or strength of a dependency relationship between the two words. By scoring all possible word pairs in the sentence, a complete dependency score matrix is ​​obtained. Finally, to ensure that the output is a valid syntactic tree—that is, each word has only one dominant headword, or parent node, excluding the root node—the model employs a maximum spanning tree algorithm from graph theory (e.g., the Chu-Liu / Edmonds algorithm) to decode the globally optimal dependency syntactic graph from the score matrix.The parameters such as the weights and biases of this dependency parsing model are obtained through supervised learning on a large-scale artificial treebank manually annotated with dependency relations by linguistic experts. When the normalized text "I want to order a meal for 18:00 tonight. By the way, do you have a parking space?" is input into this model, the dependency parsing graph generation subunit will output a specific dependency parsing graph G=(V,E). Among them, the vertex set V is all the lexical nodes in the sentence, such as {"I", "want", "order", "tonight", "18:00", "of", "meal", ",", "by the way", "ask", "you", "have", "parking space", "?", "?"}. The edge set E represents the syntactic dependency relations between the words. For example, the graph will contain some directed edges like: an edge from "want" to "I" labeled as the subject relation; an edge from "order" to "want" labeled as the verb-core relation, indicating that ordering a meal is the content of "want"; an edge from "order" to "meal" labeled as the object relation; an edge from "meal" to "18:00" forming a modifier-head relation through a preposition, indicating the time modification of the meal. Similarly, in the part after the comma, "ask" is the core verb, "parking space" is the object of "have", and the whole "do you have a parking space" as an object clause has its dependency relation pointing to "ask".

[0031] Next, this structured dependency parsing graph G is passed to the semantic core segmentation subunit 1222. The task of this subunit is to split this complex graph into several smaller and simpler subgraphs according to preset rules, with each subgraph corresponding to an independent semantic core. This process is divided into two steps: split point identification and graph segmentation algorithm execution. In the split point identification step, the subunit will traverse all the vertices in the dependency parsing graph G , and determine whether it is a segmentation point based on its word nature and dependency relation tags. According to the preset rules, the recognition conditions for segmentation points include: the word nature is a coordinating conjunction (such as "and", "with") and the dependency relation is parallel; the word nature is a conjunction of turning or connecting relation and the dependency relation is an adverbial clause; or, the node is a punctuation mark representing the internal logical segmentation of the sentence (such as a comma, and a semicolon;) and the dependency relation is punctuation. In the dependency syntax graph of this application, the lexical node "," meets the conditions of the punctuation node, so it is recognized as a key segmentation point. After all segmentation points are recognized, the graph segmentation algorithm starts to execute. This algorithm takes the root node of the graph, which is the core predicate verb of the sentence, as the starting point and performs a depth-first traversal (DFS) to construct subgraphs. During the traversal process, the algorithm will visit each node in the graph along the directed edges of the dependency relation. When the traversed path encounters a node that has been previously recognized as a segmentation point, the continuation of this path will be blocked. At this time, the set of all nodes that have been traversed through this path forms an independent subgraph. The segmentation point itself will be discarded after segmentation. The algorithm will continue this process and may start new traversals from multiple unvisited root nodes until all non-segmentation point nodes in graph G are clearly assigned to a certain subgraph. Specifically in this example, the algorithm may first start traversing from the core verb "order" in the first half of the sentence and visit all its associated nodes: {"I", "want to", "order", "tonight", "18:00", "of", "meal"}. When the traversal attempts to connect to the punctuation node "," through "order" or "meal", since "," is a segmentation point, this path is cut off. Therefore, this set of nodes forms the first subgraph. Subsequently, the algorithm will process the remaining unvisited part of the graph and start a new traversal from the core verb "ask" in the second half of the sentence, thereby visiting all related nodes: {"by the way", "ask", "next", "you", "have", "parking space", "or not"}. Finally, this complex dependency syntax graph is successfully segmented into two independent subgraphs. After the graph segmentation is completed, the last step is to recombine the lexical nodes in each subgraph into strings according to their order in the original standardized text. The nodes of the first subgraph are combined in order to generate the first clause: I want to order the meal at 18:00 tonight. The nodes of the second subgraph are also combined in order to generate the second clause: By the way, ask if you have a parking space. These two generated clauses are stored in a set to form the final clause set. This clause set {"I want to order the meal at 18:00 tonight", "By the way, ask if you have a parking space"}, compared with the original single long sentence, has a clearer structure, and each element highly focuses on a single and complete semantic intention.

[0032] Subsequently, the current context information generation subunit 1223 receives acoustic feature vectors and standardized text from the speech stream processing unit, as well as the dialogue history continuously maintained by the dialogue management module. The dialogue history is a structured data record that faithfully records each round of interaction since the beginning of the current session in chronological order. Its data structure is a list, where each element represents a round of dialogue, containing information such as the speaker (user or system), the original utterance text, the system-parsed intent, the actions performed by the system, and any identified key entities. This historical record is updated and stored in real-time by the dialogue management module after each round of interaction. This subunit generates a comprehensive current context information object through two parallel processing channels: an emotion calculation channel and a dialogue behavior analysis channel. In the emotion calculation channel, the subunit inputs the received acoustic feature vectors, calculated in the aforementioned example as [115.2,12.5,140.1,98.3,0.78,-15.4,4.5], into a pre-trained acoustic emotion classification model. This model can be a Support Vector Machine (SVM). Its basic architecture lies in finding an optimal decision hyperplane to maximize the margin between samples of different emotion categories in a high-dimensional space composed of acoustic features. To handle the complex nonlinear relationship between acoustic features and emotion, the model employs kernel function techniques, such as the Gaussian radial basis function (RBF) kernel, to map the original feature space to a higher-dimensional space, thus making originally linearly inseparable samples linearly separable. The key parameters of this SVM model, including the selection of support vectors and the specific parameters of the kernel function, are obtained through supervised learning on a large dataset containing tens of thousands of manually labeled speech segments with emotion categories (such as normal, anxious, happy, angry, etc.) and their corresponding acoustic feature vectors. When an acoustic feature vector is input, the model calculates which vector belongs to each predefined emotion category. The possibility of [the outcome]. This decision-making process can be represented by the following formula: ,in It is the set of all predefined emotion categories, for example, E = {normal, anxious, happy}. This represents a given current acoustic feature vector. Under these conditions, it belongs to the emotional category. The posterior probability. Although the output of a standard SVM is a class decision rather than a probability, its output decision distance can be transformed into a reliable probability estimate through calibration methods such as Platt scaling. For example, the model might compute the following probability distribution: =0.85, =0.10, =0.05. By taking the maximum posterior probability, The operation will select the category with the highest probability value, thus the final output sentiment label. This is normal. Meanwhile, in the dialogue behavior analysis channel, the sub-unit inputs the received standardized text and the current dialogue history into a dialogue behavior classification model. This model can employ a classifier architecture based on a pre-trained language model (such as BERT). Its workflow is as follows: First, the current dialogue history is encoded. Since this is the user's first valid question in this example, the dialogue history may only contain the system's opening remarks, and therefore can be encoded as a low-dimensional history state vector representing the initial state. Second, the standardized text is input into the BERT encoder, which, leveraging its powerful contextual understanding capabilities, generates a text vector that captures the deep semantics of the entire sentence, and the output vector corresponding to the [CLS] label is taken. Subsequently, the history state vector and the text vector are concatenated dimensionally to form a fused feature vector containing both historical and current semantics. Finally, this fused feature vector is input into a simple multilayer perceptron (MLP) classifier head, whose output layer uses the Softmax activation function to calculate whether the input belongs to each predefined dialogue behavior category. The model calculates the probabilities of actions such as initiating a new topic, clarifying, supplementing, and confirming. The model's parameters are fine-tuned on a large-scale corpus of dialogue behaviors labeled with annotations. The training objective is to minimize the cross-entropy loss between the predicted behavior and the human annotation. For the input in this example, since it's the start of a conversation, the model is highly likely to output the highest probability for the "Initiating a New Topic" category, for example: P(Initiating a New Topic|...) = 0.92, P(Clarifying|...) = 0.05, P(Supplementing|...) = 0.03. Therefore, the final output dialogue behavior label is "Initiating a New Topic". After processing through both channels, the current context information generation subunit encapsulates the obtained sentiment label (normal) and dialogue behavior label (Initiating a New Topic) into a structured current context information object and outputs it. This object is: {sentiment label: "normal", dialogue behavior label: "Initiating a New Topic"}.

[0033] Next, the encapsulation subunit 1224 receives the clause set generated in the previous steps and the newly generated current context information object. The core task of this subunit is to traverse each clause in the clause set, extract entity information for each clause, and bind the shared context information to them. At the heart of this process is a high-performance Named Entity Recognition (NER) model. This model can employ the industry-standard Bidirectional Long Short-Term Memory (BLSTM) Conditional Random Field (CRF) architecture. Its working principle is as follows: First, each word in the clause is converted into a vector through a word embedding layer; then, the BiLSTM network learns the contextual dependencies of each word in both forward and backward dimensions, generating context-aware feature representations; finally, the CRF layer learns the transition constraints between entity labels on top of the BiLSTM output (e.g., an intermediate word label representing time is highly likely to follow a starting word label representing "time"), thus decoding the globally optimal entity label sequence. The model's parameters are also trained on a large-scale corpus labeled with various entities (such as time, location, product names, etc.). The encapsulation process for the sub-unit is as follows: It first initializes an empty set of clauses for the additional context. Then, it begins to iterate through the input set of clauses. For the first clause: "I want to order dinner at 18:00 tonight," it calls the NER model to extract entities. The model will recognize that "18:00 tonight" is a time entity and "dinner" is an entity referring to an item. Therefore, the generated entity list is [{'entity':'18:00 tonight','type':'time'},{'entity':'dinner','type':'item'}]. Subsequently, a new clause object for the additional context is created and the data is encapsulated: the clause string is assigned to the text content field, the entity list is assigned to the entity list field, and the shared current context information object is assigned to the context information field. This complete object is then added to the empty set. For the second clause: "By the way, do you have parking spaces?", the NER model will recognize that "parking space" is a facility entity. The generated entity list is [{'entity':'parking space','type':'facility'}]. Similarly, a new clause object with additional context is created and populated with the corresponding data, including the same current context information object, and then added to the collection. After traversing all clauses, the encapsulation unit finally outputs a complete set of clauses with additional context.This is a list structure containing two deeply augmented objects: [{text content:"I'd like to order dinner for 6 PM tonight", entity list:[{'entity':'6 PM tonight','type':'time'},{'entity':'dinner','type':'item'}], context information:{emotion tag:"normal", dialogue behavior tag:"starting a new topic"}},{text content:"By the way, do you have parking spaces?", entity list:[{'entity':'parking space','type':'facilities'}], context information:{emotion tag:"normal", dialogue behavior tag:"starting a new topic"}}].

[0034] Before processing, the intent analysis unit 123 requires a pre-built, predefined intent library. This intent library is the core knowledge asset of the entire system, carefully designed and maintained by domain experts in conjunction with business needs. It is stored in the form of structured data, with each entry representing an independent, identifiable intent template. An intent template contains the following information: a unique intent identifier, such as querying order status, booking a table, or modifying a reservation; one or more semantic vectors, obtained by encoding intent names or typical example sentences through a pre-trained language model for semantic similarity calculation; a set of keywords for keyword matching; and a set of context rules for adjusting its matching weight under specific context conditions. For example, the intent template for booking a table might contain semantic vectors, keywords ["book", "reservation", "appointment", "table"], and context rules {"dialogue behavior label":{"initiate new topic":1.2,"confirm":1.1}}. The first stage of the intent analysis unit 123 is to traverse the clause C of each additional context in the clause set of additional contexts. Taking the output of the previous stage as an example, the clause set contains two clause objects. For each clause C, the system will further traverse each intent template I in the predefined intent library and use a weighted scoring function. Calculate the matching score between clause C and intent template I. This scoring function comprehensively considers three dimensions: semantic similarity, keyword matching degree, and contextual relevance, and is defined as follows: ,in, , , These are preset weight coefficients, which sum to 1. These weight coefficients are determined through cross-validation and hyperparameter tuning (e.g., grid search or Bayesian optimization) on a large-scale, manually annotated training corpus. The goal is to maximize the model's intent recognition accuracy on the validation set. For example, It can be set to 0.5 (emphasizing semantic understanding). Set to 0.3 (to ensure keyword triggering). Set to 0.2 (considering contextual influences). Semantic similarity. : This item measures the semantic proximity between the clause text and the intent template. The calculation method is as follows: First, use a pre-trained semantic encoder, such as the BERT model based on the Transformer architecture, to encode the text content of clause C into a high-dimensional real number vector. This BERT model can capture the deep semantic information of the text, and its weights and biases are pre-trained on a large amount of general corpus through self-supervised learning (such as masked language model, next sentence prediction), and fine-tuned on specific tasks (such as NLI, STS). The intent template I also obtains its semantic vector through the same semantic encoder during construction. Then, use the cosine similarity formula to calculate the similarity between these two vectors, and the value ranges from -1 (completely opposite) to 1 (completely the same). For example, for clause C1 = "I want to order a meal at 18:00 tonight", the cosine similarity between its semantic vector and the semantic vector of intent template I1 = "Order reservation" may be as high as 0.95, while the similarity with the semantic vector of I2 = "Parking space query" may be only 0.1. Keyword matching degree : This item evaluates whether the text of clause C contains the triggering keywords predefined in intent template I. It can be implemented by constructing a trie tree of the Aho-Corasick string matching algorithm or directly performing fuzzy matching. If the clause text contains the key trigger words of template I, then the value is set to 1.0; otherwise it is 0. For example, for clause C1, it contains the character "订", which is in the keyword list of I1 = "Order reservation", so (C1, I1) = 1.0. And C1 does not contain any keywords of I2 = "Parking space query", so (C1, I2) = 0. Similarly, for clause C2 = "By the way, do you have a parking space?", it contains keywords such as "停车位", which matches I2 = "Parking space query", then (C2, I2) = 1.0. Context relevance : This item adjusts the intent matching score according to the current context information carried by clause C, including emotion labels and dialogue act labels. Context rules are preset in intent template I. For example, if the context rule of intent template I1 = "Order reservation" specifies that when the dialogue act label is initiating a new topic, its factor is 1.2, then when the dialogue act label of the context information of clause C1 is initiating a new topic, (C1, I1) is set to 1.2. If the intent template has no relevant context rules, or the rules do not match, then the factor is set to the default value of 1.0. For example, since the dialogue act label in the context information of C1 is initiating a new topic, and the context rule of intent I1 (order reservation) has an addition for this, then (C1, I1) = 1.2. While intent I2 (parking space query) has no specific rules, to avoid unwarranted penalties, The score is still set to 1.0. After all scoring is completed, the system selects the intent template with the highest score that exceeds the preset activation threshold as its final matching intent for each clause C. This activation threshold is a preset floating-point value, such as 0.7 or 0.8, which is iteratively optimized on the labeled dataset to balance the precision (avoiding false recognition) and recall (not missing true intent). If the scores of all intent templates are below this threshold, the clause may be marked as having no intent or requiring clarification. Taking the example in this application: for the clause C1="I want to order dinner at 18:00 tonight": Score(C1,I1="Order Reservation")=0.5×0.95+0.3×1.0+0.2×1.2=1.015. Score(C1,I2="Parking Space Inquiry")=0.5×0.1+0.3×0+0.2×1.0=0.25. Clearly, Score(C1,I1) is the highest, and its 1.015 is significantly higher than the activation threshold of 0.7. Therefore, C1 matches the intent I1="Order Booking". For the clause C2="By the way, do you have parking spaces?", Score(C2,I2="Parking Space Inquiry") = 0.5 × 0.9 + 0.3 × 1.0 + 0.2 × 1.0 = 0.95. Score(C2,I1="Order Booking") = 0.5 × 0.05 + 0.3 × 0 + 0.2 × 1.0 = 0.225. Again, Score(C2,I2) is the highest, and its 0.95 is significantly higher than the activation threshold of 0.7. Therefore, C2 matches the intent I2="Parking Space Inquiry". Finally, an intent object is generated for each successfully matched clause C. An intent object is a structured data unit that binds together the matched intent identifier, the entity list extracted from the clause, and the context information carried by the clause. For example: Intent identifier: directly uses the ID of the final matched intent. Entity list: directly populates the named entities contained in C. Context information: directly populates the current context information carried in C. Finally, all generated intent objects will be collected into a list to form the final intent set. For this embodiment, the output intent set will be: [{"Intent Identifier":"Order Booking","Entity List":[{"Entity":"Tonight 18:00","Type":"Time"},{"Entity":"Meal","Type":"Item"}],"Context Information":{"Emotion Label":"Normal","Dialogue Behavior Label":"Initiate New Topic"}},{"Intent Identifier":"Parking Space Query","Entity List": [{"Entity":"Parking Space","Type":"Facilities"}],"Context Information":{"Emotion Label":"Normal","Dialogue Behavior Label":"Initiate New Topic"}}]. This set of intents represents the final result of multi-intent parsing of the user's original speech stream.

[0035] Specifically, the logical fragment matching module 130 is used to perform matching and retrieval of the intent set based on service logical topology to obtain a logical fragment set. It should be understood that these discrete intent objects are essentially just atomic descriptions of user needs, such as booking a table or querying parking spaces. They do not inherently contain the specific execution flow consisting of a series of interconnected actions necessary to fulfill these needs. A real business process, such as a complete booking service, requires sequentially executing multiple logical steps, such as querying availability, confirming time, calling the booking interface, and processing the booking result. If the identified user intent cannot be accurately mapped to these pre-defined service flows containing inherent logical dependencies, the system will be unable to generate a coherent and orderly execution plan, leading to interruptions and failures in the interaction. Therefore, in order to transform discrete intent objects into executable action sequences containing business logic, this application performs service logic topology matching and retrieval on the intent set to accurately find the corresponding entry point for each identified intent from a pre-built logical network that depicts the complete business process, and extract the complete logical subgraph, i.e., logical fragment, required to complete the intent, thereby laying the foundation for the subsequent construction of accurate and efficient composite instructions.

[0036] Figure 4 This is a schematic diagram illustrating the data flow of the logical segment matching module in a multimodal system integrating dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning according to an embodiment of this application. Figure 4 As shown, in a feasible technical solution, the logical fragment matching module 130 includes: a first intent object extraction unit 131, used to extract a first intent object from an intent set; a logical topology search unit 132, used to search in a predefined logical topology library using the first intent object as a query condition to obtain an intent-node correlation score set; an entry node determination unit 133, used to select the node corresponding to the largest intent-node correlation score from the intent-node correlation score set as the entry node; and a depth-first search unit 134, used to perform a depth-first search along the dependency edges of the predefined logical topology library starting from the entry node to obtain the first logical fragment.

[0037] Specifically, the logical fragment matching module 130 processes the following: This module matches these intent objects with a pre-built, predefined logical topology library and extracts a corresponding logical fragment for each intent object. The predefined logical topology library is the core knowledge base of this technical solution, manually built and continuously optimized by domain experts based on specific business processes and rules. It is constructed as a directed graph GL=(N,El), where N is the set of logical unit nodes and El is the set of edges representing dependencies between nodes. Each logical unit node n∈N in the graph is a structured object containing the following key fields: Unique identifier: for example, node_check_availability (check availability node), node_confirm_time (confirmation time node), node_execute_booking_api (execute booking API call node), node_query_parking_info (query parking information node). Triggering keywords: a set of words strongly related to the function of the node, used for keyword matching. For example, the keywords for the check availability node might include ["available", "Are there any spaces available", "Can I book?"]. Node Functional Semantic Vector: A fixed-length real-valued vector representing the core function of a node, obtained by inputting its textual description or typical application scenario into a pre-trained language model (such as BERT). Context Affinity Label: A label used to improve the matching priority of a node in a specific dialogue context. For example, the affinity label of a node used to handle booking failures could be set to {"emotion label": "disappointment"}. Specific Execution Content: This is the core value of a node; it stores the specific instruction template that needs to be executed when the logical unit is triggered. This can be in the form of natural language instructions (used to generate prompts for larger models), function signatures (used to call internal or external APIs), or directly executable code snippets. For example, the content of a node executing a booking interface call might be an executable template describing "what to do" (booking a table) and "what information is needed" (time and number of people). Directed edges e∈El in the graph represent logical dependencies between nodes. For example, an edge from the availability check node to the time confirmation node indicates that an availability check must be performed before proceeding to the time confirmation step.

[0038] The processing flow of the logical fragment matching module is completed collaboratively by its four internal sub-units: The first intent object extraction unit 131 first extracts the first intent object from the input intent set in sequence as the first intent object, namely {"intent identifier": "order reservation", "entity list": [{"entity": "tonight 18:00", "type": "time"}, ...], "context information": {"dialogue behavior label": "start a new topic"}}.

[0039] The logical topology search unit 132 receives the first intent object as the query condition and performs a comprehensive search in a predefined logical topology library. It traverses all logical unit nodes n in the library and calculates the intent-node correlation score between them using a multi-dimensional weighted scoring model. The scoring function is defined as follows:

[0040] Here , , These are configurable weight coefficients that sum to 1. Their values ​​are determined through experimental tuning on an labeled validation set, aiming to maximize the accuracy of entry node selection. For example, in a scenario where semantic understanding is crucial, the following weights can be set: =0.6, =0.2, =0.2. Semantic similarity This calculation is intended for the following objects. The semantic vectors of node n are averaged or weighted averaged with their intent templates, i.e., the semantic vectors of order bookings, and then cosine similarity is calculated with the functional semantic vectors of node n. For example, The semantics of the node for checking availability are highly related to the semantics of the node. The score could be as high as 0.92, while the semantic similarity with the query parking space information node could be as low as 0.15. Keyword matching degree This calculation is intended for the following objects. The degree of overlap between the entity types contained and the triggering keywords of node n. For example, Entities containing time-related keywords will receive a higher match score (e.g., 0.8) if the keywords for usability checks include time and date. Nodes without time-related keywords will have a match score of 0. Contextual affinity enhancements. This item is based on The score is awarded based on the match between the contextual information and the affinity label of node n. For example, an availability check node might have set an affinity for the action of initiating a new topic in conversation. The dialogue behavior label is for initiating a new topic. An enhancement factor of, for example, 1.2 can be obtained; otherwise, it is 1.0. The logical topology search unit will be... Calculate scores with all nodes in the library and output an intent-node relevance score set, which is a list containing (node, score) pairs.

[0041] In particular, in multi-mode TCP protocol interaction systems, accurately mapping discrete user intent objects to predefined service execution flows (i.e., logical segments) is crucial for achieving intelligent responses. However, when calculating the matching degree of this mapping, treating multiple dimensions such as semantic information of intent, entity keywords, and dialogue context as parallel evaluation objects and performing a simple weighted sum reveals significant technical flaws in complex scenarios. For example, a logical unit completely unrelated to the user's intent, such as booking a table, such as canceling an order, might receive a non-zero score simply because its keyword list contains an entity mentioned by the user, such as an order. This can easily lead the system to retrieve the wrong execution flow, causing subsequent interactions to fail. This pseudo-matching phenomenon is an inherent problem that traditional scoring models struggle to overcome. Therefore, to fundamentally solve the intent drift problem caused by local features such as keyword overlap and ensure that the retrieved logical units are highly consistent with the user's fundamental intent in terms of core functionality, this application preferably proposes a hierarchical conditional modulation method for intent-node correlation. This method improves the scoring process from a flat, weighted summation model to a hierarchical modulation model. Its core idea is to first establish a baseline confidence score representing the core functional matching degree using semantic similarity. Then, it dynamically and multiplicatively enhances or suppresses this baseline using specific details from the user's intent (i.e., entity lists and context). This approach aims to ensure that only logical units that are semantically highly relevant to the user's core intent have the opportunity to have their scores significantly amplified by specific details, thereby eliminating erroneous retrieval caused by local feature matching at the source and greatly improving the accuracy and first-response hit rate of logical fragment matching.

[0042] Based on this, in a feasible preferred technical solution, the logical topology search unit 132 is used for:

[0043] The functional semantic vector of the first logical unit node is extracted from a predefined logical topology library. The cosine similarity between the functional semantic vector of the first intent object and the first logical unit node is calculated to obtain a basic affinity level score. This process builds a solid semantic foundation, namely, evaluating the consistency between the core function of the logical unit node and the generalized topic of the user intent. It should be understood that a user's colloquial expression, such as "reserve me a seat," may differ significantly from a standardized description of the logical unit, such as "make a table reservation." Directly comparing the two would introduce a lot of noise. Therefore, by comparing the semantic relationship between the more generalized intent template extracted from the user intent, such as "order reservation," and the logical unit, we can penetrate the differences in surface text and establish a stable and reliable core matching degree metric that is unaffected by the user's colloquial or non-standard expressions, serving as the absolute benchmark for subsequent scoring. This step first extracts the currently evaluated first logical unit node from the predefined logical topology library, such as checking the pre-stored functional semantic vector of the availability node. Simultaneously, it extracts the functional semantic vector from the input first intent object... The intent is to retrieve the semantic vector of the order booking intent template defined in a predefined intent library, which is associated with the order booking intent. Then, the cosine similarity formula is used to calculate the similarity between these two vectors, obtaining a basic affinity level score. Specifically, this calculation is the one described in the above embodiment. Preferably, because The computation relies on the semantic vectors of standardized intent template objects, rather than directly using the user's original, variable utterance vectors. This makes the model's core matching logic immune to interference from users' colloquial and non-standard expressions. As a result, the system has a better consistent response to different users' various expressions of the same core intent (such as "book a seat" and "I want to make a reservation"), greatly improving the system's generalization ability and stability.

[0044] The keyword matching degree and contextual affinity enhancement terms of the functional semantic vectors of the first intent object and the first logical unit node are calculated, and these enhancement terms are weighted to obtain a specificity enhancement level score. This process combines the base score and the moderating score in a way that reflects the principle of prioritizing the base score. Simple addition cannot achieve a veto effect. A multiplicative modulation method is used to ensure that only the base affinity score is considered as a priority. At sufficiently high levels, specificity is enhanced. Only then can the positive effects be significantly amplified. This way, even if a node's keywords and context match perfectly ( (Very high), as long as its core function does not match the user's intention ( Even if the score is very low, the final score will still be suppressed to a very low level, thus effectively filtering it. Specifically, this step comprehensively considers keyword matching and contextual affinity. First, the first intent object is calculated. The keyword match between the entity list ([{'entity':'tonight 18:00','type':'time'}]) and the first logical unit node, i.e., the preset keyword list of the availability check node (e.g., ["time","vacancy"]). Simultaneously, a context affinity enhancement term is calculated between the context information of the first intent object ({dialogue behavior label: "start a new topic"}) and the context affinity label of the first logical unit node (e.g., {affinity: "start a new topic"}). Finally, the two items are weighted and summed according to preset weights to obtain the specificity enhancement level score. .Right now Among them, weight and It is a preset value (e.g.) =0.6, =0.4, whose sum is 1), obtained through tuning on the validation dataset, used to balance the importance of entity information and contextual information. This weighted sum structure makes... This allows for the quantification of a node's readiness for a specific request. Specifically, the results for keyword matching and context affinity enhancements are the same as in the above implementation and will not be described further. By making the context's operational mechanism more rational, the system can respond more precisely to the flow of the dialogue. For example, through context affinity enhancements… With proper adjustment, the system can accurately distinguish whether a request is an initial reservation or a modification of an existing reservation, thereby selecting a more precise and suitable logical flow for the current stage of the dialogue, effectively improving the intelligence level of human-computer interaction and the smoothness of the service process.

[0045] The intent-node relevance score of the first node is obtained based on the base affinity level score and the specificity enhancement level score. Preferably, simple addition cannot achieve a veto effect. A multiplicative modulation approach is used to ensure that only the base affinity (…) score is considered. When the specificity is sufficiently high, the specificity is enhanced. Only then can the positive effects of [the keyword matching process] be significantly amplified. The effect is that even if a node's keywords and context match perfectly (…), [the positive effects will be amplified]. (Very high), as long as its core function does not match the user's intention ( Even if the score is very low, its final score will still be suppressed to a very low level, thus effectively filtering it. This step is based on the calculated... and The final intent-node relevance score of the first node is calculated using a multiplicative modulation formula. Formula: Specifically, Scenario 1 (Correct Match): Evaluate the intent order booking and check the availability of the node. As calculated... Very high, for example, 0.92, indicating a high degree of semantic matching. If positive, such as 0.6, the entity and context match well. =0.92*(1+0.6)=1.472. A very high score. Scenario 2 (Pseudo-match filtering): Evaluating the intent order booking versus the node order cancellation. The node also contains the order keyword, but the semantics are not consistent, therefore... It's very low, for example, 0.35. Even if it gains some benefit due to keyword matching... For example, 0.2. =0.35*(1+0.2)=0.42. Since... The underlying structure, acting as a foundation, resulted in a score far lower than the correctly matched node, leading to its successful exclusion. Through this hierarchical conditional modulation method, the system can find the most suitable execution entry point for each intent object in the logical topology library in a robust yet refined manner, laying a solid foundation for subsequently building a high-quality set of logical fragments. Finally, by traversing and applying the hierarchical conditional modulation calculation to all logical unit nodes in the predefined logical topology library, a complete intent-node correlation score set can be generated.

[0046] Entry node determination unit 133 receives this score set and selects the node with the highest score from it as the entry node for the intent object. For example, after calculation, the availability check node is selected based on... It achieved the highest overall score of 0.85 in the matching process, far exceeding other nodes, and was therefore identified as the entry node for the intention to book an order.

[0047] Depth-First Search (DFS) unit 134 receives this entry node, i.e., the availability check node, and begins a DFS in the logical topology graph. DFS is an algorithm used to traverse or search a tree or graph. This algorithm delves deep along a branch path of the graph until it reaches the end of that path, then backtracks and continues exploring the next unvisited branch. In this scenario, dependency edges determine the direction of traversal. 1. The search begins at the entry node, which the algorithm marks as visited and adds to the currently constructed logical segment. 2. The algorithm checks all outgoing edges (nodes that depend on it) of this node. If the availability check node has an edge pointing to the confirmation time node, and the confirmation time node has not been visited, the algorithm recursively performs a DFS on the confirmation time node. 3. The confirmation time node is added to the logical segment, and the algorithm continues to check its outgoing edges, finding an edge pointing to the node that executes the subscription interface call. 4. The node that executes the subscription interface call is added to the logical segment. If it has no outgoing edges, the path exploration is complete, and the algorithm backtracks to the confirmation time node. 5. If it is confirmed that there are no other unvisited outgoing edges at the time node, the algorithm continues backtracking to the availability check node. 6. If the availability check node has another outgoing edge pointing to `node_handle_no_availability` (the node handling the no-availability case), the algorithm will explore this branch. 7. This process continues until all reachable nodes from the entry node and their inter-node dependencies have been extracted. These extracted nodes and edges together form a connected, independent logical subgraph, i.e., the first logical segment. This segment fully depicts all the steps and logical relationships required to execute the order booking intention.

[0048] This unit stores the generated first logical fragment into an initially empty logical fragment set. Subsequently, the first intent object extraction unit extracts the next intent object from the intent set, i.e., {“intent identifier”:“parking space query”,...}. The entire process described above, from logical topology search to depth-first search, will... Repeat the process once. For example, querying parking space information nodes might be recognized. The entry node is used, and since querying a parking space is a single action, depth-first search may only extract the logical fragment containing that single node. After processing all intent objects in the intent set, the depth-first search unit finally outputs a complete set of logical fragments. This set now contains multiple logical fragments, two in this example, each corresponding to an intent in the input intent set, and depicts in detail the complete workflow required to execute that intent in the form of a subgraph.

[0049] Specifically, the composite prompt word construction module 140 is used to construct composite prompt words with a hybrid structure based on a set of logical fragments. Correspondingly, the successful operation of the logical fragment matching module finds corresponding abstract workflows containing business logic for each of the user's multiple intentions, i.e., a set of logical fragments. However, these logical fragments are merely generalized, unspecific operation templates. While they define what to do and in what order, they lack the crucial data necessary to execute these operations, such as which time to book a meal or which product's inventory to check. This crucial data exists in the user's original request and has been extracted as entities. Without binding these specific entity values ​​to the abstract logical templates, subsequent execution units cannot generate specific, operable instructions. Therefore, in order to transform the abstract business logic blueprint into a unique, immediately executable sequence of instructions specific to the current user request, a composite prompt word with a hybrid structure is constructed based on the set of logical fragments.

[0050] In a feasible technical solution, the composite prompt word construction module 140 includes: a logic fragment parameterization unit 141, used to parameterize a set of logic fragments to obtain a set of instantiated logic subgraphs; an execution graph generation unit 142, used to resolve conflicts in the set of instantiated logic subgraphs to obtain a fused execution graph; and an execution graph serialization unit 143, used to serialize the fused execution graph to obtain the composite prompt word with the hybrid structure.

[0051] Specifically, the processing of the compound prompt word construction module 140 is as follows: The processing flow of the logical fragment parameterization unit 141 begins by traversing each logical fragment in the logical fragment set. First, it processes the first logical fragment, namely the logical subgraph associated with the order booking intent. This subgraph contains a series of related logical unit nodes, such as the availability check node and the confirmation time node. For each logical unit node n in this logical fragment, the parameterization unit examines its node execution content field in detail. This field stores an instruction template to be instantiated. For example, the node execution content of the availability check node might be a function signature template: check_availability(time="{slot_time}", a function call named availability check, whose time parameter time is specified by a placeholder named time slot {slot_time}). To find a specific value for this {slot_time} placeholder, the unit traces back to the source that generated this logical fragment, namely the first intent object in the intent set - the order booking intent object, and searches its entity list. The entity list is [{'Entity':'Tonight 18:00','Type':'Time'}, {'Entity':'Meal','Type':'Item'}]. The parameterized unit searches for entities whose type matches the slot requirement. In this example, {slot_time} requires a time type value, so the unit will match {'Entity':'Tonight 18:00','Type':'Time'}. After finding a matching entity, the unit extracts its entity value "Tonight 18:00" and may perform normalization based on the objective function's requirements, such as extracting only 18:00, and then uses this specific value to replace the placeholder in the node's execution content. After this step, the instruction template for the availability check node is successfully instantiated from check_availability(time="{slot_time}") to check_availability(time="18:00"). This parameterization process is applied to all nodes in this logic segment. For example, the template for the subsequent "confirm number of diners" node might be `confirm_party_size(size="{slot_size}")`, but since the number of diners wasn't mentioned in the user's original intent, there's no entity of type "number of diners" in the corresponding entity list. In this case, the slot will remain unfilled, indicating that the system needs to actively request the missing information from the user in the subsequent execution phase. After all nodes in the first logical segment have undergone parameterization attempts, this logical subgraph containing some or all of the instantiated instructions becomes an instantiated logical subgraph. Next, the parameterization unit continues to process the second logical segment in the logical segment set, namely the one corresponding to the parking space query intent.If the fragment contains only one node, a "query parking information" node, its execution content is `query_parking_info()`, and it does not contain any parameterized slots. In this case, the parameterization unit does not need to perform any population operations, and the instruction content of the node remains unchanged. After all logical fragments in the set have completed parameterization processing, the unit finally outputs an instantiated set of logical subgraphs. Each subgraph in this set has been transformed from a generalized business process template into a logical process in a quasi-execution state, specifically for the current user's request, where the data is partially or fully ready.

[0052] In a feasible technical solution, the execution graph generation unit 142 includes: a graph fusion subunit 1421, used to perform graph fusion on the instantiated set of logical subgraphs to obtain an initial fused execution graph; a conflict identification subunit 1422, used to perform conflict identification on the initial fused execution graph to obtain a set of conflict node pairs; and a priority adjudication execution subunit 1423, used to perform priority adjudication and adjudication execution on each pair of conflict nodes in the set of conflict node pairs based on the intent priority matrix to obtain the fused execution graph.

[0053] Graph fusion subunit 1421 receives a set containing two instantiated logical subgraphs and creates an empty graph structure, namely the initial fusion execution graph. Then, it completely and without modification merges all nodes and their internal dependency edges from the two instantiated logical subgraphs into this new graph. After this step, the resulting initial fusion execution graph is structurally a simple union of the two subgraphs, which may still be two independent connected components in the graph. Next, conflict identification subunit 1422 performs a detailed examination of this initial fusion execution graph. This subunit detects whether there are logically conflicting nodes in the graph according to predefined business rules. A conflict refers to two or more nodes attempting mutually exclusive write operations on the same system state or resource. For example, creating an order and canceling the same order is a typical conflict. In this embodiment, `check_availability(time="18:00")`, an operation to check if there are available parking spaces at 18:00, and `query_parking_info()`, an operation to query parking information, operate in different business domains and are both read operations, so there is no conflict between them. In this case, the output of the conflict identification subunit, "set of conflicting node pairs," will be empty. To fully illustrate the function of this unit, another conflicting scenario is given: if the user's intent is "I want to cancel order A and then immediately replace an item in order A," this will generate two instantiated nodes, such as `node_cancel(order_id="A")` and `node_modify_item(order_id="A", new_item="B")`. The conflict identification subunit, based on its internal conflict rule base (predefined by domain experts, indicating which operation combinations are mutually exclusive), identifies these two nodes as a pair of conflicting nodes and sets `{node_cancel(order_id="A", new_item="B")`. The pair {node_cancel, node_modify_item} is added to the conflict node pair set and output. Finally, the priority adjudication execution subunit 1423 receives the conflict node pair set and is responsible for resolving these conflicts. For the current order reservation and parking space query embodiment, since the conflict node pair set is empty, this subunit does not perform any operation and directly outputs the unmodified initial fusion execution graph as the fused execution graph. In the above-mentioned conflict scenario, the processing flow of this subunit is as follows: For each pair of conflict nodes in the conflict node pair set, such as {node_cancel, node_modify_item}, it queries a pre-configured, crucial knowledge base, namely the intent priority matrix. This matrix is ​​manually constructed by business analysts and domain experts according to the inherent requirements of the business logic, defining the relative priority or forced execution order between any two potentially conflicting original intents. For example, the matrix may stipulate that the priority of canceling an order intent is higher than that of modifying an order intent, or there is a logical dependency between the two that cancels first and then modifies.After obtaining the priority relationships, the sub-unit begins to execute the adjudication. Based on a pre-defined adjudication strategy, it modifies the graph structure to eliminate conflicts. One strategy is to remove the node corresponding to the lower-priority intent and all its subsequent dependent nodes. Another more common strategy is to add a new dependency edge between two conflicting nodes to enforce their execution order. In the example of canceling and modifying an order, based on the adjudication of the intent priority matrix, the system adds a directed edge between `node_cancel` and `node_modify_item`, pointing from `node_cancel` to `node_modify_item`. This ensures that in the final execution plan, the cancellation operation is always executed before the modification operation, thus resolving the logical conflict. After resolving all conflicts, the priority adjudication execution sub-unit outputs the final merged execution graph. This graph is a single, coherent, directed acyclic graph without internal logical contradictions. It integrates the user's multiple original intents and their corresponding instantiated execution steps into a unified, executable global workflow.

[0054] The first step of the execution graph serialization unit 143 is topological sorting. It applies a topological sorting algorithm to the input fused execution graph. Topological sorting is a linear sorting algorithm for directed acyclic graphs (DAGs), resulting in a sequence of vertices where all vertices satisfy the condition that for any directed edge from vertex u to vertex v, u always appears before v in the sequence. In this embodiment, the availability check node is the starting point of the order booking process; it has no prerequisite nodes, i.e., an in-degree of 0. Similarly, the parking space query node is the starting point of the parking space query process and also has no prerequisites. Therefore, they can both serve as starting nodes for sorting. The topological sorting algorithm (e.g., Kahn's algorithm) first finds all nodes with an in-degree of 0 and places them in a queue. Then, the algorithm removes a node from the queue, adds it to the final linear sequence, and removes all its outgoing edges from the graph, i.e., decrements the in-degree of all its adjacent nodes by 1. If this process makes the in-degree of an adjacent node 0, then that adjacent node is also added to the queue. This process is repeated until all nodes in the graph have been processed. For the execution graph in this example, a possible topology sorting execution sequence is: [Query parking space information node, check availability node, confirm time node, ..., execute reservation interface call node]. This sequence ensures that the operation of querying parking spaces can be executed independently, while the availability check in the reservation process must be executed before the confirmation time. If a loop is found in the graph during the sorting process (e.g., A depends on B, and B depends on A), the topology sorting cannot be completed, indicating a logical deadlock, which will trigger the exception handling mechanism. After obtaining the linear execution sequence, the unit initializes an empty compound prompt word object. This object is a structured container with three main parts: a set of natural language instructions for storing human-oriented dialogue content, a list of executable functions for storing precise background operations, and a context wrapper for encapsulating context-aware strategies. Next, the unit enters the serialization and encapsulation stage, traversing each logical unit node n in the execution sequence. For the first node in the sequence, the query parking space information node, the unit extracts its node execution content, i.e., query_parking_info(). Since this is an explicit function call, it is appended to the list of executable functions of the compound prompt word object. The second node in the sequence, the availability check node, executes `check_availability(time="18:00")`. This is also an instantiated function call, which is appended to the list of executable functions. This process continues until the entire execution sequence has been traversed. During traversal, the unit also checks whether the execution content of each node contains other types of instructions.For example, the execution content of the "Check Availability" node might be a more complex object, such as `{'function':'check_availability(time="18:00")','nl_instruction':'Okay, let me check if there are any slots available at 6 PM tonight.','context_strategy':{'on_failure': 'suggest_alternative_time'}}`. In this case: `'check_availability(time="18:00")'` is appended to the list of executable functions. `'Okay, let me check if there are any slots available at 6 PM tonight.'` is appended to the natural language instruction set. `{'on_failure': 'suggest_alternative_time'}`, a strategy indicating alternative times to suggest if failure occurs, is integrated into the context wrapper. After traversing all nodes, the compound prompt object is constructed. The unit may also refine the aggregated natural language instruction set to make it more fluent and coherent. Finally, the compound prompt word with a mixed structure output by the execution graph serialization unit is a structured instruction package containing complete execution logic and context information. Its content may look like this: {"Natural Language Instruction Set":["Okay, I'm processing your reservation and parking space query.","First, I'll check if there are any spaces available at 6 PM tonight...","At the same time, I'll also check the restaurant's parking situation..."],"List of Executable Functions":["query_parking_info()","check_availability(time=\"18:00\")","confirm_time(time=\"18:00\")","execute_booking_api(...)"],"Context Wrapper":{"global_emotion":"positive","error_handling_strategies":{"check_availability":"suggest_alternative_time"}}}. This compound prompt describes an ordered sequence of background function executions: querying parking spaces, checking and confirming availability at 18:00, and finally making a reservation. It also encapsulates a contextual strategy that, while maintaining a positive emotional tone, requires implementing an error handling scheme that suggests an alternative time if the availability check fails.

[0055] Specifically, the execution result generation module 150 is used to submit the composite prompt word with a hybrid structure to a large language model to obtain the execution result, which includes the draft response text and real-time data. It should be understood that the composite prompt word with a hybrid structure is merely a static, descriptive set of instructions. It details what to do and how to do it, but it does not actually execute these actions, nor can it obtain dynamically changing real-time world information that can only be known after the actions are executed, such as querying the actual inventory balance at a certain moment or the return result after calling an external interface. If this meticulous plan cannot interact with the real world and perform dynamic reasoning based on the interaction results, the final response generated by the system will be based on prediction rather than fact, failing to meet the user's core needs for real-time performance and accuracy. Therefore, in order to bridge the gap between the preset plan and dynamic reality, and to transform the static execution plan into an intelligent final result containing real-world feedback, this application submits this composite prompt word to a large language model for final reasoning and execution result generation.

[0056] In a feasible technical solution, the execution result generation module 150 processes the following: The core task of this module is to drive a large language model to consume the instruction package and interact with external real-time data interfaces, such as the application programming interfaces of a restaurant's customer relationship management system, order system, and inventory system, ultimately generating a structured execution result containing a draft response text and real-time data. The core of this module is a large language model with specially enhanced capabilities. In terms of technical architecture, this language model can adopt a Transformer-based encoder-decoder architecture or a decoder-only architecture. Its core lies in the self-attention mechanism, which enables the model to weigh and integrate the information of all other words in the sequence when processing any lexical unit in the input sequence, thereby deeply understanding long-distance contextual dependencies. This LLM is not a general-purpose base model, but rather fine-tuned for targeted tool usage or function call capabilities. This fine-tuning process is performed on a specialized corpus containing a large number of (input instruction, expected function call) data pairs. Through supervised learning on this corpus, the model learns to recognize the intent to call a specific tool implied in the text and can generate corresponding function call instructions in precise grammatical format. The model's weights and bias parameters are optimized during this fine-tuning process to ensure that the function calls it generates are completely consistent with those in the gold standard data.

[0057] The specific processing flow of the execution result generation module is as follows: First, after receiving a compound prompt word with a mixed structure, the module submits it completely to a large language model that has been loaded and has function call capabilities, as the initial context for its inference. Next, the large language model begins to parse this compound prompt word. When it processes the list of executable functions, its fine-tuned function call capability is activated. The model recognizes that check_availability(time="18:00") in the list is an external function call that it is authorized to understand. At this point, the model itself does not directly execute this function, but outputs a request to call the function in a predefined, structured format (such as JSON or XML). The execution result generation module captures this request and acts as a secure and reliable intermediary. Based on the request content, it actually initiates a network call to the external real-time data interface. For example, it initiates an HTTP GET request to the API endpoint of the restaurant reservation system. After receiving this request, the external system queries its real-time database and then returns interface response data containing the current real-time status. For example, if the 18:00 slots are fully booked, the API might return JSON data like: {"status":"unavailable","slots":["18:15","18:30"]}, indicating an unavailable time slot and providing 18:15 and 18:30 as alternative time slots. The module processes the `query_parking_info()` function in the same way and obtains the API response, such as {"parking_status":"available"}. Then, the module integrates this fresh API response data and adds it as new context information to the large language model's current inference environment. At this point, the large language model's knowledge base is dynamically updated; it not only knows the user's original request and the pre-defined execution plan but also the actual results of key steps in the execution plan. Based on this, the large language model enters the comprehensive inference phase. It integrates the obtained interface response data (18:00 full, but spaces available at 18:15 and 18:30; parking spaces available) with a set of natural language instructions in compound prompts, such as "Okay, I'm processing your reservation and parking space query," and context wrappers, such as strategies for suggesting alternative times in case of failure. Based on its powerful natural language understanding and generation capabilities, the model begins end-to-end reasoning: it understands that directly fulfilling the user's 18:00 reservation plan failed, but the interface returned feasible alternatives, and the parking space query plan succeeded. Following a pre-set principle of helping users solve problems, the model decides to generate a response that both informs the user of the facts and proactively offers solutions.

[0058] Finally, the large language model generates a structured execution result. This result object contains two main parts. The first part is the draft response text, a natural and fluent response text dynamically generated based on real-time data. For example: "Hello sir, I have found that the restaurant has free parking. However, the 6 PM reservation you wish to make is full, but there are seats available at 6:15 PM and 6:30 PM. Would these times be convenient for you?" The second part is the real-time data, which contains all the information returned by the `check_availability()` and `query_parking_info()` queries after execution. For example: `{"availability_check":{"status":"unavailable","slots":["18:15","18:30"]},"parking_info":{"status": "available"}}`. However, since these are read operations, the parts describing permanent changes in system status (such as the order ID field) are currently empty. This section will only be populated with content like {"booking_id":"B12345", "status":"confirmed"} once a subsequent user confirms the booking and successfully executes the write operation.

[0059] Specifically, the audio conversion module 160 is used to convert the execution result into a synthesized audio response after the execution result has been synthesized. That is, the execution result generation module has successfully converted the user's intent into an intelligent text response draft based on real-time data. However, the entire interaction process begins with the user's voice input and should end with the system responding with natural and clear voice, thus closing the voice-centric interaction loop. If the system only remains at the text generation level, it cannot effectively convey information in purely auditory interaction scenarios such as telephone customer service, leading to a fragmented and interrupted user experience. Therefore, in order to seamlessly convert the system's internal text-based intelligent decision-making results into a final output form that users can intuitively perceive and that conforms to human communication habits, the execution result is converted into a playable synthesized audio response after final synthesis.

[0060] In a feasible technical solution, the audio conversion module 160 processes the audio as follows: The audio synthesis step is the core user-facing function of this module. It first extracts the text content of the draft response text field from the received execution result object. This text is then fed into a text-to-speech (TTS) engine. This TTS engine is a complex model based on deep neural networks, aiming to generate speech that highly approximates human pronunciation in timbre, rhythm, pauses, and emotion. In terms of technical architecture, this TTS engine can adopt an industry-leading two-stage model, such as an architecture combining the Tacotron 2 model with a HiFi-GAN vocoder. In the first stage, the Tacotron 2 model is responsible for converting the input text sequence into a Mel spectrogram. A Mel spectrogram is a two-dimensional time-frequency representation of an audio signal that mimics the non-linear auditory characteristics of the human ear on the frequency axis, enabling compact encoding of speech content and prosodic information. The Tacotron 2 itself consists of an encoder, an attention-based alignment module, and a decoder. Its encoder, a bidirectional long short-term memory network, transforms the input text character or phoneme sequence into a context-aware feature representation. The attention mechanism dynamically determines which part of the input text to focus on at each decoding step, crucial for generating natural, fluent rhythms. The decoder is an autoregressive network that generates Mel spectrograms frame by frame. In the second stage, the HiFi-GAN vocoder receives the Mel spectrogram generated by Tacotron 2 and translates it into the final, playable original audio waveform. HiFi-GAN is a model based on generative adversarial networks (GANs). Its generator is an efficient fully convolutional network that upsamples the low temporal resolution Mel spectrogram to the same temporal resolution as the original audio (e.g., 16,000 samples per second) through a series of upsampling and non-linear processing, guessing and recovering missing phase information and sonic details in the spectrogram during this process. Its discriminator is trained to distinguish between real audio waveforms and waveforms synthesized by the generator. Through this adversarial training, the generator is forced to produce increasingly indistinguishable high-fidelity audio from real recordings. The weights and bias parameters of both models were obtained through supervised learning and adversarial training on a large corpus containing tens of thousands of hours of high-quality (text, audio) data pairs recorded by professional voice actors.

[0061] When the text "Hello sir..." is input into the TTS engine, the engine generates a high-quality digital audio stream (e.g., 16-bit PCM audio) that perfectly matches the text content. This synthesized audio response is then transmitted in real time by the system through the underlying communication protocol stack (such as SIP) and played back to the user on the telephone line.

[0062] Meanwhile, the state management steps are executed in parallel, ensuring the long-term memory and robustness of the entire dialogue system. First, the real-time data portion is extracted from the execution result object. This data is the latest information obtained after interaction with external systems, such as {status:'unavailable',slots:['18:15','18:30']}. Then, it obtains the context information of the current session, a complete data structure maintained by the dialogue management system that records all states prior to this round of dialogue. The state management unit merges the new real-time data with the old context information to generate updated real-time data, which is then used to overwrite the old session context. This merging process is not merely a simple append; it may also include updating and overwriting states. For example, the old context might record that the user requested a reservation for 18:00, while the new real-time data indicates that 18:00 is unavailable, and the system has proactively provided options for 18:15 and 18:30. This information is precisely recorded in the new session context. This state management enables the system to remember every detail of the interaction, thus supporting complex, non-linear dialogues. For example, after hearing the system's suggestion, a user might say, "No, I misspoke about the time; I want tomorrow night." Because the system has already recorded the failed attempt to book 18:00 and the complete status of the alternatives provided in the updated real-time data, it can accurately understand that the user's statement is not a modification of 18:15 or 18:30, but rather a retrospective correction of the earlier core information about tonight. Therefore, the system can correctly discard the current booking process and, based on the new core information of tomorrow night, restart a complete round of intent understanding and logical orchestration, demonstrating high intelligence and robustness. Through parallel processing of audio synthesis and state management, the audio conversion module not only successfully conveys the system's intelligent decisions to the user in the most natural way, completing the interaction loop, but also, through rigorous state maintenance, is fully prepared for potentially more complex dialogue processes in the future.

[0063] In summary, a multimodal system 100 integrating dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning, based on embodiments of this application, is explained. First, by performing deep analysis of the user's continuous raw speech stream, multiple concurrent intents are accurately identified. Then, the system matches and retrieves these intents against a pre-constructed service logic topology that depicts the complete business context, dynamically orchestrating them into a set of logical fragments. Finally, these logical fragments are constructed into a hybrid structured prompt word that integrates natural language, function calls, and code instructions. Submitting this highly structured prompt word to a large language model guides it to perform deterministic reasoning and seamless cross-platform collaborative operations, thereby ensuring flexible dialogue in real-time interaction and accurate and reliable completion of complex tasks such as queries and order placement. This effectively solves the problems of illusion and insufficient robustness in traditional solutions.

Claims

1. A multimodal system integrating dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning, characterized in that, include: The user's original voice stream acquisition module is used to acquire the user's original voice stream. The speech stream multi-intent parsing module is used to perform multi-intent parsing on the user's original speech stream to obtain an intent set; The logical fragment matching module is used to perform matching and retrieval of the intent set based on service logical topology to obtain a logical fragment set. It includes: a first intent object extraction unit for extracting a first intent object from the intent set; a logical topology search unit for searching a predefined logical topology library using the first intent object as a query condition to obtain an intent-node correlation score set; an entry node determination unit for selecting the node corresponding to the largest intent-node correlation score from the intent-node correlation score set as the entry node; and a depth-first search unit for performing a depth-first search along the dependency edges of the predefined logical topology library starting from the entry node to obtain the first logical fragment. A composite prompt word construction module is used to construct composite prompt words with a hybrid structure based on a set of logical fragments. It includes: a logical fragment parameterization unit for parameterizing the set of logical fragments to obtain an instantiated set of logical subgraphs; an execution graph generation unit for resolving conflicts in the instantiated set of logical subgraphs to obtain a fused execution graph; and an execution graph serialization unit for serializing the fused execution graph to obtain the composite prompt word with the hybrid structure. The execution result generation module is used to submit compound prompt words with mixed structures to the large language model to obtain the execution result, which includes the response draft text and real-time data. The audio conversion module is used to convert the execution result into a synthesized audio response after the response synthesis.

2. The multimodal system integrating dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning as described in claim 1, characterized in that, The multi-intent parsing module for the audio stream includes: The speech stream processing unit is used to perform speech stream preprocessing and multimodal feature extraction on the user's original speech stream to obtain standardized text and acoustic feature vectors; An additional context generation unit is used to perform dependency parsing and context information fusion on standardized text, acoustic feature vectors, and dialogue history to obtain a set of clauses with additional context. The intent analysis unit is used to perform multi-intent weighted scoring and structured intent generation on a set of clauses with additional context to obtain an intent set.

3. The multimodal system integrating dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning according to claim 2, characterized in that, The voice stream processing unit includes: The speech segmentation subunit is used to input the user's original speech stream into the speech activity detection module to obtain speech segments; The text acoustic feature extraction subunit is used to input speech segments into the speech recognition engine and the acoustic feature extractor respectively to obtain standardized text and acoustic feature vectors.

4. The multimodal system integrating dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning according to claim 2, characterized in that, The additional context generation unit includes: The dependency parsing graph generation subunit is used to input standardized text into the dependency parsing model to obtain a dependency parsing graph; Semantic core segmentation subunits are used to perform semantic core segmentation on dependency parsing graphs to obtain clause sets; The current context information generation subunit is used to perform dual-channel context feature vectorization on standardized text, acoustic feature vectors, and dialogue history to obtain the current context information. The encapsulation subunit is used to perform entity recognition on the clause set, and then perform context fusion encapsulation with the current context information to obtain a clause set with additional context.

5. The multimodal system integrating dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning according to claim 1, characterized in that, The logical topology search unit is used for: Extract the functional semantic vector of the first logical unit node from the predefined logical topology library; Calculate the cosine similarity between the functional semantic vectors of the first intent object and the first logical unit node to obtain the basic affinity level score; Calculate the keyword matching degree and context affinity enhancement term of the functional semantic vector of the first intent object and the first logical unit node, and weight the keyword matching degree and context affinity enhancement term to obtain the specificity enhancement level score; The intent-node relevance score of the first node is obtained based on the basic affinity level score and the specificity enhancement level score.

6. The multimodal system integrating dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning according to claim 1, characterized in that, The logic segment parameterization unit is used for: Traverse the logical unit nodes in the logical fragment set and identify the instruction templates to be instantiated and parameter placeholders in the node execution content; Search the entity list that matches the parameter placeholder type in the list of entities that generated the intent object corresponding to the logical fragment; Replace the parameter placeholders with the values ​​of the matched entities to convert the instruction template into an instantiated instruction.

7. The multimodal system integrating dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning according to claim 1, characterized in that, The execution graph generation unit includes: The graph fusion subunit is used to perform graph fusion on the instantiated set of logical subgraphs to obtain the initial fusion execution graph. The conflict identification subunit is used to identify conflicts in the initial fusion execution graph to obtain a set of conflict node pairs. The priority adjudication execution subunit is used to perform priority adjudication and adjudication execution on each pair of conflicting nodes in the conflicting node pair set based on the intent priority matrix to obtain the fused execution graph.

8. The multimodal system integrating dynamic semantic orchestration and cross-platform intelligent agent collaborative reasoning according to claim 1, characterized in that, The execution graph serialization unit is used for: A topological sorting algorithm is applied to the merged execution graph to generate a linear execution sequence; Initialize the compound prompt word object, which contains a set of natural language instructions, a list of executable functions, and a context wrapper; It iterates through the logical unit nodes in the execution sequence, appends the function calls in the node execution content to the executable function list, appends the natural language description to the natural language instruction set, and encapsulates the strategy configuration into the context wrapper.