Multimodal semantic network-driven intelligent agent context understanding method and system
By constructing a multimodal semantic network, receiving and encoding image, voice and text data, generating and annotating semantic fragments, and forming a contextual semantic network, the problem of insufficient semantic generalization ability in the intelligent agent's context understanding is solved, and efficient parsing and accurate response to complex instructions are achieved.
Patent Information
- Application Number
- CN202510767913.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing technologies have problems with insufficient semantic generalization capabilities and insufficient context modeling in agent context understanding, which leads to task execution deviation or failure, especially when processing ambiguous instructions containing temporal and historical space references.
By constructing a multimodal semantic network, the system receives raw interactive data from images, speech, and text, performs modal encoding, generates a sequence of semantic segments, and performs semantic label expansion and coreference resolution to form a set of contextual semantic segments. These segments are converted into semantic nodes, with initial memory weights and time labels set. Directed connections are generated based on temporal relationships and semantic associations to construct a contextual semantic network. Based on user instructions, the system extracts temporal, spatial, and target entity words, matches semantic nodes in the network, calculates a comprehensive matching score, dynamically updates memory weights, generates contextual semantic paths, and outputs structured information.
It significantly improves the intelligent agent's ability to parse the context of complex instructions in multi-round tasks, enhances its ability to reason about contextual semantics, improves the accuracy and coherence of instruction responses, and solves the problem of task failure caused by weak semantic generalization ability and insufficient context modeling in existing technologies.
Smart Images

Figure CN120297287B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for intelligent agent context understanding driven by a multimodal semantic network. Background Art
[0002] In existing technologies, semantic parsing methods based on single-modal or multimodal fusion are often used to enhance an intelligent agent's ability to understand its environment. Multimodal methods extract features from multiple sources of data, such as images, speech, and text, and use attention mechanisms or graph neural networks to build inter-modal associations to form a unified semantic representation. In intelligent agent systems, this semantic representation serves as the basis for perception and reasoning, enabling task instruction understanding and execution. However, most of these methods rely on static semantic embedding models, making it difficult to dynamically adapt to contextual changes, resulting in insufficient semantic generalization capabilities.
[0003] For example, when a service robot executes cross-room commands, for example, when a user issues a command like "Go to the kitchen and get the medicine I just put on the table," existing multimodal fusion technologies struggle to effectively integrate temporal and spatial contextual information such as "just now," "table," and "medicine" during semantic reasoning. Because current systems lack a dynamic modeling mechanism for conversational context, they often interpret "table" as the most relevant object in the current location, rather than its historical context, leading to execution path deviations or task failures. This lack of contextual understanding is particularly prominent in scenarios where tasks are continuously interactive, hindering the agent's ability to reason coherently. Summary of the Invention
[0004] The purpose of the present invention is to provide a multimodal semantic network driven intelligent agent context understanding method and system, aiming to solve the problems mentioned in the background technology.
[0005] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0006] In a first aspect, a multimodal semantic network-driven agent context understanding method, the method comprising:
[0007] Receive raw interaction data containing images, voice, and text, perform modal encoding on each, extract features of each modality, and then synchronize and align the temporal information to generate a sequence of semantic segments;
[0008] Perform semantic label expansion and coreference resolution on the semantic fragment sequence, annotate the time anchor point and coreference target of each semantic fragment, and obtain the contextual semantic fragment set;
[0009] Based on the set of contextual semantic fragments, each contextual semantic fragment is converted into a semantic node, and the initial memory weight, time label and source modality are set. Based on the temporal relationship and semantic association between semantic nodes, directed connections are generated to construct a contextual semantic network.
[0010] Based on user instructions, extract the time words, space words and target entity words, match them with the relevant semantic nodes in the context semantic network, and calculate the comprehensive matching score;
[0011] Compare the comprehensive matching score with the preset comprehensive matching score threshold, screen the candidate semantic nodes, and generate a candidate semantic node set;
[0012] Based on the candidate semantic node set, the current memory weight is dynamically updated through the comprehensive matching score of the semantic node to form an updated contextual semantic network;
[0013] According to the updated contextual semantic network, the semantic node sequence with the highest current memory weight is extracted as the contextual semantic path, and the semantic entities in the path are structured and output.
[0014] Preferably, semantic tag expansion and coreference resolution are performed on the semantic segment sequence, and the time anchor point and coreference target of each semantic segment are marked to obtain a context semantic segment set, including:
[0015] According to the sequence of semantic segments, the temporal expression in each semantic segment is identified, a mapping relationship between the segment and the standard time axis is established, and a temporal anchor point label is generated.
[0016] According to the sequence of semantic segments, the pronouns in each semantic segment are identified to determine the candidate referent entities;
[0017] Based on temporal proximity, semantic relevance, and modal confidence scores, the reference confidence of each candidate referent entity is calculated, and the candidate referent entity with the highest reference confidence is determined as the referent target of the referent;
[0018] The referent target and time anchor annotations are attached to the corresponding semantic segments to form contextual semantic segments with time anchor labels and referent target labels, and a contextual semantic segment set is obtained.
[0019] Preferably, generating directed connections based on the temporal relationship and semantic association between semantic nodes to construct a contextual semantic network includes:
[0020] According to the time label of each semantic node, it will be sorted in chronological order to form a time-ordered node sequence;
[0021] According to the time-ordered node sequence, for any two semantic nodes, the degree of semantic association is calculated by comparing the directional consistency and content similarity between their semantic features;
[0022] When the semantic correlation between two semantic nodes is greater than the preset correlation threshold and the former semantic node is earlier than the latter semantic node in time, a unidirectional connection is established between them, with the direction pointing to the later semantic node in time, forming a contextual semantic network.
[0023] Preferably, based on the user instruction, the time words, space words and target entity words are extracted, and the relevant semantic nodes in the context semantic network are matched, and the comprehensive matching score is calculated, including:
[0024] Perform lexical item division and semantic classification on user instructions, extract time expression words, spatial location words and target entity words, and generate a lexical item set;
[0025] According to the word set, the three categories of words extracted are semantically matched with the elements of the corresponding categories in each semantic node in the context semantic network, and the similarity of the three categories of word matches is calculated to obtain a similarity set;
[0026] Based on the similarity set, the three types of similarity results are weighted according to the preset proportional coefficient to form a comprehensive matching score.
[0027] Preferably, based on the candidate semantic node set, the current memory weight of the semantic node is dynamically updated by the comprehensive matching score of the semantic node to form an updated contextual semantic network, including:
[0028] According to the candidate semantic node set, obtain the current memory weight value and comprehensive matching score value of each candidate semantic node;
[0029] Multiply the current memory weight value by the first update coefficient, multiply the matching score value by the second update coefficient, and then sum the two to obtain the new memory weight of the candidate semantic node;
[0030] According to all non-candidate semantic nodes, the time difference between their time labels and the current time is calculated, and the time difference is input into the preset time decrement function to generate the attenuation coefficient;
[0031] Multiply the current memory weight of the non-candidate node by the decay coefficient to obtain the new memory weight of the non-candidate node;
[0032] The new memory weights of all semantic nodes are written into the corresponding semantic nodes to form an updated contextual semantic network. The updated memory weights are used for subsequent contextual semantic path extraction.
[0033] Preferably, the reference confidence of each candidate referent entity is calculated based on temporal proximity, semantic relevance, and modal confidence score, including:
[0034] Calculate the distance between each candidate referent entity and the referent word on the time axis according to their appearance positions in the semantic segment sequence, and convert it into a temporal neighbor score, where the temporal neighbor score decreases as the distance increases;
[0035] Compare the semantic content between the candidate referent entity and the referent word, extract their semantic word embedding information or context word vector, calculate their matching degree in the semantic dimension, and generate a semantic relevance score;
[0036] Identify the modal type from which the candidate referent entity originates, match it based on the preset modal priority table, and generate a modal confidence score;
[0037] The temporal neighbor score, semantic relevance score and modal confidence score are weighted respectively, and weighted superposition is performed to obtain the reference confidence of the candidate referent entity;
[0038] The candidate entity with the highest reference confidence is used as the reference target of the pronoun.
[0039] Preferably, according to the time-ordered node sequence, for any two semantic nodes, the degree of semantic association is calculated by comparing the directional consistency and content similarity between their semantic features, including:
[0040] According to the time-ordered node sequence, for any two semantic nodes, the degree of semantic association is calculated by comparing the directional consistency and content similarity between their semantic features, including:
[0041] According to the time-ordered node sequence, the semantic direction features and semantic content features are extracted from the semantic nodes; the following processing is performed on any two semantic nodes:
[0042] Compare the semantic directional features, map them into directional vectors, and then calculate the angle or cosine similarity between the two directional vectors as the directional consistency score, with a value range of 0 to 1;
[0043] Compare semantic content features and calculate the semantic co-occurrence degree between the referent entities contained in the two semantic nodes, the matching degree of contextual keywords, and the consistency of the contextual background to form a semantic content similarity score;
[0044] Multiply the direction consistency score and the semantic content similarity score by their respective contribution coefficients, and then sum them up to obtain the semantic association degree between the two semantic nodes;
[0045] When the degree of semantic association is higher than the set association threshold and the node time sequence satisfies the order from early to late, a unidirectional connection relationship is established between the two nodes from the earlier node to the later node as part of the contextual semantic network structure.
[0046] In the second aspect, a multimodal semantic network-driven intelligent agent context understanding system, the system includes:
[0047] The data acquisition and processing module is used to receive raw interactive data including images, voice, and text, perform modal encoding on each, extract features of each modality, and then synchronize and align the temporal information to generate a sequence of semantic segments;
[0048] The semantic annotation module is used to perform semantic label expansion and coreference resolution on the semantic segment sequence, annotating the time anchor point and coreference target of each semantic segment to obtain a contextual semantic segment set;
[0049] The node construction module is used to convert each contextual semantic fragment into a semantic node based on the contextual semantic fragment set, set the initial memory weight, time label and source modality, generate directed connections based on the temporal relationship and semantic association between semantic nodes, and construct a contextual semantic network;
[0050] The instruction parsing module is used to extract the time words, space words and target entity words based on the user instruction, match the relevant semantic nodes in the context semantic network, and calculate the comprehensive matching score;
[0051] A candidate node screening module is used to compare the comprehensive matching score with a preset comprehensive matching score threshold, screen candidate semantic nodes, and generate a candidate semantic node set;
[0052] The weight update module is used to dynamically update the current memory weight of the candidate semantic node set based on the comprehensive matching score of the semantic node to form an updated contextual semantic network, in which the semantic nodes with high comprehensive matching scores are given higher weights, and the semantic nodes further away from the current moment are reduced in weight according to the attenuation coefficient;
[0053] The structured output module is used to extract the semantic node sequence with the highest current memory weight as the context semantic path based on the updated context semantic network, and structure the output of the semantic entities in the path for intelligent agent task understanding and response generation.
[0054] The above solution of the present invention includes at least the following beneficial effects:
[0055] By constructing a multimodal semantic network and introducing a dynamic contextual semantic understanding mechanism, this invention significantly improves the agent's ability to contextually parse complex instructions across multiple rounds of tasks. Unlike existing technical solutions that rely on static semantic embedding models, this invention modally encodes and temporally aligns the raw interaction data of images, speech, and text to form a time-ordered sequence of semantic fragments. These fragments are then further structured into semantic nodes containing time labels and memory weights, generating directed connections with temporal causal directions between the semantic nodes, thereby forming a dynamically evolving contextual semantic network.
[0056] The present invention also comprehensively matches the time words, space words and target entity words in the user's current instruction with the nodes in the semantic network, combines the instruction semantics with the context history for joint calculation, screens out a set of candidate semantic nodes with high correlation, and dynamically adjusts the node memory weights. In particular, the present invention introduces a matching enhancement mechanism and a time decay mechanism in the weight update process, so that the system can strengthen the current semantic focus and gradually dilute historical irrelevant information, thereby achieving dynamic focusing on semantic information. This processing strategy effectively avoids the task execution deviation caused by "reference dislocation" or "semantic jump" in the prior art, and is particularly suitable for processing ambiguous instructions containing temporal and historical spatial references such as "the medicine just placed on the table".
[0057] Therefore, the present invention can significantly improve the modeling ability of multimodal complex semantics and the reasoning ability of contextual semantics in scenarios such as service robots, intelligent voice assistants, and multi-round question-and-answer systems, enhance the accuracy and coherence of command responses, and solve the problem of task failure caused by weak semantic generalization ability and insufficient context modeling in the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a flowchart of the multimodal semantic network-driven intelligent agent context understanding method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0059] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0060] like Figure 1 As shown, an embodiment of the present invention proposes a multimodal semantic network driven agent context understanding method, the method comprising:
[0061] S100, receiving raw interaction data including images, voice, and text, performing modal encoding on each, extracting features of each modality, and then synchronously aligning temporal information to generate a sequence of semantic segments;
[0062] S200, performing semantic label expansion and coreference resolution processing on the semantic segment sequence, marking the time anchor point and coreference target of each semantic segment, and obtaining a context semantic segment set;
[0063] S300: Based on the context semantic segment set, each context semantic segment is converted into a semantic node, and an initial memory weight, time label, and source modality are set. Directed connections are generated based on the temporal relationship and semantic association between the semantic nodes to construct a context semantic network.
[0064] S400, based on the user instruction, extracting the time words, space words and target entity words, matching the relevant semantic nodes in the context semantic network, and calculating the comprehensive matching score;
[0065] S500, comparing the comprehensive matching score with a preset comprehensive matching score threshold, screening candidate semantic nodes, and generating a candidate semantic node set;
[0066] S600: Based on the candidate semantic node set, dynamically update the current memory weight of the semantic node by the comprehensive matching score of the semantic node to form an updated contextual semantic network, wherein the semantic node with a high comprehensive matching score has its weight increased, and the semantic node further away from the current moment has its weight reduced according to the attenuation coefficient;
[0067] S700. According to the updated context semantic network, the semantic node sequence with the highest current memory weight is extracted as the context semantic path, and the semantic entities in the path are structured and output for intelligent agent task understanding and response generation.
[0068] In an embodiment of the present invention, the overall process includes four stages: multimodal perception input, context modeling, semantic matching reasoning, and structured output. The system first receives raw interaction data including images, voice, and text. These data sources may include the user's verbal commands to the intelligent agent, the captured image content, and the input instruction text. These three types of data are input into the corresponding modal encoders respectively. The image data is extracted from the spatial layout and target entity features through the visual encoding module. The voice data is converted into a time-stamped text stream through the voice recognition and semantic extraction process, and the text data is directly semantically embedded. After encoding, the system uniformly aligns the time information of each modality, synchronously maps the data from different times to the standard timeline, and generates a series of semantic fragment sequences sorted by time for subsequent context understanding.
[0069] Based on the sequence of semantic fragments, the system further performs semantic label expansion and reference resolution. Each fragment is identified to see whether it contains time expressions and whether there are pronouns, such as "it", "there", "then", etc. The system binds time anchors to fragments containing time words, points pronouns to semantic entities that have appeared in the previous text, and completes label annotations to form a set of contextual semantic fragments. Each semantic fragment is converted into a semantic node as an independent node and is assigned an initial memory weight, a time label, and modal source information. The system establishes directed connections between nodes by judging the temporal order and semantic content similarity between nodes, thereby constructing a dynamically evolving contextual semantic network.
[0070] When a user issues a new interactive command, the system extracts time terms (such as "just now" and "afterwards"), spatial terms (such as "kitchen" and "next to the table"), and target entity terms (such as "medicine" and "key"). These terms are fed into the matching module as a keyword set and compared with each node in the semantic network. By scoring and weighting the multi-dimensional matching degree of time, space, and target terms, a comprehensive matching score is obtained for each semantic node and the current command. The system compares all scores with a set matching threshold to select a set of candidate semantic nodes.
[0071] For these candidate nodes, the system dynamically updates their memory weights based on their comprehensive matching scores, with those with higher scores receiving higher weights. Nodes whose time tags are farther from the current time have their memory weakened according to a preset attenuation coefficient. Non-candidate nodes only undergo attenuation to reduce interference. After completing the weight update, the system identifies a group of semantic nodes with the highest memory weights in the semantic network as the contextual semantic path in the current interaction. The semantic entities contained in this path will be structured and output for the agent to generate the final execution plan or response statement, such as determining the specific item to be retrieved and its spatial location, thereby achieving context-consistent agent understanding capabilities.
[0072] The process involves receiving raw interactive data including images, voice, and text, performing modal encoding on each, extracting features from each modality, and then aligning temporal information to generate a sequence of semantic segments. This includes:
[0073] During the operation of an actual intelligent system, users may use voice commands, video images captured by the camera, and text input on the interactive terminal to form a complete round of command interaction. To effectively analyze this heterogeneous data, this method first separates the raw interaction data into three modal channels: image data, voice data, and text data. Each type of data is fed into the corresponding modal encoder through an independent preprocessing module. For example, image data can be used to extract visual object features and their spatial position relationships through a convolutional neural network; voice data is first transcribed into text through a speech recognition module, and then a semantic embedding vector is extracted, while the time of occurrence is recorded; text data is directly extracted through a text embedding model to extract a high-dimensional vector representing its semantic structure.
[0074] After the above encoding operation, each modal data is still in its own original time track. In order to unify the modeling of this information, this method implements modal alignment processing through a time synchronization module. Specifically, the system uses a unified time axis as a reference and aligns data segments that are close in time according to the receiving timestamps or frame marks of each modal data. For example, the image frame corresponding to a certain voice phrase "Put the medicine on the table" is a scene of the table being photographed, and the text input may be "The medicine has been put away." The system aligns the three according to their triggering time points (such as 10.2 seconds for voice, 10.5 seconds for image, and 10.4 seconds for text) to a synchronized time point segment and divides it into a "semantic segment."
[0075] In this way, the system forms a chronological sequence of semantic fragments, each of which integrates semantic information from different modalities and serves as the fundamental unit for subsequent context mapping, instruction parsing, and semantic reasoning. This process not only preserves modality specificity but also addresses the understanding bias caused by modal asynchrony in existing technologies, providing the system with a temporally consistent foundation for cross-modal semantic understanding.
[0076] Among them, according to the updated context semantic network, the semantic node sequence with the highest current memory weight is extracted as the context semantic path, and the semantic entities in the path are structured and output, specifically including:
[0077] The contextual semantic network is a graph structure consisting of a set of semantic nodes and directed connections between them. Each node represents a semantic segment with a labeled time anchor and a referential target, and has a dynamically changing memory weight. This memory weight reflects the node's activity in the current interaction context. A higher weight indicates that the node is more likely to be semantically related to the current user instruction.
[0078] After the contextual semantic network completes a round of memory weight updates, the system sorts all semantic nodes from high to low according to their memory weight. After sorting, the system starts from the node with the current highest score and traces back to its adjacent nodes in chronological order based on the directed connection relationship, recursively forming a semantic link layer by layer. To avoid path deviation, the system also considers whether the memory weight remains above the preset threshold when selecting adjacent nodes. If the weight of a node is significantly lower than the threshold, the path backtracking is interrupted to ensure that the path only contains content that is strongly related to the current semantics.
[0079] This path ultimately forms a contextual semantic path from the remote semantic history to the current semantic focus. The system extracts semantic entities from each semantic node in the path, including participants, action terms, spatial locations, time representations, and other content, and outputs them according to the semantic label structure. For example, if the path contains the two nodes "put the medicine on the table 10 minutes ago" and "the kitchen light is on", the system will extract information such as "medicine" (physical object), "table" (spatial location), and "10 minutes ago" (time expression), and output it in the form of a structured dictionary or key-value pair, such as: {"object":"medicine","location":"table","time":"10 minutes ago"}.
[0080] This structured output allows the agent to perform actions based on clear semantic entities, such as navigating to the "table," recognizing the specific appearance of "medicine," and performing the "retrieve" action. Furthermore, this structured semantic information is easily accessible to other modules. For example, the natural language generation module can convert it into natural language feedback such as "I'm going to the kitchen to get medicine," thereby improving the system's human-computer interaction performance and the accuracy of its semantic responses.
[0081] In a preferred embodiment of the present invention, semantic tag expansion and coreference resolution are performed on the semantic segment sequence, and the time anchor point and coreference target of each semantic segment are marked to obtain a context semantic segment set, including:
[0082] According to the sequence of semantic segments, the temporal expression in each semantic segment is identified, a mapping relationship between the segment and the standard time axis is established, and a temporal anchor point label is generated.
[0083] According to the sequence of semantic segments, the pronouns in each semantic segment are identified to determine the candidate referent entities;
[0084] Based on temporal proximity, semantic relevance, and modal confidence scores, the reference confidence of each candidate referent entity is calculated, and the candidate referent entity with the highest reference confidence is determined as the referent target of the referent;
[0085] The referent target and time anchor annotations are attached to the corresponding semantic segments to form contextual semantic segments with time anchor labels and referent target labels, and a contextual semantic segment set is obtained.
[0086] In this embodiment of the present invention, after encoding and temporally aligning the multimodal input data, a sequence of semantic segments with a temporal order is generated. To further enhance the alignment of the semantic segments with the real-world context, the system performs semantic label expansion and coreference resolution on this sequence, enabling the agent to effectively restore the real-world semantic context when faced with complex, ambiguous, or omitted user language.
[0087] Specifically, the system first identifies whether each semantic segment contains time-expressing elements, such as keywords like "last night," "ten minutes later," and "now." For each segment containing a time-expressing element, the system maps it to a unified timeline and generates a corresponding time anchor tag to mark the moment when the semantic meaning occurred. This time anchor provides a positioning reference for subsequent temporal relationship determination and path tracing.
[0088] Next, the system identifies and resolves pronouns, mainly targeting semantically ambiguous phrases such as "he," "it," and "that place." The system first determines temporal proximity based on the position of the fragment, calculates the distance between the pronoun and the candidate entity in the time series, and reflects this distance as a temporal neighbor score. Secondly, the system compares the semantic information of the pronoun and the candidate entity, such as the semantic distance between word embedding vectors, co-occurrence context, etc., and calculates the semantic relevance score. At the same time, considering that entities may come from different modalities, the system assigns different modal confidence scores according to the source of the entity (image, voice, or text) through a table lookup, reflecting its information reliability.
[0089] The system combines the above three scores according to a preset weighted ratio to form the final reference confidence. For each referential word, the one with the highest confidence score is selected from all candidate entities as the referential target and bound to the fragment. Ultimately, a set of contextual semantic fragments containing time anchor points and referential target labels is formed, providing high-precision, context-consistent semantic units for subsequent semantic node graph construction and matching processing. This step greatly improves the agent's ability to understand omitted instructions, skip instructions, or anaphoric expressions, and has a significant effect in multi-round dialogue interactions and scene memory tasks.
[0090] Among them, according to the semantic segment sequence, the time expression in each semantic segment is identified, a mapping relationship between it and the standard time axis is established, and a time anchor point label is generated, which specifically includes:
[0091] Semantic fragment sequences consist of time-aligned data units from multimodal input data (images, speech, and text). Each semantic fragment typically corresponds to a specific user expression or environmental event. Because semantic fragments often include descriptions of time, such as "just now," "ten minutes later," "two o'clock in the afternoon," or "yesterday," time expressions in natural language are diverse and implicit, necessitating standardized processing through a time parsing mechanism.
[0092] First, the system scans the text components in the semantic segment based on a time expression recognition model to identify whether it contains terms or phrases indicating time. This model can be a rule-based recognizer (such as a regular expression matching "X o'clock in the morning," "yesterday," etc.) or a semantic model-based classifier to detect ambiguous expressions such as "just now," "afterwards," and "at that time."
[0093] After recognizing time expressions, the system needs to map these natural language time markers to a unified standard timeline. This standard timeline can be defined as the system time at which the agent receives input, expressed as either absolute time (e.g., "May 27, 2025, 14:30") or relative time offset (e.g., "current time minus 10 minutes"). For example, if the current system time is 14:30, the mapping result is 14:20, corresponding to a specific anchor point on the standard timeline.
[0094] The mapping process needs to account for contextual shifts. For example, a term like "at that time" needs to be inferred from the time tags of previous semantic nodes to refer to the time point. This can be achieved by introducing a time-backtracking module that compares the most recent explicit time tag in the context segment. If the previous explicit time for "at that time" was "10:00 AM," the system can map "at that time" to the time period corresponding to 10:00 AM.
[0095] Finally, the system adds this mapping result as a structured time anchor label for the current semantic segment. This label can be a key-value pair, such as {"time_anchor":"2025-05-27 14:20"}, or it can be added as a timestamp field to the semantic segment data structure. This approach ensures that the semantic network constructed subsequently has a time basis that is sortable and reasonable, resolving the existing problem of a lack of alignment between semantic temporal representation and system understanding.
[0096] Among them, according to the sequence of semantic fragments, the pronouns in each semantic fragment are identified to determine the candidate referent entities, which specifically includes:
[0097] During continuous interactions or multi-round conversations, users often use numerous pronouns to replace entity names that appeared previously, such as "it," "he," "there," "then," "that thing," and so on. If these pronouns aren't correctly parsed, they can lead to semantic fragmentation, seriously impacting the construction of semantic networks and the correct execution of commands. Therefore, the present invention incorporates pronoun recognition and candidate entity screening mechanisms into the semantic fragment processing process to recover the real objects referred to by the original semantics.
[0098] First, the system performs part-of-speech tagging and dependency syntactic analysis on the text within the semantic segment, identifying lexical items with referential grammatical functions. This process can be implemented using natural language processing tools, such as conditional random field models or sequence labeling networks (such as BiLSTM-CRF), to identify pronoun-like items and, based on dependency structures, determine whether they belong to key positions such as subject, object, or adverbial. For example, in the sentence "Put it on the table," "it" is identified as a third-person referent in the object position.
[0099] After identifying the pronoun, the system needs to search for possible candidate entities in the historical part of the semantic segment sequence based on the time anchor of the current semantic segment. Candidate entities must meet two prerequisites: (1) appear in the semantic segment before the pronoun; (2) be semantically referable, that is, belong to a perceptible object or place. The system traverses the entity information in the historical semantic segments, including object nouns (such as "medicine", "book", "key") and place names (such as "kitchen", "table"), as a list of candidate pronoun entities.
[0100] After the candidate entity is determined, the system further calculates its relevance to the pronoun through the following dimensions: (1) Temporal proximity: If the candidate entity is closer to the temporal segment of the pronoun, it is more likely to be selected as the referent target; (2) Syntactic pointing consistency: If the candidate entity and the pronoun play the same grammatical role in the context, it indicates high semantic consistency; (3) Modal consistency: If the two entities come from the same modality (such as both are obtained by image detection), the confidence level is improved.
[0101] Combining the above scoring dimensions, the system assigns a correlation score to each candidate entity and selects the entity with the highest score as the final reference target. This score is recorded in the current semantic segment for subsequent semantic graph construction. For example, if the sentence in the current semantic segment is "Please bring it over here" and the previous segment is "I put the medicine on the table," then "it" will be judged to refer to "medicine." The system will add a reference annotation to the current segment: {"coreference":"medicine"}.
[0102] Through this processing step, the system can effectively solve the problems of entity omission and reference that frequently occur in conversations, achieve entity tracking and context consistency across semantic segments, enhance the ability to handle multiple rounds of fuzzy instructions, and significantly improve the semantic reasoning stability of intelligent agents in actual scenarios.
[0103] In a preferred embodiment of the present invention, generating directed connections based on the temporal relationship and semantic association between semantic nodes to construct a contextual semantic network includes:
[0104] According to the time label of each semantic node, it will be sorted in chronological order to form a time-ordered node sequence;
[0105] According to the time-ordered node sequence, for any two semantic nodes, the degree of semantic association is calculated by comparing the directional consistency and content similarity between their semantic features;
[0106] When the semantic correlation between two semantic nodes is greater than the preset correlation threshold and the former semantic node is earlier than the latter semantic node in time, a unidirectional connection is established between them, with the direction pointing to the later semantic node in time, forming a contextual semantic network.
[0107] In this embodiment of the present invention, to achieve coherent reasoning of contextual semantics, the semantic segments in the contextual semantic segment set must be constructed as semantic nodes, and a reasonable connection structure must be established between the semantic nodes. To ensure that the semantic transmission in the network structure has temporal logic and semantic consistency, the system generates directed connections based on the temporal relationships and semantic associations between semantic nodes, thus forming a contextual semantic network.
[0108] First, the system sorts all nodes chronologically based on the time tag information of each semantic node, forming a temporally ordered node sequence. This sequence not only defines the directionality of connections but also serves as an important reference for path tracing during reasoning. The system ensures that connections are made only between temporally adjacent or subsequent nodes, preventing semantic chain instability caused by temporal dislocations.
[0109] Within the sorted, time-ordered sequence of nodes, the system analyzes any two semantic nodes, first comparing whether the semantic features they represent are directionally consistent. Directional consistency refers to whether two nodes point in the same direction of semantic behavior development. For example, there is continuity in the direction of semantic migration between "walking out of the room" and "entering the hallway." The system uses verb analysis and semantic scene structure recognition to determine whether there is a continuation or connection in the action trends between the two nodes. If there is consistency, the system proceeds to the next comparison step.
[0110] For pairs of nodes with directional consistency, the system further compares their semantic content for similarity, primarily focusing on entity type, keyword overlap, and scene context consistency. For example, if "the medicine on the table" and "the things I just put down" share a location description and similar entities, the system will determine that they have a semantically similar relationship. Through this comparison, the system can determine the degree of semantic connection between the two nodes.
[0111] Once the degree of association between two semantic nodes exceeds a set threshold, and the temporal order satisfies the requirement that the preceding node precedes the succeeding node, the system establishes a directed connection between them. This connection is directed from the earlier node to the later node, ensuring a temporal causal basis for contextual reasoning. The resulting contextual semantic network not only expresses the physical connections between semantic nodes but also embodies the semantic chain structure of behavioral evolution, facilitating subsequent semantic path screening and behavioral prediction.
[0112] In a preferred embodiment of the present invention, based on user instructions, time words, space words and target entity words are extracted, and relevant semantic nodes in the context semantic network are matched, and a comprehensive matching score is calculated, including:
[0113] Perform lexical item division and semantic classification on user instructions, extract time expression words, spatial location words and target entity words, and generate a lexical item set;
[0114] According to the word set, the three categories of words extracted are semantically matched with the elements of the corresponding categories in each semantic node in the context semantic network, and the similarity of the three categories of word matches is calculated to obtain a similarity set;
[0115] Based on the similarity set, the three types of similarity results are weighted according to the preset proportional coefficient to form a comprehensive matching score.
[0116] In an embodiment of the present invention, after receiving a user interaction instruction, the system needs to extract instruction keywords in order to understand the user's intention and locate the semantic node corresponding to the instruction in the context, and match the context semantic network based on these keywords, so as to calculate the comprehensive matching score between each semantic node and the instruction.
[0117] Natural language commands typically contain temporal elements, spatial components, and target entities. The system first performs word segmentation and part-of-speech tagging on the input commands. Using a semantic classification mechanism, it extracts temporal terms (e.g., "just now," "last night"), spatial terms (e.g., "kitchen," "behind the sofa"), and target entity terms (e.g., "medicine bottle," "remote control"), unifying these three categories into a single term set to drive the subsequent semantic matching module.
[0118] During the matching phase, the system traverses each semantic node in the contextual semantic network and extracts the time tag, spatial location description, and target entity tag contained in the node. The system performs semantic similarity matching on the three types of terms and the corresponding node elements. Matching methods may include word vector similarity calculation, synonym expansion comparison, and similar semantic attribution analysis. For example, the word "medicine" in the instruction can establish a high entity similarity with the "capsule" and "prescription medicine" in the node; "kitchen" can form a spatial relationship link with "cooking area" and "stove".
[0119] After calculating the three types of similarity, the system weights and combines the similarity results of each type of match according to a preset ratio. This weighting can be set based on different task scenarios, such as giving a higher weight to entity words in an item search task, while prioritizing spatial words in a path planning task. Ultimately, by combining the three sub-matching scores, the system determines the overall match score for the semantic node to the current instruction, providing a basis for subsequent node screening and memory weight updates.
[0120] This matching process can greatly improve the ability to align the semantics of instructions with historical context. It is particularly suitable for dealing with ambiguous semantic problems caused by omitted subjects, skipped instructions or multi-round conversations, and effectively improves the accuracy of the intelligent agent's understanding and response to natural language instructions in complex scenarios.
[0121] The user instructions are divided into terms and semantically classified, and the time expression words, spatial location words and target entity words are extracted to generate a term set, which specifically includes:
[0122] During the interaction between an agent and a user, the user expresses operational instructions in natural language. These instructions may be verbal or textual, and their structure may exhibit characteristics such as semantic condensation, subject omission, and incomplete instruction patterns. Therefore, to achieve accurate semantic parsing, instruction sentences require lexical item segmentation and semantic classification, breaking them down into semantically functional keyword phrases.
[0123] First, the system uses a word segmentation tool to segment the instruction text into terms. This tool uses a bidirectional LSTM+CRF-based sequence tagging network, combined with a part-of-speech tagging model, to extract all structured language units, such as nouns, verbs, and prepositional phrases. For example, when faced with the instruction "Bring the medicine I just put on the table," the system recognizes the four phrases "just now," "on the table," "medicine," and "bring it over."
[0124] Next, the system performs semantic classification on the divided terms, which mainly include the following three categories:
[0125] Time expressions: used to describe the time when an event occurs or is referenced in an instruction, such as "just now", "yesterday", "five minutes later", "at that time", etc.
[0126] Spatial location words: used to represent the space where the target object or operation is located, such as "on the table", "in the kitchen", "in the cabinet", "at the door", etc.
[0127] Target entity words: These are the key operation objects mentioned by the user, usually specific objects, people, or concepts, such as "medicine," "remote control," "key," "cat," etc.
[0128] Classification is typically achieved through a combination of keyword matching and context analysis. For example, after "table" is identified as a noun, if it appears in combination with positional prepositions such as "up" or "down," it is classified as a spatial positional word. Similarly, if "medicine" is preceded by a verb such as "take" or "get," it is labeled as a target entity word.
[0129] The system organizes the above classification results into a structured data structure to form a set of terms, such as:
[0130] {
[0131] "time_terms":["just now"],
[0132] "space_terms":["on the table"],
[0133] "entity_terms":["medicine"]
[0134] }
[0135] This set will serve as the input data source for subsequent semantic matching calculations, ensuring that the semantic alignment operation has a clear three-dimensional vector basis.
[0136] Among them, according to the word set, the three categories of extracted words are semantically matched with the elements of the corresponding category in each semantic node in the context semantic network, and the similarity of the three categories of word matches is calculated respectively to obtain a similarity set, which specifically includes:
[0137] The purpose of semantic matching is to determine whether the time, space and entity information expressed by the current user instruction exists in the context semantic network, so as to achieve effective alignment of context instructions.
[0138] After generating a set of terms, the system matches the three categorical terms in that set with the corresponding semantic information recorded in each semantic node in the contextual semantic network. Each semantic node contains a time tag, a spatial description field, and a target entity identifier. This information is automatically extracted and recorded by the system during the fragment generation and graph construction phases.
[0139] The specific matching method is as follows:
[0140] Time word matching: Compare the time expression with the time tag in the semantic node. If the time expression is relative (such as "just now"), the system converts it to a relative offset in the system time and determines whether the node's time tag falls within this time window. For example, if "just now" refers to 2 to 5 minutes before the current time, the system determines whether the node time falls within this range.
[0141] Spatial word matching: Compares spatial location words with the location expressions recorded in semantic nodes. The system calculates the similarity between spatial descriptions based on geographic dictionaries, indoor object topology, and language embedding models. For example, if "on the dining table" and "next to the table" have high semantic overlap, the system marks them as a high match.
[0142] Target entity word matching: The semantic proximity between the entity noun and the node target entity is calculated using the entity noun's language embedding vector. For example, after using the Word2Vec or BERT model to generate the vector, the cosine similarity between the two is calculated. If the semantic distance is less than a preset threshold, the two are considered a match.
[0143] The system records the similarity results of the three types of matching for each semantic node, forming a similarity set. For example, for a node, its temporal matching is 0.8, spatial matching is 0.6, and target entity matching is 0.9; the system stores this set as:
[0144] {
[0145] "node_id":"N15",
[0146] "time_score":0.8,
[0147] "space_score":0.6,
[0148] "entity_score":0.9
[0149] }
[0150] This collection will be used in subsequent steps to calculate the comprehensive matching score, perform candidate node screening, and generate context paths.
[0151] Through the above-mentioned semantic matching mechanism, the system can ensure that the correspondence between user instructions and context has a semantic basis and structural support, solve the problem of misunderstanding caused by "term drift" or "phrase ambiguity" in traditional instruction parsing, and make semantic reasoning more robust and controllable.
[0152] In a preferred embodiment of the present invention, based on the candidate semantic node set, the current memory weight of the semantic node is dynamically updated by the comprehensive matching score of the semantic node to form an updated contextual semantic network, including:
[0153] According to the candidate semantic node set, obtain the current memory weight value and comprehensive matching score value of each candidate semantic node;
[0154] Multiply the current memory weight value by the first update coefficient, multiply the matching score value by the second update coefficient, and then sum the two to obtain the new memory weight of the candidate semantic node;
[0155] According to all non-candidate semantic nodes, the time difference between their time labels and the current time is calculated, and the time difference is input into the preset time decrement function to generate the attenuation coefficient;
[0156] Multiply the current memory weight of the non-candidate node by the decay coefficient to obtain the new memory weight of the non-candidate node;
[0157] The new memory weights of all semantic nodes are written into the corresponding semantic nodes to form an updated contextual semantic network. The updated memory weights are used for subsequent contextual semantic path extraction.
[0158] In this embodiment of the present invention, after the system matches the current instruction with the semantic nodes of the contextual semantic network, it dynamically updates the memory weights of the resulting set of candidate semantic nodes. This mechanism strengthens the current semantic focus area and suppresses distracting information, allowing the agent to more accurately focus on relevant content during subsequent behavior generation.
[0159] First, for each candidate semantic node, the system obtains its current memory weight and the matching score for the current round of commands. It then multiplies the current memory weight by the first update coefficient and the matching score by the second update coefficient. The sum of the two is then used as the node's new memory weight. The first coefficient primarily controls memory retention, while the second coefficient reflects the degree of matching enhancement. By linearly superimposing the two, we can both preserve historical semantic traces and enhance the current interaction's main semantic thread.
[0160] For semantic nodes that are not selected as candidates, the system applies a time-decreasing function to generate a decay coefficient based on the time difference between the time tag and the current time. The decay function can be set as a linear or exponential function, and its function is to control the natural decay of node memory over time. The longer the time, the smaller the generated decay coefficient. The system then multiplies this decay coefficient by the corresponding node's current memory weight to obtain its new memory weight value after decay.
[0161] After processing all nodes, the system writes the new memory weights back to the corresponding nodes in the semantic network, forming an updated contextual semantic network. When the next round of instructions arrives, this network will use the updated memory state as a basis for selecting semantic paths. Through this update mechanism, the system achieves bidirectional regulation of memory enhancement and interference suppression, giving the intelligent agent adaptive capabilities closer to human contextual memory. This is particularly suitable for multi-task scenarios, such as intelligent service robots executing user instructions and dialogue systems maintaining semantic consistency.
[0162] The current memory weight value is multiplied by the first update coefficient, the matching score value is multiplied by the second update coefficient, and the two are summed to obtain the new memory weight of the candidate semantic node, specifically including:
[0163] When updating the memory state of candidate semantic nodes, this paper introduces a weighted overlay mechanism, allowing the historical semantic state and current interaction results to jointly influence the node's importance. This mechanism essentially fuses these two information sources, "historical memory" and "current match," through coefficient adjustment to form a new semantic weight for subsequent context path screening.
[0164] First, for a candidate semantic node, the system extracts its memory weight value in the current context semantic network and records it as the basic memory value, which reflects the node's activity level in past semantics. At the same time, after matching the node with the current instruction, the system has calculated a comprehensive matching score for it, indicating its response level to the current semantic intention.
[0165] Next, the system multiplies the basic memory value by the first update coefficient and the matching score value by the second update coefficient. The product of the two is summed to obtain the new memory weight of the node. The first update coefficient is used to control the degree of retention of historical information, and the second update coefficient represents the strength of the current interaction semantics to modify the node state. Usually, the sum of these two coefficients does not exceed 1 to keep the total memory weight stable. For example: if the original memory weight is 0.6 and the matching score is 0.9, the first update coefficient is set to 0.7 and the second update coefficient is set to 0.3, then the new memory weight is (0.6×0.7)+(0.9×0.3)=0.69.
[0166] Through this dynamic fusion mechanism, the system can quickly focus on the semantic nodes associated with the current interaction while maintaining the continuity of historical context, strengthen its influence in the network, and make the semantic basis based on which the intelligent agent responds to instructions closer to the current user intention.
[0167] Among them, according to all non-candidate semantic nodes, the time difference between their time labels and the current time is calculated, and the time difference is input into the preset time decrement function to generate the attenuation coefficient, which specifically includes:
[0168] To prevent irrelevant nodes in the contextual semantic network from retaining high memory values for extended periods, leading to a biased memory of irrelevant content, this paper proposes a natural decay mechanism for memory weights. This mechanism dynamically adjusts the weights of semantic nodes based on their temporal properties, thereby maintaining the network's semantic sensitivity to the current interactive task.
[0169] First, the system extracts the timestamps of all non-candidate nodes (i.e., nodes not identified as associated by the current instruction) and compares them with the current system time. The difference between the two is calculated in seconds, minutes, or timestamp intervals. For example, if the node time is 14:00 and the current time is 14:30, the time difference is 30 minutes.
[0170] The system then inputs the time difference into a preset time-decreasing function to generate the memory decay coefficient of the node. The time-decreasing function can be designed as:
[0171] Linear function: Memory decreases linearly over time, which is suitable for interactive scenarios with sparse information and frequent updates; for example:
[0172] ;
[0173] : Indicates the current linearly decreasing weight of the node;
[0174] : Linear deceleration rate coefficient, which controls how much it decreases per unit time;
[0175] : The difference between the current time and the node time label;
[0176] : Ensure that the decreasing weight is not negative, that is, the lowest is 0.
[0177] Example usage scenarios:
[0178] Suitable for modeling tasks on medium time scales, such as object tracking or semantic responses within a few minutes.
[0179] Exponential function: Memory decays exponentially over time. It is suitable for high-density semantic networks, causing nodes far away from the interaction focus to be quickly "forgotten". For example:
[0180] ;
[0181] : represents the decay factor of the semantic node at the current time point;
[0182] : Initial attenuation amplitude constant (usually set to 1, i.e. no attenuation at the beginning);
[0183] : Time decay rate coefficient, which determines the speed of decay. The larger the value, the faster the weight decays.
[0184] : The time difference between the current system time and the semantic node time tag, the unit can be seconds, minutes, etc.;
[0185] e: A natural constant, approximately equal to 2.718.
[0186] Example usage scenarios:
[0187] It is suitable for time-sensitive short-term interactive tasks, such as "the action just now", and can quickly reduce the influence of old nodes.
[0188] Piecewise function: retains the original memory value for nodes within a specific time window, and decays sharply after exceeding the threshold; for example:
[0189] ;
[0190] f(t): represents the current decay factor of the semantic node;
[0191] : First time threshold (e.g. 30 seconds);
[0192] : The second time threshold (such as 90 seconds), after which the message is considered expired;
[0193] : Fixed weight value in the retention phase, usually 0.5 or 0.3;
[0194] : The interval between the current time and the node timestamp.
[0195] Example usage scenarios:
[0196] It is suitable for task switching scenarios in multi-round conversations to prevent sudden information drops, retain short-term memory but completely remove it outside the time limit.
[0197] The specific value is achieved by configuring the system's adjustable parameters. This process does not require user intervention and is automatically calculated and generated by the system after each round of interaction.
[0198] Through this time-driven decay mechanism, the system can ensure that the semantic network maintains a dynamic contraction state, reduce the computing resources occupied by useless nodes, and at the same time encourage the intelligent agent to always reason and judge around recent semantics, thereby improving the system's context processing efficiency and semantic accuracy.
[0199] The current memory weight of the non-candidate node is multiplied by the decay coefficient to obtain the new memory weight of the non-candidate node, which specifically includes:
[0200] After obtaining the time decay coefficient for each non-candidate node, the system multiplies the node's current memory weight by the corresponding decay coefficient to calculate the node's new memory state. This process is equivalent to performing a "soft deletion" or "degradation" on the time dimension, preserving some historical information while suppressing its influence on subsequent semantic reasoning.
[0201] For example, if a node's current memory weight is 0.5 and its time decay coefficient is 0.4, its updated memory weight will be 0.2. If a node has been inactive for an extended period of time, its decay coefficient may drop to 0.1 or even lower, and its memory weight will approach 0, causing it to be automatically excluded from semantic path extraction. The system performs this multiplication operation during each update cycle to ensure that the semantic network maintains a constant focus.
[0202] Furthermore, to prevent important semantic nodes (such as key entities referenced multiple times) from being decayed prematurely, the system can set a "minimum retention value" or a "memory lock flag." If a node has appeared as a member of a context path in several past rounds of interaction, its memory value must not fall below a set lower limit, or decay operations must be suspended, thereby improving the "historical memory" of the semantic network.
[0203] Through the above steps, the system realizes the dynamic management of semantic nodes, so that the entire semantic network can always revolve around the core of the current task, while naturally eliminating invalid information, providing the intelligent agent with a semantic graph structure foundation with time perception capabilities.
[0204] In a preferred embodiment of the present invention, the reference confidence of each candidate referent entity is calculated based on temporal proximity, semantic relevance, and modal confidence score, including:
[0205] Calculate the distance between each candidate referent entity and the referent word on the time axis according to their appearance positions in the semantic segment sequence, and convert it into a temporal neighbor score, where the temporal neighbor score decreases as the distance increases;
[0206] Compare the semantic content between the candidate referent entity and the referent word, extract their semantic word embedding information or context word vector, calculate their matching degree in the semantic dimension, and generate a semantic relevance score;
[0207] Identify the modal type from which the candidate referent entity originates, match it based on the preset modal priority table, and generate a modal confidence score;
[0208] The temporal neighbor score, semantic relevance score and modal confidence score are weighted respectively, and weighted superposition is performed to obtain the reference confidence of the candidate referent entity;
[0209] The candidate entity with the highest reference confidence is used as the reference target of the pronoun.
[0210] In this embodiment of the present invention, to make the coreference resolution process more accurate, especially when user expressions contain ambiguity, personal pronouns, omitted references, and other contexts, the system introduces a comprehensive scoring mechanism composed of temporal proximity, semantic relevance, and modal confidence scores to calculate the coreference confidence of candidate referent entities. This mechanism efficiently locates the semantic object most likely to be referred to by the user within a multimodal semantic segment and uses it to construct semantic nodes in the contextual semantic network.
[0211] During processing, the system first records the occurrence position of the pronoun in the semantic fragment sequence and extracts the occurrence position of all candidate pronoun entities according to the time sorting structure. Then, the system measures the distance between each candidate pronoun entity and the target pronoun relative position on the time axis. The smaller the distance, the greater the temporal proximity between the two in the dialogue context. The system reflects this distance as a temporal neighbor score, which is usually set as the inverse of the distance or a logarithmic function value to reflect its attenuation characteristics, that is, the smaller the time interval, the higher the score.
[0212] Next, the system extracts the semantic content features of each candidate entity, constructs a corresponding semantic representation vector, and calculates similarity with the vector of the target pronoun. Semantic embedding models (such as contextual language models) can be used to obtain semantic vectors. These vectors are then compared using cosine similarity or vector projections to generate a semantic relevance score. This score reflects the degree of semantic consistency between the candidate entity and the pronoun in terms of semantic concepts, categories, and attributes.
[0213] In addition, the system also considers the modal attributes of the candidate entity's source, that is, whether the entity originates from visual images, speech recognition, or plain text input. Due to differences in data quality and confidence across modalities, for example, textual content that is directly input typically has higher semantic clarity than ambiguous object labels extracted from image recognition. Therefore, based on a preset modality priority table, the system assigns corresponding confidence factors to different modal sources to calculate a modality confidence score.
[0214] The system ultimately assigns different weights to the temporal proximity score, semantic relevance score, and modal confidence score, and calculates the reference confidence value for each candidate entity using a weighted superposition method. Among all candidate entities, the one with the highest confidence score is selected as the target entity pointed to by the current pronoun. This processing method avoids misreferences caused by modal bias or syntactic ambiguity and is particularly suitable for linguistic features such as omissions, reference jumps, and context compression that often appear in users' natural expressions. It helps to enhance the semantic accuracy and context consistency understanding capabilities of the intelligent agent.
[0215] To facilitate system processing and scoring, the system sets the following modal priority table:
[0216]
[0217] This modality confidence score will be used as a sub-score item of the candidate entity and participate in the comprehensive calculation of the reference confidence. The higher the score, the more credible the source of the entity is in the current scenario. For example, when the user's instruction "take it away" appears, if the system detects that the text clearly mentions "medicine" and the word originates from the text modality, the corresponding modality confidence score will be 1.0; while the confidence score of "water bottle" appearing in the image is only 0.7, and its priority in the overall reference judgment will be automatically lowered.
[0218] Furthermore, to adapt to different application scenarios, the system allows developers to customize priority strategies. For example, in scenarios without text input, such as a purely image-based visual agent system, the confidence weight of the image modality can be increased to 0.9 to reflect its primary information source.
[0219] Through the design of this modal priority table, the system can model the credibility differences of candidate entities from different sources, achieve systematic control of modal uncertainty, improve the stability and rationality of reference resolution, and solve the problem of "reference errors caused by unclear sources" in existing technologies.
[0220] In a preferred embodiment of the present invention, based on a time-ordered node sequence, for any two semantic nodes, the degree of semantic association is calculated by comparing the directional consistency and content similarity between their semantic features, including:
[0221] According to the time-ordered node sequence, for any two semantic nodes, the degree of semantic association is calculated by comparing the directional consistency and content similarity between their semantic features, including:
[0222] According to the time-ordered node sequence, the semantic direction features and semantic content features are extracted from the semantic nodes; the following processing is performed on any two semantic nodes:
[0223] Compare the semantic directional features, map them into directional vectors, and then calculate the angle or cosine similarity between the two directional vectors as the directional consistency score;
[0224] Compare semantic content features and calculate the semantic co-occurrence degree between the referent entities contained in the two semantic nodes, the degree of matching of contextual keywords, and the consistency of the contextual background to form a semantic content similarity score;
[0225] Multiply the direction consistency score and the semantic content similarity score by their respective contribution coefficients, and then sum them up to obtain the semantic association degree between the two semantic nodes;
[0226] When the degree of semantic association is higher than the set association threshold and the node time sequence satisfies the order from early to late, a unidirectional connection relationship is established between the two nodes, from the earlier node to the later node, as part of the contextual semantic network structure.
[0227] In this embodiment of the present invention, to improve the accuracy of node connections and the coherence of contextual expressions in semantic networks, the system analyzes the directional consistency and content similarity between any two semantic nodes in a time-ordered node sequence during the construction of the contextual semantic network. This calculation determines the degree of semantic association between the two nodes and determines whether a directed connection should be established between them. This process not only considers the node content itself but also the evolution of the nodes' semantic behavior trends. This ensures that the network structure is more consistent with the actual logic of semantic change, improving the accuracy of subsequent path extraction and semantic reasoning.
[0228] First, when constructing a time-ordered sequence of nodes, the system extracts the semantic directional features and semantic content features of each semantic node. The semantic directional feature refers to the dynamic directional information of the event or action expressed in the semantic node, which is commonly seen in the directionality of verbs or verb phrases, such as "put down," "pick up," "move to," and "hand over." The system can use the language parsing module to analyze the subject-predicate structure of the sentence in the node, identify the semantic backbone, and use semantic embedding technology to map the verb and its object combination into a directional vector to represent the "directionality" of the action in the semantic space. For example, "put the medicine on the table" indicates the direction of action "changing position to a stationary state," while "take away" indicates the opposite direction of "moving away from the current position." The directional vector can be represented as a continuous value in a high-dimensional space, supporting subsequent vector calculations.
[0229] When the system selects two semantic nodes for pairing in chronological order, it first compares their semantic directional features. The cosine similarity or angle between the two directional vectors is calculated to determine the degree of consistency between the two in terms of action trends. If the two represent semantically converging behaviors, such as "put down" and "place", their directional vectors are close and the similarity is close to 1; if they represent behaviors in opposite directions, such as "put down" and "take away", the similarity is close to 0 or a negative value. The system uses this calculation result as a directional consistency score to measure the degree of similarity between the two nodes in terms of behavioral evolution trends.
[0230] Next, the system compares the semantic content features between the two semantic nodes. This comparison process includes three sub-dimensions: First, the system analyzes whether the two nodes involve the same referent entity or its equivalent entity. If the two nodes share the same target entity (such as "medicine" and "that bottle"), the co-occurrence degree is high. Second, the system matches the keywords appearing in the sentence and calculates the semantic proximity between the contextual keywords. For example, "kitchen" and "on the table" may have overlapping backgrounds in the spatial semantic scene. Third, the system determines whether the two nodes appear in the same context, such as whether they belong to the same larger semantic scene of "user picking up an item." The scores of the three dimensions are weighted and fused to form a semantic content similarity score, which is used to measure the degree of semantic fit between nodes at the information level.
[0231] The system then performs a weighted combination of the directional consistency score and the semantic content similarity score by multiplying each by its contribution coefficient to the overall relevance. The two products are then summed to form a comprehensive semantic relevance score. For example, if the directional consistency score is 0.9 and the content similarity score is 0.7, and the system presets a 1:1 contribution ratio between the two, the final semantic relevance score is 0.8. The contribution coefficient can be dynamically adjusted based on task requirements, such as prioritizing content consistency in object interaction tasks and directional coherence in process monitoring tasks.
[0232] The system compares the degree of semantic association with a set threshold. If the score is higher than the threshold, and the two nodes meet the condition that the former's time label is earlier than the latter in terms of time order, the system will establish a one-way connection between the two nodes. The direction of the connection is from the earlier semantic node to the later node, which conforms to the temporal advancement law and semantic development logic of natural language expression. This connection structure is incorporated into the semantic network structure and stored as part of the network graph to support subsequent processing processes such as path extraction, intent understanding, and entity backtracking.
[0233] Through the above method, the present invention not only realizes semantic organization based on temporal sequence, but also introduces directional trend modeling and content matching control, so that the structure of the contextual semantic network has stronger semantic logical coherence and behavioral evolution interpretability, which is particularly suitable for multi-round dialogue understanding and semantic anaphora analysis in complex task scenarios, greatly improving the system's semantic composition ability and reasoning accuracy in continuous contexts.
[0234] Among them, the semantic content features are compared, and the semantic content similarity score is formed by measuring the degree of semantic co-occurrence between the referent entities contained in the two semantic nodes, the degree of matching of contextual related keywords, and whether the context background is consistent. Specifically, it includes:
[0235] In order to judge the content similarity of two semantic nodes in semantic information, the present invention introduces three content-level comparison indicators to compare the semantic proximity between the two nodes from the three dimensions of entity co-occurrence, keyword overlap and contextual background consistency, and integrate them into a content similarity score as one of the semantic bases for establishing connections between semantic nodes.
[0236] First, the system extracts and compares the referent entities contained in the two semantic nodes. The system uses a referent chain tracking mechanism to uniformly map the entities mentioned in the semantic nodes to their real representative objects. For example, the core entity pointed to by "medicine", "capsule", "it", etc. may be "cold medicine". After determining the main entity in the semantic node through the referent resolution module, the system performs a semantic co-occurrence analysis on the main entities of the two nodes. If the same or semantically equivalent entities appear in the two nodes (such as "water cup" and "glass"), the co-occurrence degree is considered to be high. The system can use the word embedding model to vectorize the entity words and measure their semantic distance, such as calculating their co-occurrence strength through the similarity of word vectors. If the two entities are close in the semantic space, the system will set the co-occurrence degree to a higher score; if the semantic distance is far, the score will be lower.
[0237] Secondly, the system extracts the set of contextual keywords that appear in the two semantic nodes. Keywords refer to supplementary semantic information other than the main entity, such as "place", "pick up", "kitchen", "desktop", etc. The system uses TF-IDF, TextRank or named entity recognition algorithms to extract keywords and construct a keyword set. The overlap rate between the two sets is then calculated, for example, by measuring the ratio of their intersection to their union through Jaccard similarity. If two nodes share multiple semantic keywords, it means that their semantic scenes have a strong overlap and a high degree of matching.
[0238] Finally, the system determines whether the two semantic nodes belong to the same context. The context can be identified through the context window or classified through semantic labels. For example, "behavior in the kitchen" and "operation on the stove" can be divided into the same "kitchen work" context, while "writing in the study" and "picking up food in the kitchen" belong to different contexts. The system can add background labels to semantic nodes. The label source can be manual annotation, a topic model obtained by corpus training, or automatic identification based on a scene classification network. If two nodes have the same context label or are in the same context window (such as a short time interval and continuous before and after), the context consistency score is higher.
[0239] Ultimately, the system performs a weighted fusion of the three sub-scores (co-occurrence, keyword matching, and contextual consistency) to form a content similarity score for the semantic node pair. This score serves as a key indicator in determining semantic connectivity, assessing information overlap and behavioral continuity between nodes.
[0240] The direction consistency score and the semantic content similarity score are multiplied by their respective contribution coefficients and then summed to obtain the semantic association degree between the two semantic nodes, including:
[0241] To achieve a comprehensive assessment of both directionality and content, this paper introduces the concept of a "contribution coefficient" to regulate the relative weights of the "directional consistency score" and "semantic content similarity score" in the overall relevance score. This coefficient is a system design parameter, set according to different task requirements or operating environments, reflecting the importance of each factor to the semantic relationship.
[0242] The system sets the contribution coefficients to two normalized values, one for directional consistency and the other for content similarity. The sum of these two coefficients is generally 1 (or some other normalized total) to prevent scale drift in the overall score. The default setting is an equal weighting configuration, with a directional consistency contribution coefficient of 0.5 and a content similarity contribution coefficient of 0.5, indicating that the system pays equal attention to both semantic relationships.
[0243] In applications, this coefficient can also be adjusted based on the usage scenario. For example, in tasks with strong descriptive language, diverse verbs, and drastic scene changes (such as service robot instruction parsing), the system can increase the contribution of direction consistency to 0.7. In scenarios such as object search and object tracking, where the semantic content of the target entity is more important, the contribution of content similarity can be increased to 0.8.
[0244] For example, if a node pair's directional consistency score is 0.9 and its content similarity score is 0.6, and the system sets the directional contribution coefficient to 0.6 and the content contribution coefficient to 0.4, the overall semantic relevance is: (0.9 × 0.6) + (0.6 × 0.4) = 0.78. This overall score will be used as one of the criteria for determining whether to establish a connection.
[0245] By introducing the contribution coefficient, the system not only has the ability to flexibly weight different semantic factors, but also makes the semantic network construction process task-adaptive. It can dynamically adjust the structure generation strategy according to the semantic evolution characteristics, thereby enhancing the system's robustness and generalization capabilities in a variety of semantic environments.
[0246] An embodiment of the present invention further provides a multimodal semantic network driven intelligent agent context understanding system, the system comprising:
[0247] The data acquisition and processing module is used to receive raw interactive data including images, voice, and text, perform modal encoding on each, extract features of each modality, and then synchronize and align the temporal information to generate a sequence of semantic segments;
[0248] The semantic annotation module is used to perform semantic label expansion and coreference resolution on the semantic segment sequence, annotating the time anchor point and coreference target of each semantic segment to obtain a contextual semantic segment set;
[0249] The node construction module is used to convert each contextual semantic fragment into a semantic node based on the contextual semantic fragment set, set the initial memory weight, time label and source modality, generate directed connections based on the temporal relationship and semantic association between semantic nodes, and construct a contextual semantic network;
[0250] The instruction parsing module is used to extract the time words, space words and target entity words based on the user instruction, match the relevant semantic nodes in the context semantic network, and calculate the comprehensive matching score;
[0251] A candidate node screening module is used to compare the comprehensive matching score with a preset comprehensive matching score threshold, screen candidate semantic nodes, and generate a candidate semantic node set;
[0252] The weight update module is used to dynamically update the current memory weight of the candidate semantic node set based on the comprehensive matching score of the semantic node to form an updated contextual semantic network, in which the semantic nodes with high comprehensive matching scores are given higher weights, and the semantic nodes further away from the current moment are reduced in weight according to the attenuation coefficient;
[0253] The structured output module is used to extract the semantic node sequence with the highest current memory weight as the context semantic path based on the updated context semantic network, and structure the output of the semantic entities in the path for intelligent agent task understanding and response generation.
[0254] It should be noted that this system is a system corresponding to the above method, and all implementation methods in the above method embodiment are applicable to this embodiment and can achieve the same technical effects.
[0255] An embodiment of the present invention further provides a computing device comprising: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the above-described method. All implementations in the above-described method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0256] The embodiment of the present invention further provides a computer-readable storage medium storing instructions, which, when executed on a computer, causes the computer to execute the above-described method. All implementations in the above-described method embodiment are applicable to this embodiment and can achieve the same technical effects.
[0257] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A multimodal semantic network driven agent context understanding method, characterized by: The method comprises: Receive raw interaction data containing images, voice, and text, perform modal encoding on each, extract features of each modality, and then synchronize and align the temporal information to generate a sequence of semantic segments; According to the sequence of semantic segments, the temporal expression in each semantic segment is identified, a mapping relationship between the segment and the standard time axis is established, and a temporal anchor point label is generated. According to the sequence of semantic segments, the pronouns in each semantic segment are identified to determine the candidate referent entities; Based on temporal proximity, semantic relevance, and modal confidence scores, the reference confidence of each candidate referent entity is calculated, and the candidate referent entity with the highest reference confidence is determined as the referent target of the referent; The referent and temporal anchor annotations are appended to the corresponding semantic segments to form contextual semantic segments with temporal anchor labels and referent labels, thus obtaining a contextual semantic segment set. Based on the set of contextual semantic fragments, each contextual semantic fragment is converted into a semantic node, and the initial memory weight, time label and source modality are set. Based on the temporal relationship and semantic association between semantic nodes, directed connections are generated to construct a contextual semantic network. Based on user instructions, extract the time words, space words and target entity words, match them with relevant semantic nodes in the context semantic network, and calculate the comprehensive matching score; Compare the comprehensive matching score with the preset comprehensive matching score threshold, screen the candidate semantic nodes, and generate a candidate semantic node set; Based on the candidate semantic node set, the current memory weight is dynamically updated through the comprehensive matching score of the semantic node to form an updated contextual semantic network; According to the updated contextual semantic network, the semantic node sequence with the highest current memory weight is extracted as the contextual semantic path, and the semantic entities in the path are structured and output.
2. The multimodal semantic network driven agent context understanding method according to claim 1, characterized in that: Generate directed connections based on the temporal relationship and semantic association between semantic nodes to build a contextual semantic network, including: According to the time label of each semantic node, it will be sorted in chronological order to form a time-ordered node sequence; According to the time-ordered node sequence, for any two semantic nodes, the degree of semantic association is calculated by comparing the directional consistency and content similarity between their semantic features; When the semantic correlation between two semantic nodes is greater than the preset correlation threshold and the former semantic node is earlier than the latter semantic node in time, a unidirectional connection is established between them, with the direction pointing to the later semantic node in time, forming a contextual semantic network.
3. The multimodal semantic network driven agent context understanding method according to claim 1, characterized in that: Based on user instructions, extract the time words, space words, and target entity words, match them with relevant semantic nodes in the contextual semantic network, and calculate the comprehensive matching score, including: Perform lexical item division and semantic classification on user instructions, extract time expression words, spatial location words and target entity words, and generate a lexical item set; According to the word set, the three categories of words extracted are semantically matched with the elements of the corresponding categories in each semantic node in the context semantic network, and the similarity of the three categories of word matches is calculated to obtain a similarity set; Based on the similarity set, the three types of similarity results are weighted according to the preset proportional coefficient to form a comprehensive matching score.
4. The multimodal semantic network driven agent context understanding method according to claim 1, characterized in that: Based on the candidate semantic node set, the current memory weight is dynamically updated through the comprehensive matching score of the semantic node to form an updated contextual semantic network, including: According to the candidate semantic node set, obtain the current memory weight value and comprehensive matching score value of each candidate semantic node; Multiply the current memory weight value by the first update coefficient, multiply the matching score value by the second update coefficient, and then sum the two to obtain the new memory weight of the candidate semantic node; According to all non-candidate semantic nodes, the time difference between their time labels and the current time is calculated, and the time difference is input into the preset time decrement function to generate the attenuation coefficient; Multiply the current memory weight of the non-candidate node by the decay coefficient to obtain the new memory weight of the non-candidate node; The new memory weights of all semantic nodes are written into the corresponding semantic nodes to form an updated contextual semantic network. The updated memory weights are used for subsequent contextual semantic path extraction.
5. The multimodal semantic network driven agent context understanding method according to claim 1, characterized in that: Based on temporal proximity, semantic relevance, and modal confidence scores, the reference confidence of each candidate referent entity is calculated, including: Calculate the distance between each candidate referent entity and the referent word on the time axis according to their appearance positions in the semantic segment sequence, and convert it into a temporal neighbor score, where the temporal neighbor score decreases as the distance increases; Compare the semantic content between the candidate referent entity and the referent word, extract their semantic word embedding information or context word vector, calculate their matching degree in the semantic dimension, and generate a semantic relevance score; Identify the modal type from which the candidate referent entity originates, match it based on the preset modal priority table, and generate a modal confidence score; The temporal neighbor score, semantic relevance score and modal confidence score are weighted respectively, and weighted superposition is performed to obtain the reference confidence of the candidate referent entity; The candidate entity with the highest reference confidence is used as the reference target of the pronoun.
6. The multimodal semantic network driven agent context understanding method according to claim 2, characterized in that: According to the time-ordered node sequence, for any two semantic nodes, the degree of semantic association is calculated by comparing the directional consistency and content similarity between their semantic features, including: According to the time-ordered node sequence, for any two semantic nodes, the degree of semantic association is calculated by comparing the directional consistency and content similarity between their semantic features, including: According to the time-ordered node sequence, the semantic direction features and semantic content features are extracted from the semantic nodes; the following processing is performed on any two semantic nodes: Compare the semantic directional features, map them into directional vectors, and then calculate the angle or cosine similarity between the two directional vectors as the directional consistency score, with a value range of 0 to 1; Compare semantic content features and calculate the semantic co-occurrence degree between the referent entities contained in the two semantic nodes, the matching degree of contextual keywords, and the consistency of the contextual background to form a semantic content similarity score; Multiply the direction consistency score and the semantic content similarity score by their respective contribution coefficients, and then sum them up to obtain the semantic association degree between the two semantic nodes; When the degree of semantic association is higher than the set association threshold and the node time sequence satisfies the order from early to late, a unidirectional connection relationship is established between the two nodes from the earlier node to the later node as part of the contextual semantic network structure.
7. A multimodal semantic network driven agent context understanding system, characterized by: Applied to the method according to any one of claims 1 to 6, the system comprises: The data acquisition and processing module is used to receive raw interactive data including images, voice, and text, perform modal encoding on each, extract features of each modality, and then synchronize and align the temporal information to generate a sequence of semantic segments; The semantic annotation module is used to perform semantic label expansion and coreference resolution on the semantic segment sequence, annotating the time anchor point and coreference target of each semantic segment to obtain a contextual semantic segment set; The node construction module is used to convert each contextual semantic fragment into a semantic node based on the contextual semantic fragment set, set the initial memory weight, time label and source modality, generate directed connections based on the temporal relationship and semantic association between semantic nodes, and construct a contextual semantic network; The instruction parsing module is used to extract the time words, space words and target entity words based on the user instruction, match the relevant semantic nodes in the context semantic network, and calculate the comprehensive matching score; A candidate node screening module is used to compare the comprehensive matching score with a preset comprehensive matching score threshold, screen candidate semantic nodes, and generate a candidate semantic node set; The weight update module is used to dynamically update the current memory weight of the candidate semantic node set based on the comprehensive matching score of the semantic node to form an updated contextual semantic network, in which the semantic nodes with high comprehensive matching scores are given higher weights, and the semantic nodes further away from the current moment are reduced in weight according to the attenuation coefficient; The structured output module is used to extract the semantic node sequence with the highest current memory weight as the context semantic path based on the updated context semantic network, and structure the output of the semantic entities in the path for intelligent agent task understanding and response generation.
8. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Cooperative disambiguation method based on deep semantic neighbor and multivariate entity association
CN112883199A
Large model deployment method and system based on multi-level cache mechanism
CN119739809A