Memory tag-based memory management method and mobile terminal
Patent Information
- Application Number
- CN202610911652.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-06-24
AI Technical Summary
但该类方案通常没有将交互时间、交互地点、兴趣点属性、交互主题、用户意图、环境参数等信息统一形成与交互记忆绑定的记忆标签集,导致后续按照地点、主题、意图等条件检索历史记忆时,检索效率和准确性较低
[0015]如上所述,本公开实施例中提供基于记忆标签的记忆管理方法及移动终端,方法包括:获取智能体与用户交互过程中产生的当前的交互数据段;基于所述交互数据段生成对应的记忆单元;其中,所述记忆单元被添加至一记忆数据库;基于所述交互数据段进行语义识别,得到交互语义信息;获取所述用户在形成所述交互数据段时的定位信息,并得到所述定位信息对应场所作为兴趣点,并获取所述兴趣点对应的兴趣点属性信息,所述兴趣点属性信息包括兴趣点类型和兴趣点位置坐标;基于所述兴趣点类型同所述交互语义信息之间的匹配得到所述记忆单元的语义置信度,并根据所述兴趣点位置坐标同所述定位信息之间距离得到所述记忆单元的距离置信度;根据所述记忆单元的语义置信度和距离置信度确定所述记忆单元的语义-地理置信校验结果,并将所述语义-地理置信校验结果作为一记忆标签加入至所述记忆单元的记忆标签集;响应于用户的当前交互请求,在所述记忆数据库中查询具有与所述当前交互请求匹配的目标记忆标签的目标记忆单元,并据以形成对所述当前交互请求的响应信息。本公开实施例通过对记忆单元生成包含语义-地理置信校验结果的记忆标签集,使用户能够根据语义-地理置信校验结果判断记忆单元中的交互语义与对应的地点信息之间是否匹配,从而本公开提高记忆单元展示结果的可信度和用户交互体验。
Smart Images

Figure CN122470825B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of memory management technology, and in particular to a memory management method and mobile terminal based on memory tags. Background Technology
[0002] With the widespread adoption of intelligent agent applications such as smart assistants, smart speakers, and mobile voice assistants, a large amount of interactive data is continuously generated between users and these agents. This interactive data typically includes text, voice, images, user behavior, location information, and environmental awareness information. Structured management of this interactive data, and enabling the system to quickly retrieve historical interaction content based on criteria such as time, location, topic, and intent, is a crucial issue in intelligent agent memory management.
[0003] In existing technologies, one type of solution generates tags for multimedia files such as photos, videos, and audio recordings. For example, after a user uploads a multimedia file to the cloud, the cloud server generates tags such as time, location, and scene. This type of solution relies on cloud processing, which suffers from high latency, privacy risks due to uploading raw data, and its processing objects are usually not the interactive memories between the user and the intelligent agent. Another type of solution mainly generates simple tags based on terminal sensors. For example, smart bracelets or mobile terminals can generate motion tags or location tags based on accelerometers, positioning data, etc. This type of solution usually only focuses on motion state or geographical location, lacking structured processing of interactive content such as dialogue topics, user intentions, image semantics, and environmental parameters. A third type of solution is the dialogue log system. This type of system usually saves raw text or voice records in chronological order and supports searching by time or keywords. However, this type of solution usually does not unify information such as interaction time, interaction location, point of interest attributes, interaction topics, user intentions, and environmental parameters into a memory tag set bound to the interaction memory, resulting in low retrieval efficiency and accuracy when subsequently searching historical memories by location, topic, intention, etc.
[0004] Furthermore, in existing solutions, location information and interaction semantic information are usually independent of each other, lacking consistency verification. When there is drift in terminal positioning or inaccurate matching of points of interest, the system may directly write incorrect location tags into the memory data, leading to false recalls in subsequent location-based memory retrieval. Therefore, existing technologies suffer from problems such as a single dimension of memory tags, lack of verification between location information and interaction semantic information, separation of dialogue memory and structured tags, and insufficient retrieval accuracy. Summary of the Invention
[0005] In view of the shortcomings of the prior art described above, the purpose of this disclosure is to provide a memory management method and mobile terminal based on memory tags, and to solve the problems in the related technologies. The first aspect of this disclosure provides a memory management method based on memory tags, comprising: acquiring a current interaction data segment generated during interaction between an agent and a user; generating a corresponding memory unit based on the interaction data segment; wherein the memory unit is added to a memory database; performing semantic recognition based on the interaction data segment to obtain interaction semantic information; acquiring the user's location information when the interaction data segment is formed, and obtaining the location corresponding to the location information as a point of interest, and acquiring the point of interest attribute information corresponding to the point of interest, the point of interest attribute information including point of interest type and point of interest location coordinates; and acquiring the point of interest type and the corresponding memory unit based on the memory tag. The semantic confidence of the memory unit is obtained by matching the semantic information between the interactive information, and the distance confidence of the memory unit is obtained based on the distance between the location coordinates of the point of interest and the location information. The semantic-geographic confidence verification result of the memory unit is determined based on the semantic confidence and distance confidence of the memory unit, and the semantic-geographic confidence verification result is added as a memory tag to the memory tag set of the memory unit. In response to the user's current interaction request, the target memory unit with the target memory tag that matches the current interaction request is queried in the memory database, and a response information to the current interaction request is formed accordingly.
[0006] In the first aspect of the embodiment, based on the interactive semantic information corresponding to each memory unit, the semantic correlation degree between the multiple memory units is obtained, and based on the time information and the location information corresponding to each memory unit, the spatiotemporal correlation degree between the multiple memory units is obtained; based on the semantic correlation degree and the spatiotemporal correlation degree, the memory correlation degree between each memory unit is determined; the memory units whose memory correlation degree with the target memory unit satisfies the correlation degree condition are determined as the associated memory units corresponding to the target memory unit; the associated memory units are sorted along the temporal direction to form the temporal context sequence of the target memory unit, and the temporal context sequence is added as a memory tag to the memory tag set of the target memory unit.
[0007] In an embodiment of the first aspect, the memory association links are formed based on the temporal connections between memory units in the same temporal context sequence, and the set of memory association links forms a memory spatiotemporal network.
[0008] In an embodiment of the first aspect, feature extraction processing is performed on the interactive data segment according to the modality type of the interactive data segment to obtain modality feature information corresponding to the interactive data segment; the modality type is at least one of text, speech, and image; tag recognition processing is performed on the modality feature information to obtain the modality feature tag corresponding to the modality feature information, and the modality feature tag is added to the memory tag set corresponding to the memory unit; wherein, the modality feature tag includes a conversation tag based on text or speech and / or an image tag, and the conversation tag includes a topic tag and / or an intent tag.
[0009] In a first aspect embodiment, obtaining the semantic confidence of the memory unit based on the matching between the interest point type and the interaction semantic information includes: performing semantic recognition on modal feature tags extracted from one or more modal types within the same time period to obtain corresponding tag semantic information; in response to the existence of tag semantic information of multiple modal types, fusing the tag semantic information to obtain fused tag semantic information; performing semantic recognition on the interest point information corresponding to the interest point to obtain interest point semantic information; and obtaining the semantic confidence of the memory unit based on the semantic similarity between the fused tag semantic information and the interest point semantic information.
[0010] In an embodiment of the first aspect, the interactive data segment further includes an environmental perception data segment, and the method further includes: mapping the environmental perception data segment to obtain an environmental perception tag, and adding the environmental perception tag to the memory tag set corresponding to the memory unit; wherein, the environmental perception data segment includes at least one environmental parameter collected by an environmental sensor.
[0011] In a first aspect embodiment, key tags are determined based on user instructions; the memory tag set is stored in a tag index library; wherein the tag index library establishes at least one index based on the key tags.
[0012] In an embodiment of the first aspect, determining the semantic-geographic confidence verification result of the memory unit based on the semantic confidence and distance confidence of the memory unit includes: performing a weighted fusion of the distance confidence and the semantic confidence to determine the semantic-geographic confidence verification result.
[0013] In an embodiment of the first aspect, the step of querying the memory database for a target memory unit with a target memory tag matching the current interaction request in response to a user's current interaction request includes: in response to the search conditions corresponding to the current interaction request including location search conditions, filtering a set of candidate memory tags matching the location search conditions from the memory tag set corresponding to a plurality of memory units; determining the location search confidence level corresponding to the candidate memory tag set based on the semantic-geographic confidence verification result of the candidate memory tag set; determining the target memory tag set from the candidate memory tag set based on the location search confidence level; and determining the target memory unit based on the memory unit identity identifier corresponding to the target memory tag set.
[0014] A second aspect of this disclosure provides a mobile terminal, comprising: a processor and a memory; the memory storing a computer program or instructions; the processor being configured to run the computer program or instructions to perform the memory management method based on memory tags as described in the first aspect.
[0015] As described above, this disclosure provides a memory management method and mobile terminal based on memory tags. The method includes: acquiring a current interaction data segment generated during the interaction between an agent and a user; generating a corresponding memory unit based on the interaction data segment; wherein the memory unit is added to a memory database; performing semantic recognition based on the interaction data segment to obtain interaction semantic information; acquiring the user's location information when the interaction data segment is formed, and obtaining the location corresponding to the location information as a point of interest, and acquiring the point of interest attribute information corresponding to the point of interest, the point of interest attribute information including the point of interest type and the point of interest location coordinates; and based on the point of interest... The semantic confidence of the memory unit is obtained by matching the type with the interactive semantic information, and the distance confidence of the memory unit is obtained based on the distance between the location coordinates of the point of interest and the location information. The semantic-geographic confidence verification result of the memory unit is determined based on the semantic confidence and distance confidence, and the semantic-geographic confidence verification result is added as a memory tag to the memory tag set of the memory unit. In response to the user's current interaction request, a target memory unit with a target memory tag matching the current interaction request is queried in the memory database, and a response information to the current interaction request is formed accordingly. This embodiment of the present disclosure generates a memory tag set containing semantic-geographic confidence verification results for the memory unit, enabling the user to determine whether the interactive semantics in the memory unit match the corresponding location information based on the semantic-geographic confidence verification results. This improves the credibility of the memory unit display results and the user interaction experience. Attached Figure Description
[0016] Figure 1This diagram illustrates a scenario application of the memory management method based on memory tags according to one embodiment of the present disclosure.
[0017] Figure 2 A flowchart illustrating the memory management method based on memory tags in an embodiment of this disclosure is shown.
[0018] Figure 3 A flowchart illustrating a modal feature label generation method in yet another embodiment of this disclosure is shown.
[0019] Figure 4 A schematic diagram of a data processing structure based on memory tags is shown in one embodiment of this disclosure.
[0020] Figure 5 This illustration shows a flowchart of a method for generating temporal context sequences based on memory association in one embodiment of the present disclosure.
[0021] Figure 6(a) shows a schematic diagram of the structure of the target memory association link in the memory spatiotemporal network in one embodiment of the present disclosure.
[0022] Figure 6(b) shows a schematic diagram of the structure of the target memory association link in the memory spatiotemporal network in another embodiment of this disclosure.
[0023] Figure 7 A flowchart illustrating a method for adding new memory units in a memory-space-time network according to an embodiment of this disclosure is provided.
[0024] Figure 8 A schematic diagram of a memory management system based on memory tags is shown in one embodiment of this disclosure.
[0025] Figure 9 A schematic diagram of the structure of a mobile terminal is shown in one embodiment of this disclosure. Detailed Implementation
[0026] The following specific examples illustrate the implementation of this disclosure. Those skilled in the art can easily understand other advantages and effects of this disclosure from the information disclosed herein. This disclosure can also be implemented or applied through other different specific embodiments, and various details in this disclosure can be modified or changed according to different viewpoints and application modules without departing from the spirit of this disclosure. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this disclosure can be combined with each other.
[0027] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings, so that those skilled in the art to which this disclosure pertains can readily implement it. This disclosure may be embodied in many different forms and is not limited to the embodiments described herein.
[0028] In this disclosure, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic represented in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. Furthermore, the specific features, structures, materials, or characteristics represented may be combined in any suitable manner in any one or a group of embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples represented in this disclosure, as well as the features of those different embodiments or examples.
[0029] Furthermore, the terms "first" and "second" are used for illustrative purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the representation of this disclosure, "a set" means two or more, unless otherwise explicitly specified.
[0030] For the purpose of clarity, devices unrelated to the description are omitted, and the same or similar components throughout the specification are given the same reference numerals.
[0031] Throughout this specification, when it is said that a device is "connected" to another device, this includes not only "direct connection" but also "indirect connection" by placing other components in between. Furthermore, when it is said that a device "comprises" a certain constituent element, unless otherwise stated otherwise, this does not exclude other constituent elements, but rather implies that other constituent elements may be included.
[0032] While the terms first, second, etc., are used in some examples herein to refer to various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, first interface and second interface, etc., are used. Furthermore, as used herein, the singular forms “a,” “an,” and “the” are intended to also include the plural forms unless the context indicates otherwise. It should be further understood that the terms “comprising,” “including,” indicate the presence of the stated feature, step, operation, element, module, item, kind, and / or group, but do not exclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, modules, items, kinds, and / or groups. The terms “or” and “and / or” as used herein are interpreted as inclusive, or mean any one or any combination thereof. Thus, “A, B, or C” or “A, B, and / or C” means “any one of: A; B; C; A and B; A and C; B and C; A, B, and C.” Exceptions to this definition will only occur if the combination of elements, functions, steps, or operations is inherently mutually exclusive in some way.
[0033] The technical terms used herein are for reference only to specific embodiments and are not intended to limit the scope of this disclosure. The singular form used herein includes the plural form unless the statement explicitly indicates otherwise. The word "comprising" as used in this specification means to specify a particular characteristic, region, integer, step, operation, element, and / or component, and does not exclude the presence or addition of other characteristics, regions, integers, steps, operations, elements, and / or components.
[0034] Although not explicitly defined, all terms, including technical and scientific terms used herein, shall have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Terms defined in commonly used dictionaries shall be further interpreted as having a meaning consistent with the relevant technical literature and the message of the present disclosure, and shall not be over-interpreted as having an ideal or overly formulaic meaning unless otherwise defined.
[0035] Currently, with the widespread adoption of intelligent agent applications such as smart assistants, smart speakers, and mobile phone voice assistants, a large amount of interactive data is continuously generated between users and these intelligent agents. This type of interactive data typically includes text, voice, images, user behavior, location information, and environmental awareness information. Structuring and managing this interactive data, and enabling the system to quickly retrieve historical interaction content based on criteria such as time, location, topic, and intent, is a crucial issue in intelligent agent memory management.
[0036] In existing technologies, one type of solution generates tags for multimedia files such as photos, videos, and audio recordings. For example, after a user uploads a multimedia file to the cloud, the cloud server generates tags such as time, location, and scene. This type of solution relies on cloud processing, which suffers from high latency, privacy risks due to uploading raw data, and its processing objects are usually not the interactive memories between the user and the intelligent agent. Another type of solution mainly generates simple tags based on terminal sensors. For example, smart bracelets or mobile terminals can generate motion tags or location tags based on accelerometers, positioning data, etc. This type of solution usually only focuses on motion state or geographical location, lacking structured processing of interactive content such as dialogue topics, user intentions, image semantics, and environmental parameters. A third type of solution is the dialogue log system. This type of system usually saves raw text or voice records in chronological order and supports searching by time or keywords. However, this type of solution usually does not unify information such as interaction time, interaction location, point of interest attributes, interaction topics, user intentions, and environmental parameters into a memory tag set bound to the interaction memory, resulting in low retrieval efficiency and accuracy when subsequently searching historical memories by location, topic, intention, etc.
[0037] Furthermore, in existing solutions, location information and interaction semantic information are usually independent of each other, lacking consistency verification. When there is drift in terminal positioning or inaccurate matching of points of interest, the system may directly write incorrect location tags into the memory data, leading to false recalls in subsequent location-based memory retrieval. Therefore, existing technologies suffer from problems such as a single dimension of memory tags, lack of verification between location information and interaction semantic information, separation of dialogue memory and structured tags, and insufficient retrieval accuracy.
[0038] Therefore, this disclosure provides a memory management method based on memory tags. As an example, the memory management method can be applied to the following scenarios.
[0039] refer to Figure 1 The diagram illustrates a scenario application of the memory management method based on memory tags. User 100 can input interactive data through a terminal 200 connected to an intelligent agent. The terminal 200 can include electronic devices such as mobile phones, tablets, smartwatches, and in-vehicle terminals. The terminal 200 generates corresponding memory units based on the interactive data segments, and each memory unit can be associated with a memory tag. The terminal 200 identifies the semantic-geographic confidence verification result by matching the interactive semantic information and location information of the interactive data segments. This result serves as a memory tag of one dimension in the memory tag set to describe the corresponding memory unit. The memory unit and its associated memory tag set are stored in database 300.
[0040] like Figure 2 The diagram illustrates a flowchart of a memory management method based on memory tags, as shown in this embodiment. The steps in the memory management method can be executed locally on the terminal 200 connected to the intelligent agent, or they can be uploaded to the server for execution based on the interaction data between the terminal 200 connected to the intelligent agent and the user 100. Alternatively, some steps can be performed on the terminal 200 connected to the intelligent agent, and some steps on the server.
[0041] exist Figure 2 The memory management method based on memory tags includes:
[0042] Step S110: Obtain the current interaction data segment generated during the interaction between the agent and the user.
[0043] In some embodiments, the user may actively input the current interaction data segment to the agent, or the agent may collect the current interaction data segment with the user's permission during the interaction between the user and the agent.
[0044] In some embodiments, the interactive data segment is a multimodal type, and the modal type of the interactive data segment includes at least one of text, voice, and image.
[0045] In some examples, when the modal type of the interaction data segment is text, the current interaction data segment includes a text data segment, and the agent can directly use the text data segment as usable text; wherein, the text data segment can be obtained by acquiring the text content entered by the user in the human-computer interaction interface of the agent.
[0046] In other examples, when the modality type of the interaction data segment is speech, the interaction data segment includes a speech data segment, which can be acquired by an audio acquisition device on the agent.
[0047] In some other examples, when the modality type of the interaction data segment is image, the interaction data segment includes image data, which can be acquired by an image acquisition device on the agent or obtained by receiving image data sent by the user.
[0048] Step S120: Generate a corresponding memory unit based on the interactive data segment; wherein the memory unit is added to a memory database.
[0049] In some embodiments, when a start signal for a memory unit is received from the segmentation decision engine, the high-precision clock interface of the terminal operating system is invoked to record the absolute start time of the memory unit and generate a millisecond-level integer timestamp. When an end signal for the memory unit is received, its end time is recorded, and the total duration of the memory unit is calculated. The segmentation decision engine can be implemented as an independent module, responsible for determining the start and end of memory units using strategies such as silent timeout, semantic boundaries, and user-initiated operations. The timestamp of the memory unit's start or end time can be added to the memory tag set as a memory timestamp tag.
[0050] Step S130: Perform semantic recognition based on the interactive data segment to obtain interactive semantic information.
[0051] In some embodiments, the agent can perform semantic recognition on the interaction data segment according to the modality type of the interaction data segment to obtain the interaction semantic information corresponding to the memory unit. Thus, the interaction semantic information of the original interaction data segment can be obtained.
[0052] In some embodiments, Figure 3 A flowchart illustrating the modal feature label generation method is provided. Step S130 includes the following steps:
[0053] Step S131: Perform feature extraction processing on the interactive data segment according to the modal type of the interactive data segment to obtain the modal feature information corresponding to the interactive data segment; the modal type is at least one of text, speech and image.
[0054] In some embodiments, in response to an interaction data segment including a speech data segment, the agent may invoke a speech recognition model, such as a speech recognition model based on a recurrent neural network, to perform speech-to-text processing on the speech data segment, obtaining speech-to-text, and using the speech-to-text as usable text. Further, in response to an interaction data segment also including a text data segment, the agent may merge the text data segment and the speech-to-text in the current interaction data segment to obtain the session data in the current interaction data segment.
[0055] Step S132: Perform tag recognition processing on the modal feature information to obtain the modal feature tag corresponding to the modal feature information, and add the modal feature tag to the memory tag set corresponding to the memory unit.
[0056] In some embodiments, the agent may invoke a language understanding model (e.g., a lightweight BERT-Tiny model) to perform label recognition processing on the conversation data to obtain the conversation label corresponding to the current interaction data segment. The language understanding model can be configured as a classification model, and its output category is determined by a predefined label category system. For example, conversation labels include topic labels and / or intent labels. Topic labels are used to characterize the category of interaction content corresponding to the current interaction data segment, such as scheduling, casual conversation, emotional expression, and dining consumption. Intent labels are used to characterize the user's interaction purpose in the current interaction data segment, such as asking about the weather, sharing experiences, requesting modifications, setting reminders, or seeking comfort.
[0057] In some embodiments, if the interaction data segment includes image data, the agent can extract image feature information based on the image data and perform image semantic recognition on the image feature information to obtain corresponding image labels. Image labels are used to characterize the target object or scene category in the image data, such as scenery, food, documents, people, pets, etc. For example, a lightweight image classification model (such as MobileNetV3) is used to identify the main objects or scenes in the image data and directly output image labels.
[0058] In some embodiments, the agent can perform semantic recognition on one or more modal feature labels extracted from the same time period to obtain the label semantic information corresponding to the memory unit. For example, the label semantic information is a feature vector of the label semantics.
[0059] In some embodiments, in response to the existence of label semantic information of one modality type, the label semantic information is used as fused label semantic information. In response to the existence of label semantic information of multiple modality types, the agent can vectorize the multiple modality feature labels among topic labels, intent labels, and image labels to obtain a vectorized representation of the interaction semantic information. Specifically, the agent can input at least one of topic labels, intent labels, and image labels into a language understanding model to obtain the corresponding label semantic vectors, and perform vector fusion processing on each label semantic vector to obtain the fused semantic vector corresponding to the memory unit. The vector fusion processing includes at least one of weighted averaging, post-concatenation mapping, and attention-based weighted fusion.
[0060] For example, when the topic tag is "food and beverage consumption", the intent tag is "sharing experiences", and the image tag is "coffee drinks", the agent can convert the above tags into feature vectors of tag semantics and then fuse them to obtain the fused tag semantic vector of "coffee consumption scenario".
[0061] In some embodiments, if the interaction data segment includes an environmental perception data segment, the agent can acquire at least one environmental parameter corresponding to the environmental perception data segment and generate an environmental perception tag based on the environmental parameter. For example, the environmental parameters include light intensity and / or sound intensity. The environmental sensor includes a light sensor and / or a sound sensor. Light intensity can be acquired by the light sensor, and sound intensity can be acquired by a microphone or a sound sensor. Further, the agent can also map light intensity to light intensity tags such as dark, normal, and bright according to the numerical range to which the light intensity belongs, and can also map sound intensity to sound intensity tags such as quiet, normal, and noisy according to the numerical range to which the sound intensity belongs. For example, the light intensity tag and / or sound intensity tag are added to the memory tag set of the memory unit.
[0062] Step S140: Obtain the user's location information when the interaction data segment is formed, and obtain the location corresponding to the location information as a point of interest, and obtain the point of interest attribute information corresponding to the point of interest, the point of interest attribute information including the point of interest type and the point of interest location coordinates.
[0063] In some embodiments, based on the location information of the memory unit, corresponding candidate points of interest are determined in the point of interest database, and a matching distance between the candidate points of interest and the interactive location information is determined. In response to the matching distance satisfying a preset distance condition, the candidate points of interest are determined as the points of interest corresponding to the memory unit. The location information can characterize the user's geographical location, such as the user's original latitude and longitude information in forming the interactive data segment or its relative position information with other reference points.
[0064] In some embodiments, the location coordinates of the points of interest are obtained, and the location coordinates of the points of interest are added to the memory tag set as latitude and longitude coordinate labels.
[0065] In some embodiments, the point of interest attribute information further includes: point of interest name information, for example, the point of interest name information is "XX Community", and the point of interest name information can be added to the memory tag set as a location name tag.
[0066] For example, the location information can be obtained by acquiring the user's original latitude and longitude coordinates in the interactive data segment, and by calling the terminal's built-in reverse geocoding service, the original latitude and longitude coordinates can be converted into text address description information, such as "No. 1, Lane 1, Street 1, District A, City O". The text address description information is then added to the memory tag set as an address description tag.
[0067] Step S150: Obtain the semantic confidence of the memory unit based on the matching between the interest point type and the interaction semantic information, and obtain the distance confidence of the memory unit based on the distance between the interest point location coordinates and the positioning information.
[0068] In some embodiments, the semantic confidence can be determined by a preset semantic similarity matrix. Specifically, the agent can obtain the interest point type corresponding to the current memory unit, and obtain at least one of the topic tag, intent tag, and image tag in the interaction semantic information. Subsequently, the agent can query the preset semantic similarity matrix based on the interest point type and at least one of the topic tag, intent tag, and image tag to obtain the semantic similarity between the interest point type and the interaction semantic information, and use the semantic similarity as the semantic confidence corresponding to the current memory unit.
[0069] For example, the preset semantic similarity matrix can be used to record the degree of matching between different point-of-interest (POI) types and different semantic tags. For instance, when the POI type is "coffee shop" and the tag is "food and beverage consumption," a high semantic confidence can be obtained through the pre-defined semantic similarity matrix between the POI type and the tag. The preset semantic similarity matrix can be implemented as a pre-defined matrix M of POI type-tag type, where M∈R^{30×20}.
[0070] In other embodiments, the semantic confidence can be determined by the vector similarity between the fused feature vector and the interest point semantic vector. Specifically, the agent can input at least one of topic labels, intent labels, and image labels into the language understanding model to obtain corresponding label semantic vectors. Further, the agent can perform vector fusion processing on each of the label semantic vectors to obtain the fused semantic vector corresponding to the current memory unit. Exemplarily, the vector fusion processing may include at least one of weighted averaging, post-concatenation mapping, and attention-weighted fusion.
[0071] Furthermore, the agent can input the type of interest point into the language understanding model to obtain the semantic vector of the interest point.
[0072] Subsequently, the agent can calculate the vector similarity between the fused semantic vector and the interest point semantic vector, and determine the semantic confidence based on the vector similarity. For example, the vector similarity can be cosine similarity, and the cosine similarity calculation formula can be expressed as:
[0073] (1)
[0074] in, Indicates semantic confidence. This represents a semantic vector obtained by fusing any one or more of topic tags, intent tags, and image tags. This represents the semantic vector of interest points obtained based on the type of interest point.
[0075] For example, the current memory unit might have the following tags: topic label "food and beverage consumption," intent label "sharing experiences," image label "coffee and beverages," and interest type "coffee shop." The agent can obtain the corresponding fused semantic vector based on the topic label, intent label, and image label, and obtain the interest point semantic vector based on the interest point type. Therefore, semantic confidence can be determined based on the cosine similarity between the fused semantic vector and the interest point semantic vector, in a positive correlation.
[0076] In some embodiments, the agent can determine the distance confidence level of the memory unit based on the spatial distance between the original latitude and longitude coordinates in the positioning information and the location coordinates of the point of interest. Specifically, the agent can calculate the spherical distance between the original latitude and longitude coordinates and the location coordinates of the point of interest, and input the spherical distance into a preset distance decay function to obtain the distance confidence level. The distance confidence level is negatively correlated with the spherical distance, and the distance confidence level calculation formula can be expressed as:
[0077] (2)
[0078] Where D is the distance confidence score, and d is the spherical distance between the original latitude and longitude coordinates and the latitude and longitude coordinates of the point of interest. For example, the preset distance attenuation value can be the value of a preset distance matching threshold. For example, if the preset distance matching threshold is 20 meters, then the preset distance attenuation value is also 20 meters.
[0079] Step S160: Determine the semantic-geographic confidence verification result of the memory unit based on the semantic confidence and distance confidence of the memory unit, and add the semantic-geographic confidence verification result as a memory tag to the memory tag set of the memory unit.
[0080] In some embodiments, the distance confidence and the semantic confidence are weighted and fused to determine the semantic-geographic confidence verification result, which can be expressed as:
[0081] (3)
[0082] Where C represents the semantic-geographic confidence verification result. , To pre-determine the normalized weighting coefficients, satisfying: For example, 0.3 It is 0.7. The semantic confidence is obtained through equation (1), and the distance confidence D is calculated through equation (2).
[0083] In some embodiments, in response to the semantic-geographic confidence check result corresponding to the memory unit not meeting the semantic-geographic matching condition, a semantic-geographic conflict label is generated and added to the memory label set. The semantic-geographic matching condition includes a preset semantic-geographic matching threshold.
[0084] For example, if the point of interest category is "company" and the topic tag is "schedule", and the semantic-geographic confidence check result C ≥ 0.7 obtained according to Equation (3), it is judged as highly relevant, and C is directly output as the alignment confidence, and the semantic conflict is marked as "false". If 0.4 ≤ C < 0.7 is obtained according to Equation (3), it is judged as moderately relevant, and C is output, and the semantic conflict is marked as "false". In some other examples, if the point of interest category is "gym" and the dialogue topic is "emotional expression", if C < 0.4 is obtained according to Equation (3), it is judged as a semantic conflict. The two have no typical correlation, so the tag set of this memory unit is still retained, and C is output as the semantic-geographic confidence check result. However, the "semantic-geographic conflict tag" is set to "true" at the same time, and the conflict information is recorded. For example, the conflict information is: {"point of interest name": "gym", "topic tag": "emotional expression", "semantic-geographic confidence check result": 0.33"}.
[0085] In some embodiments, the agent may further combine at least one of topic tags, intent tags, image tags, and environmental awareness tags according to a preset tag combination rule to form a memory tag set corresponding to the current interaction data segment. For example, the memory tag set may include at least one of topic tags, intent tags, image tags, light intensity tags, and sound intensity tags. The structured perception parameter set can serve as a tag item in the memory tag set corresponding to the memory unit.
[0086] In some embodiments, key tags in the memory tag set are predetermined based on user instructions, and the memory tag set is stored in a tag index library. The tag index library is built based on at least one index corresponding to the key tags. For example, a time index can be built based on memory timestamp tags in the memory tag set; a spatial index can be built based on latitude and longitude coordinate tags in the memory tag set; a topic index can be built based on topic tags in the memory tag set; and an intent index can be built based on intent tags in the memory tag set. Based on the semantic-geographic confidence verification results and the spatial index retrieval results, the retrieved memory units can be distinguished by confidence level. Memory units whose semantic-geographic confidence verification results meet preset confidence conditions are displayed as high-confidence memory units, while memory units whose semantic-geographic confidence verification results do not meet preset confidence conditions are displayed as low-confidence memory units.
[0087] Step S170: In response to the user's current interaction request, query the memory database for target memory units with target memory tags that match the current interaction request, and form response information for the current interaction request accordingly.
[0088] In some embodiments, in response to a user’s current interaction request, a target memory tag set with a target memory tag that matches the current interaction request is queried in the tag index library, and a target memory unit is determined from the memory database based on the memory unit identity identifier corresponding to the target memory tag set, thereby forming response information to the current interaction request.
[0089] In some embodiments, the current interaction request includes a memory retrieval request. In response to the retrieval conditions corresponding to the memory retrieval request, including location retrieval conditions, a candidate memory tag set matching the location retrieval conditions is filtered from the memory tag set corresponding to the plurality of memory units. Based on the semantic-geographic confidence verification results of the candidate memory tag set, the location retrieval confidence level corresponding to the candidate memory tag set is determined. Based on the location retrieval confidence level, the target memory tag set is determined from the candidate memory tag set. Based on the memory unit identity identifier corresponding to the target memory tag set, the target memory unit is determined.
[0090] like Figure 4 The diagram shown illustrates a data processing structure based on memory tags in one embodiment of this disclosure.
[0091] In some embodiments, current interaction data segments generated during user-agent interaction are collected. These current interaction data segments may include at least one of text data segments, voice data segments, image data segments, and environmental perception data segments. Specifically, text data segments may be text content entered by the user in the interactive interface; voice data segments may be voice content entered by the user through a microphone; image data may be images uploaded by the user or images captured by a terminal connected to the intelligent agent; and environmental perception data segments may include data such as light intensity and sound intensity used to characterize the user's environment.
[0092] After acquiring the interaction data, the terminal connected to the intelligent agent processes it. Specifically, it can extract interactive semantics from text and / or voice data, map environmental parameters to environmental perception data segments, and record the location and time information corresponding to the interaction data. Further, a memory unit can be generated based on the processed interaction data. The memory unit includes a memory unit identifier, an interaction data segment summary extracted from the interaction data segment, and the original interaction data segment content. Subsequently, the terminal connected to the intelligent agent can store the memory unit in the memory database 310.
[0093] Furthermore, the terminal connected to the intelligent agent generates a corresponding memory tag set based on the processed interaction data. The memory tag set may include at least one of the following: memory unit identity identifier, topic tag, intent tag, image tag, light intensity tag, sound intensity tag, latitude and longitude coordinate tag, address description tag, location name tag, temporal context tag, semantic-geographic conflict tag, and semantic-geographic confidence verification result tag.
[0094] Simultaneously, the terminal connected to the intelligent agent can also perform semantic-geographic confidence verification. Specifically, the terminal connected to the intelligent agent can calculate semantic confidence based on the matching relationship between interactive semantic information and point-of-interest (POI) attribute information, and calculate distance confidence based on the distance between the location information and the POI location coordinates. Then, based on the semantic confidence and distance confidence, the semantic-geographic confidence verification result is obtained. The semantic-geographic confidence verification result can also be added to the memory tag set as a memory tag. If the semantic-geographic confidence verification result does not meet the preset matching conditions, a semantic-geographic confidence conflict tag can be generated and added to the memory tag set.
[0095] The terminal connected to the intelligent agent stores the memory tag set in the tag index library 320. For example, it can be stored in a local spatiotemporal index library (e.g., an SQLite database). The tag index library stores the correspondence between memory unit identifiers and memory tag sets. Specifically, a memory tag set can be written into the tag index library as a tag record and associated with the corresponding memory unit in the memory database 310 through the memory unit identifier.
[0096] The tag index library 320 can build an index based on key tag items in the memory tag set. Key tag items can be selected from a pre-set memory tag set; for example, the key tag items include at least one of time tags, spatial tags, topic tags, and intent tags. Accordingly, the tag index library can include indexes corresponding to key tags such as time indexes, spatial indexes, topic indexes, and intent indexes to support subsequent memory retrieval based on time, location, topic, or intent.
[0097] In some embodiments, when user 100 initiates a current interaction request, the terminal connected to the intelligent agent can query a matching set of memory tags in the tag index library based on the current interaction request and obtain the corresponding memory unit identity identifier. Subsequently, the terminal connected to the intelligent agent can read the corresponding memory unit content from the memory database 310 according to the memory unit identity identifier, and generate response information based on the memory unit content to feed back to user 100.
[0098] For example, at 14:32:15 on May 9, 2026, a user had a continuous conversation with a smart assistant on their smartphone at home (in a residential community). The conversation was as follows: User (14:32:15, text): "Check the weather for me tomorrow." Smart agent (14:32:20): "Tomorrow will be cloudy, with temperatures ranging from 19 to 30 degrees Celsius, suitable for going out." User (14:32:27, text): "What about the day after tomorrow? Will it rain?" Smart agent (14:32:32): "The day after tomorrow will be sunny turning cloudy, with temperatures ranging from 19 to 33 degrees Celsius." User (14:32:35, voice): "Okay, thank you." The entire conversation lasted 20 seconds, with the user using both text and voice input, and the smart agent providing multiple rounds of responses. The user also sent a photo of a cloudy day outside their window. The user's speech was converted to text, and all text input was merged. The dialogue topic was identified as "information query" and the intent as "weather query." The image sent by the user was identified as "scenery." The light sensor readings (120 lux) and microphone background noise (45 dB) were read and mapped to the "normal" level. Location information such as GPS coordinates was then obtained, and the specific address was obtained through reverse geocoding. This address was matched against "X neighborhood" in the offline point of interest database, with the point of interest category being "residential," and the distance between the point of interest location coordinates and the location information being 12 meters. The semantic-geographic confidence check result calculated using the above information was 0.82, indicating no semantic conflict. All the obtained tags were assembled into a memory tag set, which, for example, can be set to JSON format. The memory tag set was inserted as a record into the local tag index library, using the memory unit's identity as the primary key, and indexes such as time, location, topic, and intent were automatically created. Therefore, the multi-dimensional memory tag set corresponding to the memory unit has been completely stored. Through this memory tag set, relevant information of the memory unit (such as time, place, topic, intention, environment, etc.) can be obtained. If you need to view the full text of the original dialogue, you can query the full text of the original dialogue stored in the memory database through the memory unit's identity identifier.
[0099] In summary, in this embodiment, the memory database 310 is used to store the content of memory units, and the tag index library is used to store the memory tag set and its index relationship. The intelligent agent can first quickly locate the identity of the target memory unit through the tag index library 320, and then read the content of the target memory unit from the memory database 310, thereby improving the efficiency of retrieval and response generation based on memory tags.
[0100] Figure 5 This illustration shows a flowchart of another embodiment of a method for generating temporal context sequences based on memory association, where memory units can construct context sequences of each other based on memory association. Compared to previous embodiments, this embodiment describes the process of constructing a memory spatiotemporal network from memory units stored in a memory database. The method includes the following steps:
[0101] Step S210: Based on the interactive semantic information corresponding to each memory unit, obtain the semantic correlation degree between the multiple memory units, and according to the time information and the location information corresponding to each memory unit, obtain the spatiotemporal correlation degree between the multiple memory units.
[0102] The spatiotemporal correlation degree includes: temporal proximity degree and spatial correlation degree.
[0103] In some embodiments, the temporal proximity between multiple memory units is determined based on the temporal information of each memory unit. The temporal proximity calculation formula can be expressed as:
[0104] (4)
[0105] in, This represents the value of the temporal proximity between memory unit A and memory unit B. , These are the absolute timestamps of memory units A and B, respectively. This is the time decay coefficient, which is used to control the rate at which time proximity decays as the interval increases.
[0106] Based on the location information of each memory unit, the spatial correlation degree between multiple memory units is determined. The formula for calculating the spatial correlation degree can be expressed as:
[0107] (5)
[0108] in, The value represents the spatial correlation degree between memory unit A and memory unit B. , These are location identifiers obtained by mapping the geographical coordinates of memory units A and B after spatial clustering and semantic feature inference. From historical data, from location Transfer to location frequency, count( Any) is from location Total frequency of transfers to any location.
[0109] Based on the interactive semantic information of each memory unit, the semantic association degree between multiple memory units is determined. The formula for calculating the semantic association degree can be expressed as:
[0110] (6)
[0111] in, The value represents the semantic association degree between memory unit A and memory unit B. For the semantic feature vector of memory unit A, This is the semantic feature vector of memory unit B.
[0112] In some embodiments, the agent can determine the location identifier corresponding to a memory unit based on its geographic location information. Specifically, the agent can first perform spatial clustering on the geographic location coordinates of multiple memory units to group memory units whose spatial distance meets a preset distance condition into the same location cluster. Subsequently, the agent can obtain semantic feature information corresponding to each memory unit in the same location cluster and identify location semantic clues corresponding to the location cluster based on the semantic feature information. Location semantic clues may include at least one of location name, scene keywords, scene type information, and user-defined location naming information. The agent can determine the discrete location identifier corresponding to the location cluster based on the location semantic clues and use the discrete location identifier as the location identifier corresponding to the memory unit within the location cluster.
[0113] For example, if the semantic features of multiple memory units within the same location cluster contain semantic cues such as "company," "office," "meeting room," and "project meeting," then the location identifier corresponding to that location cluster can be identified as an office location identifier. Similarly, if the semantic features of multiple memory units within the same location cluster contain semantic cues such as "hospital," "registration," "doctor," "examination," and "follow-up examination," then the location identifier corresponding to that location cluster can be identified as a medical location identifier. Thus, continuous geographic coordinates can be converted into semantically meaningful location identifiers, facilitating subsequent calculations of the spatial relationships between memory units.
[0114] Step S220: Based on the semantic correlation degree and spatiotemporal correlation degree, determine the memory correlation degree between each memory unit.
[0115] In some embodiments, the memory association degree can be the result of fusing at least two of the semantic association degree, the temporal proximity degree, and the spatial association degree, such as by calculating a weighted average sum, to obtain the memory association degree between multiple memory units.
[0116] For example, the formula for calculating memory association degree can be expressed as:
[0117] (7)
[0118] In equation (7), The value of memory association. , , To pre-determine the normalized weighting coefficients, satisfying: .
[0119] Step S230: The memory units whose memory association with the target memory unit meets the association condition are identified as the associated memory units corresponding to the target memory unit.
[0120] In some embodiments, the agent can, for any memory unit, traverse other memory units and determine whether the memory association degree between the other memory units and the memory unit satisfies the association degree condition. In response to a memory association degree greater than a preset association degree threshold, the agent identifies the corresponding other memory units as associated memory units corresponding to the memory unit.
[0121] Step S240: Sort the associated memory units along the temporal direction to form the temporal context sequence of the target memory unit, and add the temporal context sequence of the memory unit as a temporal context tag to the memory tag set of the target memory unit.
[0122] In some embodiments, the agent can sort the associated memory units temporally based on their temporal information to form a temporal context sequence corresponding to each memory unit. The temporal context sequence may include the memory unit identifier of each associated memory unit, or it may include the memory unit identifier of each associated memory unit and the time interval between adjacent associated memory units. The agent can add the temporal context sequence as a tag to the memory tag set of the memory unit, so that it can subsequently determine the associated memory units with contextual association with the memory unit based on the temporal context sequence.
[0123] In some embodiments, when a target memory unit is subsequently retrieved based on a user's interaction request, the agent may determine associated memory units that have a contextual association with the target memory unit based on the temporal context sequence, in order to assist in displaying the contextual information of the target memory unit or to assist in generating response information to the user's interaction request.
[0124] In one embodiment of this disclosure, the connection relationship between memory units in the same temporal context sequence along the temporal direction can be used to form a memory association link, and the collection of the memory association links forms a memory spatiotemporal network.
[0125] In some embodiments, each memory association link may have a corresponding link weight. The link weight is positively correlated with the memory association degree between the corresponding memory units; for example, the value of the memory association degree can be directly used as the link weight. Through the link weight, the agent can preferentially select memory association links with higher weights when traversing memories subsequently, thereby obtaining the target memory association links related to the user's current request.
[0126] Thus, the intelligent agent can not only simply arrange multi-day discrete dialogue records in chronological order, but also organize memory units related to the same matter into memory association links with upstream and downstream connections based on semantic and spatiotemporal associations.
[0127] In some embodiments, memory association links that meet the correlation criteria but are not selected into the target memory association link can be retained as branch links in the memory spatiotemporal network and displayed as expandable auxiliary networks in the response information to the memory tracing request. When the aforementioned local memory association links are not selected into the current target memory association link, they can be displayed as collapsed branch links.
[0128] In some examples, as shown in Figure 6(a), a schematic diagram of the structure of the target memory association link in the memory spatiotemporal network according to an embodiment of the present disclosure is provided. Since the memory association degree between memory unit N3 and memory unit N5 is greater than that between memory unit N2 and memory unit N5, the memory association link between memory unit N2 and memory unit N5 is folded and unfolded as a branch link, and “N2→N3→N5→N6→N8” is taken as the target memory association link B. In addition, the target memory association link B can be highlighted on the human-computer interaction interface. In other examples, as shown in Figure 6(b), a schematic diagram of the structure of the target memory association link in the memory spatiotemporal network according to another embodiment of the present disclosure is provided. Since the memory association degree between memory unit N3 and memory unit N5 is less than that between memory unit N2 and memory unit N5, the memory association links between memory unit N3 and memory unit N2 and between memory unit N3 and memory unit N5 are folded and unfolded as branch links. The memory association links between memory unit N5 and memory unit N6 and between memory unit N6 and memory unit N8 are also folded and unfolded as branch links. "N2→N5→N8" is taken as the target memory association link C. In addition, the target memory association link C can be highlighted on the human-computer interaction interface.
[0129] Therefore, the target memory association link can further compress weakly associated or redundant units based on memory association degree while retaining the original time sequence association. This allows the memory content with a higher degree of association with the target memory unit to be called first when responding to current interactive requests such as memory tracing, memory question answering, and memory retrieval, thereby improving the contextual effectiveness of the response information.
[0130] Therefore, this embodiment enables the intelligent agent to not only return the target memory unit when the user initiates an interaction request, but also to determine the upstream background and downstream development of the target memory unit based on the memory spatiotemporal network, thereby generating response information with event evolution relationship and traceability.
[0131] This disclosure provides another embodiment in which the memory spatiotemporal network can be continuously updated with new interactions between the user and the agent. Compared to previous embodiments, this embodiment describes the process of adding new memory units to the existing memory spatiotemporal network. In response to the agent generating a new memory unit for the user, this disclosure maintains the connection relationships between existing memory units in the memory spatiotemporal network unchanged, and establishes temporally sequential connection relationships between the new memory unit and existing memory units that meet the correlation conditions to update the set of memory association links, forming the updated memory spatiotemporal network.
[0132] like Figure 7 The diagram illustrates a flowchart of a method for adding new memory units in a memory-space-time network according to an embodiment of this disclosure. The method includes the following steps:
[0133] S310: In response to the agent generating a new memory unit for the user, based on the spatiotemporal information of the new memory unit, a spatiotemporal candidate unit retrieval request is initiated to at least one of the storage partitions, so that each of the storage partitions retrieves a set of first candidate memory units that are adjacent to the new memory unit in the time dimension and / or spatial dimension from the memory units already stored in each of the storage partitions based on the spatiotemporal index.
[0134] For example, newly added memory units can be retrieved in the time dimension to find candidate memory units that are temporally adjacent, and in the spatial dimension to find candidate memory units that are spatially adjacent. The union of the set of candidate memory units that are temporally adjacent and the set of candidate memory units that are spatially adjacent is obtained to get the first set of candidate memory units.
[0135] The spatiotemporal index may include at least one of a time index, a spatial index, or a combined time-space index. Since memory units can be stored distributed across different storage partitions, a spatiotemporal candidate unit retrieval request can be sent to multiple storage partitions, allowing each partition to retrieve candidate units based on its local spatiotemporal index. Therefore, even if an existing memory unit spatiotemporally adjacent to the newly added memory unit is not stored in the same storage partition, it can still be recalled as a first candidate memory unit.
[0136] S320: Based on the semantic feature information extracted from the newly added memory unit, initiate a semantic retrieval request to at least one of the storage partitions, so that each of the storage partitions can retrieve a set of second candidate memory units that match the newly added memory unit in the semantic dimension from the memory units already stored in each of the storage partitions based on the semantic feature vector index.
[0137] The semantic feature vector index can be used for retrieval based on the similarity between semantic feature vectors. For example, a second set of candidate memory units with semantic matching can be determined based on the cosine similarity between the semantic feature vector of a newly added memory unit and the semantic feature vector of an already stored memory unit. This allows for the recall of existing memory units that belong to the same event, the same task chain, or have a similar semantic theme as the newly added memory unit.
[0138] S330: Based on the first candidate memory unit set and / or the second candidate memory unit set, form a candidate memory unit set corresponding to the newly added memory unit.
[0139] In some embodiments, the candidate memory cell set can serve as a candidate range for subsequent calculation of the correlation of newly added memory cells, thereby reducing the computational overhead caused by calculating the correlation of each of all stored memory cells one by one.
[0140] In the first example, the first set of candidate memory units can be used as the set of candidate memory units corresponding to the newly added memory unit.
[0141] In the second example, the second set of candidate memory units can be used as the set of candidate memory units corresponding to the newly added memory unit.
[0142] In the third example, the first candidate memory unit set and the second candidate memory unit set can be merged and deduplicated to obtain the candidate memory unit set corresponding to the newly added memory unit. For example, the candidate memory unit set corresponding to the newly added memory unit can be obtained by taking the union of the first candidate memory unit set and the second candidate memory unit set.
[0143] In the fourth example, the priorities of the first candidate memory set and the second candidate memory set can be preset respectively. In response to the first candidate memory set having a higher priority than the second candidate memory set, when the first candidate memory set is empty, the second candidate memory set is used as the candidate memory set corresponding to the newly added memory unit. When the first candidate memory set is not empty, the first candidate memory set is used as the candidate memory set corresponding to the newly added memory unit. In response to the first candidate memory set having a lower priority than the second candidate memory set, when the second candidate memory set is empty, the first candidate memory set is used as the candidate memory set corresponding to the newly added memory unit. When the second candidate memory set is not empty, the second candidate memory set is used as the candidate memory set corresponding to the newly added memory unit.
[0144] For example, if the new memory unit corresponds to a location-strongly related event such as an offline visit or a hospital visit, the first candidate memory unit set can be selected first; if the new memory unit corresponds to a location-weakly related event such as a document modification or a project discussion, the second candidate memory unit set can be selected first.
[0145] S340: Based on the new memory association degree between the newly added memory unit and each candidate memory unit in the candidate memory unit set, determine the target associated memory unit whose new memory association degree meets the association degree condition from the candidate memory unit set.
[0146] In some embodiments, the new memory association degree can be determined based on one or more of the semantic association degree, temporal proximity, and spatial association degree between the new memory unit and the candidate memory unit. For example, the value of the new memory association degree can be calculated based on equations (4), (5), (6), and (7). If the value of the new memory association degree meets the preset edge-building threshold, the corresponding candidate memory unit is determined as the target associated memory unit; if the value of the new memory association degree does not meet the preset edge-building threshold, no memory association link is established between the two; if the value of the new memory association degree between the new memory unit and any candidate memory unit does not meet the preset edge-building threshold, the new memory unit can be added to the memory spatiotemporal network as an independent memory unit.
[0147] S350: Establish the connection relationship between the newly added memory unit and each of the target associated memory units to update the set of memory association links to form the updated memory spatiotemporal network.
[0148] In some embodiments, a temporal connection can be established based on the temporal order between the newly added memory unit and the target associated memory unit, and the association degree of the newly added memory can be used as the link weight corresponding to the connection. Therefore, after a new memory unit is added, only the vertices, memory association links, and link weights in the memory spatiotemporal network are updated, without directly triggering a main network update. When a user subsequently initiates a memory tracing request, the agent can dynamically determine the corresponding target memory association link based on the updated memory spatiotemporal network.
[0149] Therefore, this embodiment recalls candidate neighboring units from multiple storage partitions through spatiotemporal indexing and semantic feature vector indexing, and establishes a connection relationship between the new memory unit and the existing memory unit based on the new memory association degree, so that the new memory unit can be incorporated into the memory spatiotemporal network, rather than simply added to the storage as an isolated historical record, thereby improving the accuracy of subsequent memory retrieval, memory question answering and memory tracing.
[0150] like Figure 8The diagram illustrates a module schematic of a memory management system based on memory tags according to an embodiment of this disclosure. It should be noted that the principles and technical implementation of the memory management system based on memory tags can be referenced from the memory management methods based on memory tags in previous embodiments; therefore, they will not be repeated in this embodiment.
[0151] The memory management system 400 based on memory tags includes:
[0152] The acquisition module 401 is used to acquire the current interaction data segment generated during the interaction between the intelligent agent and the user, and to acquire the user's location information when the interaction data segment is formed.
[0153] The memory unit generation module 402 is used to generate corresponding memory units based on the interactive data segment; wherein the memory units are added to a memory database.
[0154] The semantic recognition module 403 is used to perform semantic recognition based on the interactive data segment to obtain interactive semantic information.
[0155] The point of interest attribute information determination module 404 is used to obtain the location corresponding to the location information as the point of interest based on the user's location information when forming the interaction data segment, and to obtain the point of interest attribute information corresponding to the point of interest, wherein the point of interest attribute information includes the point of interest type and the point of interest location coordinates.
[0156] The memory tag set generation module 405 is used to obtain the semantic confidence of the memory unit based on the matching between the interest point type and the interaction semantic information, and to obtain the distance confidence of the memory unit based on the distance between the interest point location coordinates and the location information. Based on the semantic confidence and distance confidence of the memory unit, the semantic-geographic confidence verification result of the memory unit is determined, and the semantic-geographic confidence verification result is added as a memory tag to the memory tag set of the memory unit.
[0157] The response generation module 406 is used to respond to the user's current interaction request by querying the memory database for target memory units with target memory tags that match the current interaction request, and thereby forming response information for the current interaction request.
[0158] It should be noted that, in Figure 8The various functional modules in the embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented in software, they can be implemented, in whole or in part, in the form of a computer program or instruction product. A computer program or instruction product includes one or more computer programs or instructions. When a computer program or instruction is loaded and executed on a computer, it produces, in whole or in part, the flow or function according to this disclosure. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.
[0159] and, Figure 8 The system disclosed in the embodiments can be implemented through other modularization methods. The system embodiments shown above are merely illustrative. For example, the modularization described is only a logical functional division, and in actual implementation, there may be other division methods. For example, a group of modules or modules may be combined or dynamically integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical or other forms.
[0160] in addition, Figure 8 The functional modules and sub-modules in the embodiments can be dynamically integrated within a single processing unit, or each module can exist physically independently, or two or more modules can be dynamically integrated within a single unit. These dynamic units can be implemented in hardware or as software functional modules. If these dynamic units are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a hard disk, or an optical disk, etc.
[0161] It should be specifically noted that the flowchart representations of the embodiments described above in this disclosure can be understood as representing a module, segment, or portion of code comprising one or more executable instructions configured to implement a specific logical function or process. Furthermore, the scope of the preferred embodiments of this disclosure includes other implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved.
[0162] For example, Figure 2 , 3 The order of the steps in the method embodiments 5 and 7 may vary in specific scenarios and is not limited to the above representation.
[0163] like Figure 9The diagram shown illustrates the structure of a mobile terminal in one embodiment of this disclosure.
[0164] The mobile terminal 500 may be, for example, a processing terminal, such as a server, desktop computer, laptop computer, tablet computer, smartphone, or other terminal.
[0165] The mobile terminal 500 includes a bus 501, a processor 502, and a memory 503. The processor 502 and the memory 503 can communicate via the bus 501. The memory 503 can store computer programs or instructions. The processor 502 implements the method flow or function of the previous embodiments by running the computer program or instructions in the memory 503, for example... Figure 2 , 3 5.
[0166] Bus 501 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, although only one thick line is used in the diagram, this does not indicate that there is only one bus or one type of bus.
[0167] In some embodiments, processor 502 may be implemented as a central processing unit (CPU), microprocessor unit (MCU), system on chip (System on Chip), or field-programmable array (FPGA). Memory 503 may include volatile memory for temporary data storage during program execution, such as random access memory (RAM).
[0168] The memory 503 may also include non-volatile memory for data storage, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state disk (SSD).
[0169] In some embodiments, the mobile terminal 500 may further include a communicator 504. The communicator 504 is used for communication with external devices. In specific examples, the communicator 504 may include one or more wired and / or wireless communication circuit modules. For example, the communicator 504 may include one or more of the following: a wired network card, a USB module, a serial interface module, etc. The wireless communication protocols followed by the wireless communication module include, for example, Nearfield Communication (NFC) technology, Infrared (IR) technology, Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), Bluetooth (BT), Global Navigation Satellite System (GNSS), etc., one or more of the following.
[0170] This disclosure also provides a computer-readable storage medium storing a computer program or instructions, which, when run, implement the method flow or function of any of the previous embodiments.
[0171] That is, the method steps in the above embodiments are implemented as software or computer code that can be stored in a recording medium (such as CD ROM, RAM, floppy disk, hard disk or magneto-optical disk), or implemented as computer code that is originally stored in a remote recording medium or a non-transitory machine-readable medium and will be stored in a local recording medium after being downloaded via a network, so that the method represented herein can be stored in such software processing on a recording medium using a general-purpose computer, a special processor or programmable or special hardware (such as ASIC or FPGA).
[0172] This disclosure may also provide a computer program product, comprising one or more computer programs or instructions, which, when run, perform all or part of the processes or functions described in this disclosure. The computer program product includes one or more computer programs or instructions.
[0173] Computer programs or instructions can be stored in a readable storage medium or transferred from one readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The readable storage medium can be any available medium capable of access, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; an optical medium, such as a digital video optical disc; or a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or it can include both volatile and non-volatile types of storage media.
[0174] In summary, this disclosure provides a memory management method and mobile terminal based on memory tags. The method includes: acquiring a current interaction data segment generated during the interaction between an agent and a user; generating a corresponding memory unit based on the interaction data segment; wherein the memory unit is added to a memory database; performing semantic recognition based on the interaction data segment to obtain interaction semantic information; acquiring the user's location information when the interaction data segment is formed, and obtaining the location corresponding to the location information as a point of interest, and acquiring the point of interest attribute information corresponding to the point of interest, the point of interest attribute information including the point of interest type and the point of interest location coordinates; and based on the point of interest... The semantic confidence of the memory unit is obtained by matching the type with the interactive semantic information, and the distance confidence of the memory unit is obtained based on the distance between the location coordinates of the point of interest and the location information. The semantic-geographic confidence verification result of the memory unit is determined based on the semantic confidence and distance confidence, and the semantic-geographic confidence verification result is added as a memory tag to the memory tag set of the memory unit. In response to the user's current interaction request, a target memory unit with a target memory tag matching the current interaction request is queried in the memory database, and a response information to the current interaction request is formed accordingly. This disclosure improves the user's interactive experience with the memory unit by generating a memory tag set containing semantic-geographic confidence verification results for the memory unit, enabling users to determine whether the interactive semantics in the memory unit match the corresponding location information based on the semantic-geographic confidence verification results.
[0175] The above embodiments are merely illustrative of the principles and effects of this disclosure and are not intended to limit this disclosure. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this disclosure. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this disclosure should still be covered by the protection scope of this disclosure.
Claims
1. A memory management method based on memory tags, characterized in that, include: Acquire the current interaction data segment generated during the interaction between the intelligent agent and the user; A corresponding memory unit is generated based on the interactive data segment; wherein, the memory unit is added to a memory database; Semantic recognition is performed based on the interactive data segment to obtain interactive semantic information; wherein, the semantic recognition based on the interactive data segment to obtain interactive semantic information includes: performing feature extraction processing on the interactive data segment according to the modality type of the interactive data segment to obtain modality feature information corresponding to the interactive data segment; the modality type is at least one of text, speech, and image; performing tag recognition processing on the modality feature information to obtain modality feature tags corresponding to the modality feature information, and adding the modality feature tags to the memory tag set corresponding to the memory unit; the modality feature tags include conversation tags and / or image tags based on text or speech, and the conversation tags include topic tags and / or intent tags; The location information of the user when the interaction data segment is formed is obtained, and the location corresponding to the location information is obtained as a point of interest. The point of interest attribute information corresponding to the point of interest is obtained, and the point of interest attribute information includes the point of interest type and the point of interest location coordinates. The semantic confidence of the memory unit is obtained based on the matching between the interest point type and the interaction semantic information, and the distance confidence of the memory unit is obtained based on the distance between the interest point location coordinates and the positioning information; wherein, obtaining the semantic confidence of the memory unit based on the matching between the interest point type and the interaction semantic information includes: performing semantic recognition on modal feature labels extracted from one or more modal types within the same time period to obtain corresponding label semantic information; in response to the existence of label semantic information of multiple modal types, fusing the label semantic information to obtain fused label semantic information; performing semantic recognition on the interest point information corresponding to the interest point to obtain interest point semantic information; and obtaining the semantic confidence of the memory unit based on the semantic similarity between the fused label semantic information and the interest point semantic information; The semantic-geographic confidence verification result of the memory unit is determined based on the semantic confidence and distance confidence of the memory unit, and the semantic-geographic confidence verification result is added as a memory tag to the memory tag set of the memory unit. In response to a user's current interaction request, the system queries the memory database for target memory units with target memory tags that match the current interaction request, and forms response information for the current interaction request accordingly.
2. The memory management method based on memory tags according to claim 1, characterized in that, Also includes: Based on the interactive semantic information corresponding to each memory unit, the semantic correlation degree between the multiple memory units is obtained, and based on the time information and the location information corresponding to each memory unit, the spatiotemporal correlation degree between the multiple memory units is obtained; Based on the semantic and spatiotemporal correlation, the memory correlation between each memory unit is determined; The memory units whose memory association with the target memory unit satisfies the association condition are identified as the associated memory units corresponding to the target memory unit. The associated memory units are sorted along the temporal direction to form the temporal context sequence of the target memory unit, and the temporal context sequence is added as a memory tag to the memory tag set of the target memory unit.
3. The memory management method based on memory tags according to claim 2, characterized in that, The method further includes: Based on the connection relationship between memory units in the same temporal context sequence along the temporal direction, memory association links are formed, and the collection of memory association links forms a memory spatiotemporal network.
4. The memory management method based on memory tags according to claim 1, characterized in that, The interactive data segment further includes an environment-aware data segment, and the method further includes: The environmental perception data segment is mapped to obtain an environmental perception tag, and the environmental perception tag is added to the memory tag set corresponding to the memory unit; The environmental perception data segment includes at least one environmental parameter collected by environmental sensors.
5. The memory management method based on memory tags according to claim 1, characterized in that, Also includes: Determine the key tags in the memory tag set based on the user's instructions; The memory tag set is stored in a tag index library; wherein, the tag index library establishes at least one index based on the key tags.
6. The memory management method based on memory tags according to claim 1, characterized in that, The step of determining the semantic-geographic confidence verification result of the memory unit based on the semantic confidence and distance confidence of the memory unit includes: The distance confidence and the semantic confidence are weighted and fused to determine the semantic-geographic confidence verification result.
7. The memory management method based on memory tags according to claim 1, characterized in that, The step of responding to a user's current interaction request by querying the memory database for a target memory unit with a target memory tag matching the current interaction request includes: In response to the search conditions corresponding to the current interaction request, including location search conditions, a set of candidate memory tags that match the location search conditions is filtered from the memory tag set corresponding to multiple memory units; Based on the semantic-geographic confidence verification results of the candidate memory tag set, the location retrieval credibility corresponding to the candidate memory tag set is determined; Based on the location retrieval credibility, the target memory tag set is determined from the candidate memory tag set; The target memory unit is determined based on the memory unit identity identifier corresponding to the target memory tag set.
8. A mobile terminal, characterized in that, include: Processor and memory; The memory stores computer programs or instructions; The processor is configured to run the computer program or instructions to perform the memory management method based on memory tags as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for interaction, equipment, storage medium and program product
CN120848723A
Memory enhancement and reasoning method and system for intelligent interaction
CN121998106A