A document label determination method, system and electronic device

By analyzing user interaction data with intelligent agents to generate personalized document tags, this technology solves the problem that existing tagging systems cannot dynamically adapt to user needs, improving the efficiency and accuracy of document retrieval. It is suitable for enterprise knowledge bases and academic retrieval systems.

CN122489765APending Publication Date: 2026-07-31LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LENOVO (BEIJING) LTD
Filing Date
2026-04-30
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing document tagging systems cannot dynamically and accurately generate tags based on users' actual usage habits and behavioral patterns, resulting in low retrieval efficiency and accuracy, especially when dealing with massive amounts of heterogeneous documents and personalized needs.

Method used

By analyzing user interaction data with the intelligent agent, personalized document tags are generated. Combined with the deep semantic understanding capabilities of the large language model, the tag content is dynamically updated to adapt to user needs.

Benefits of technology

It improves the efficiency and accuracy of document retrieval, reduces the cost of manual annotation, and continuously optimizes tag quality through user feedback mechanisms, making it suitable for scenarios such as enterprise knowledge bases and academic retrieval systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122489765A_ABST
    Figure CN122489765A_ABST
Patent Text Reader

Abstract

This application provides a document tag determination method, system, and electronic device. The method includes obtaining interaction behavior data associated with a target document generated by a user during interaction with an intelligent agent, wherein the target document is a document retrieved by the intelligent agent from a knowledge base; obtaining first tag content corresponding to the target document based on the interaction behavior data; and setting a first tag for the target document in the knowledge base according to the first tag content, wherein the first tag is used for retrieving the target document.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a document tag determination method, system, and electronic device. Background Technology

[0002] In the knowledge management module, the document tagging system, as a core organizational structure, directly impacts knowledge retrieval efficiency, content management effectiveness, and collaboration fluency. Especially in environments with massive amounts of document data, an efficient tagging system can transform scattered information into structured knowledge, significantly enhancing the value of information utilization. Some solutions involve pre-setting a fixed tagging system, with documents possessing fixed static tags. However, as user scenarios change, a single document tagging system struggles to adapt to user needs. Summary of the Invention

[0003] This application provides a document tag determination method, system, and electronic device.

[0004] The embodiments of this application adopt the following technical solution: a document tag determination method, the method comprising: Obtain interactive behavior data associated with a target document generated by the user during interaction with the intelligent agent, wherein the target document is a document retrieved by the intelligent agent from a knowledge base; Based on the interaction behavior data, obtain the first tag content corresponding to the target document; A first tag is set in the knowledge base based on the content of the first tag, wherein the first tag is used to retrieve the target document.

[0005] In some embodiments, obtaining the interaction behavior data associated with the target document generated by the user during interaction with the intelligent agent includes: The system obtains question data submitted by a user to the agent, wherein the target document is retrieved from the knowledge base based on the question data.

[0006] In some embodiments, obtaining the interaction behavior data associated with the target document generated by the user during interaction with the intelligent agent includes: Obtain interactive data generated by the user's actions or information feedback on the target document during interaction with the intelligent agent.

[0007] In some embodiments, obtaining the interaction behavior data associated with the target document generated by the user during interaction with the intelligent agent includes: Obtain multi-turn dialogue data between the user and the agent, wherein the multi-turn dialogue data is associated with the target document.

[0008] In some embodiments, obtaining the first tag content corresponding to the target document based on the interaction behavior data includes: The interaction data is preprocessed to remove invalid conversations and noise data between the user and the agent; The preprocessed interactive behavior data is set as an interactive sequence; The target input information is determined based on the interaction sequence; The target input information is input into the target model to obtain the first tag content generated by the target model.

[0009] In some embodiments, determining the target input information based on the interaction sequence includes: Behavioral analysis is performed on the interaction sequence to determine the usage scenario information; The target input information is determined based on the usage scenario information and the target document.

[0010] In some embodiments, the first tag of the target document can be updated based on target interaction data between the user and the target document with the first tag.

[0011] In some embodiments, based on the interaction behavior data, obtaining the first tag content corresponding to the target document includes: Based on the interaction behavior data, the user type is determined; Based on the user type, determine the first tag content corresponding to the user type for the target document.

[0012] This application also provides a document tag determination system, including: The data acquisition module is configured to acquire interactive behavior data associated with a target document generated by the user during interaction with the intelligent agent, wherein the target document is a document retrieved by the intelligent agent from a knowledge base; The tag content determination module is configured to obtain the first tag content corresponding to the target document based on the interaction behavior data. A setting module is configured to set a first tag for the target document in the knowledge base based on the content of the first tag, wherein the first tag is used to retrieve the target document.

[0013] This application embodiment provides an electronic device, including a memory and a processor. The memory stores an executable program, and the processor is used to execute: Obtain interactive behavior data associated with a target document generated by the user during interaction with the intelligent agent, wherein the target document is a document retrieved by the intelligent agent from a knowledge base; Based on the interaction behavior data, obtain the first tag content corresponding to the target document; A first tag is set in the knowledge base based on the content of the first tag, wherein the first tag is used to retrieve the target document. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 A flowchart illustrating the method for determining document tags for this application.

[0016] Figure 2 For this application Figure 1 A flowchart of one embodiment of step S2.

[0017] Figure 3 For this application Figure 2 A flowchart of one embodiment of step S23.

[0018] Figure 4 Define the system's structural framework for the document tags in this application.

[0019] Figure 5 This is a structural block diagram of the electronic device of this application. Detailed Implementation

[0020] Various embodiments and features of this application are described herein with reference to the accompanying drawings.

[0021] It should be understood that various modifications can be made to the embodiments described herein. Therefore, the above description should not be considered as limiting, but merely as an example of embodiments. Other modifications within the scope and spirit of this application will be apparent to those skilled in the art.

[0022] The accompanying drawings, which are included in and form part of this specification, illustrate embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, serve to explain the principles of the present application.

[0023] These and other features of this application will become apparent from the following description of preferred forms of embodiments given as non-limiting examples, with reference to the accompanying drawings.

[0024] It should also be understood that although this application has been described with reference to some specific examples, those skilled in the art can certainly implement many other equivalent forms of this application.

[0025] The above and other aspects, features and advantages of this application will become more apparent when taken in conjunction with the accompanying drawings and in view of the following detailed description.

[0026] Specific embodiments of this application are described thereafter with reference to the accompanying drawings; however, it should be understood that the claimed embodiments are examples of this application and can be implemented in various ways. Well-known and / or repeated functions and structures are not described in detail to avoid unnecessary or redundant details that could obscure the application. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but rather serve as the basis and representative basis for the claims to teach those skilled in the art to use this application in a variety of substantially any suitable detailed structures.

[0027] This specification may use the phrases “in one embodiment,” “in another embodiment,” “in yet another embodiment,” or “in other embodiments,” all of which may refer to one or more of the same or different embodiments according to this application.

[0028] First, let's introduce some application scenarios of this application. As mentioned in the background, document tags directly affect document retrieval efficiency and accuracy. Fixed document tags cannot meet the aforementioned document retrieval requirements. To address this issue, some methods introduce multi-dimensional tags, such as flat multi-tag classification and hierarchical classification. However, these predefined tags still struggle to fully cover user scenarios.

[0029] This application aims to address the shortcomings of existing document tag management methods, such as poor flexibility and inability to be personalized. Specifically, traditional methods struggle to dynamically and accurately generate document tags automatically based on users' actual usage habits and behavioral patterns, resulting in a rigid tag system, low retrieval efficiency, and poor retrieval accuracy, particularly when faced with massive amounts of heterogeneous documents and personalized needs. Unlike the solutions mentioned above, this application can determine a first tag for the target document corresponding to the interaction during the interaction based on user interaction data with the intelligent agent. This first tag can be used to improve retrieval efficiency and accuracy when searching for the target document.

[0030] Combination Figure 1 This application provides a method for determining document tags, the method may include the following steps: S1, obtain interactive behavior data associated with the target document generated by the user during interaction with the intelligent agent, wherein the target document is a document retrieved by the intelligent agent from the knowledge base.

[0031] For example, during user interactions with an agent, interactive behavior data is generated, and this data is associated with a target document. The target document is the document retrieved by the agent from a knowledge base during the user's interaction. A target document can be understood as one or more documents retrieved by the agent from a large pool of documents in response to a user's interaction request.

[0032] A knowledge base can be understood as a repository of information that provides answers to intelligent agents. An intelligent agent can be bound to multiple knowledge bases. Here, a knowledge base can be a knowledge base that stores documents, such as, but not limited to, after-sales knowledge bases, human resources knowledge bases, product knowledge bases, platform rules knowledge bases, etc.

[0033] It should be noted that the entity that obtains the interaction behavior data is not the intelligent agent that interacts with the user, nor is it a RAG (Retrieval-Augmented Generation) or a knowledge base. It can be a specific role, which can be an intelligent agent in form.

[0034] For example, when a user interacts with an agent, asking "What should I do if the API connection times out?", the agent performs semantic understanding of the user's interaction, searches and matches it in a knowledge base, and retrieves the target document. The target document could be something like "Solutions for API connection timeout" or "Timeout parameter adjustment schemes." Based on the content of the retrieved target document, the agent responds to the user. The user's subsequent actions related to the retrieved target document can be obtained as interaction behavior data.

[0035] It should be noted that the entity that obtains the interaction behavior data can receive the user interaction data forwarded by the intelligent agent that interacts with the user, and carry out the subsequent generation of the first tag based on this data. This does not require additional processing by the original module responsible for dialogue response or retrieval. While not occupying the computing resources of the original core process, it can also independently iterate the tag generation logic according to the user's needs without affecting the stable operation of the original interaction system.

[0036] S2, based on the interaction behavior data, obtain the first tag content corresponding to the target document.

[0037] For example, interaction behavior data can include data related to user actions such as clicking, saving, forwarding, dwell time, adopting and quoting content from the target document to generate replies, and liking replies generated based on the target document. Analyzing interaction behavior data can reveal the actual usage scenarios of the target document, and a large language model can be used to generate personalized first-tag content corresponding to the actual usage scenarios of the target document.

[0038] For example, continuing with the example in step S1, after the agent responds to the user based on the target document, the user repeatedly asks how to specifically set or adjust the API (Application Programming Interface) timeout-related parameters and repeatedly clicks on the details of the target document. The user's interaction data with the agent revolves around "API timeout" and doesn't deviate from other issues. Therefore, a first tag can be generated for the target document retrieved by the agent. This first tag can be used to characterize the degree of matching between the target document and the user's actual needs. For example, the first tag could be "Practical Guide to Solving API Timeout".

[0039] In this application, the first tag content is generated based on the user's interaction habits with the intelligent agent. By analyzing the user's interaction behavior data with the intelligent agent, the user's actual usage scenario is determined, thereby generating personalized tag content for the target document that matches the user's usage habits.

[0040] The content of the first tag can be stored in the target location, and then retrieved (e.g., read) from the target location when the target document is subsequently tagged.

[0041] S3, set a first tag for the target document in the knowledge base according to the content of the first tag, wherein the first tag is used to retrieve the target document.

[0042] For example, after obtaining the first tag content, tags are set for the target document in the knowledge base, and the first tag content is assigned to the target document. That is, the target document can obtain a first tag generated based on the user's interaction data with the agent. When the user interacts with the agent regarding API timeouts, the agent can quickly and accurately retrieve the target document with the first tag. Here, the user can be a user who has previously interacted with the agent, or another user.

[0043] This application combines dynamic user behavior analysis with the deep semantic understanding capabilities of large language models, achieving a shift in tag generation from a "document-centric" to a "user-centric" approach. Unlike traditional methods (such as clustering based on pre-defined keywords or static content), this application mines the potential value and application context of documents through actual user usage records, making tags more context-relevant and personalized. Furthermore, this application effectively reduces the cost of manual tagging and significantly improves document retrieval efficiency and user experience in the RAG system. Automatically generated personalized tags enable target documents to more accurately match user search intent and knowledge status, enhancing the automation level of the knowledge management platform. Its optimization mechanism based on real user feedback (such as positive confirmation in dialogues) also ensures continuous improvement in tag quality, making it widely applicable to scenarios requiring efficient document organization, such as enterprise knowledge bases and academic retrieval systems.

[0044] In some embodiments, obtaining the interaction behavior data associated with the target document generated by the user during interaction with the intelligent agent includes: The system obtains question data submitted by a user to the agent, wherein the target document is retrieved from the knowledge base based on the question data.

[0045] For example, during user interaction with the agent, the user may need to query content or ask one or more questions. These questions are stored as query data, such as in a business database. The query data stored in the business database can be retrieved by calling the corresponding data interface. For instance, the entity executing the interaction data can retrieve the user's questions from the business database by calling a backend API. It should be noted that the target document is retrieved from the knowledge base by the agent interacting with the user based on the query data.

[0046] For example, if a user's first question is, "I've already signed for the package, can I return it without giving a reason?", the agent performs semantic understanding of the first question, retrieves a matching document from the knowledge base, finds the target document "7-Day No-Reason Return Policy," and provides a response to the user based on the document's content. The user then asks a second question, "Can I still return the package after it's been unpacked?" The agent responds to this second question as well. After this, the user doesn't interact with the agent further regarding returns. In this case, the first and second questions posed by the user are stored as question data in the business knowledge base, which can be retrieved to analyze the user's information needs.

[0047] Of course, it is understandable that a user may only ask a question once during interaction with the intelligent agent, in which case the question data will only include the data corresponding to that single question; conversely, a user may ask multiple questions, in which case the question data will include the data corresponding to multiple questions. The scenario of a user asking two questions in the above embodiments is merely an example and is not intended to limit the scope of protection of the claims.

[0048] Additionally, it's important to note that users do not interact directly with RAG, but rather with the agent. Here, RAG's role is to retrieve relevant reference information from the RAG library after receiving a user's question. The agent then inputs the user's question and the reference information into the corresponding large model, and finally returns the generated content to the user.

[0049] In some embodiments, obtaining the interaction behavior data associated with the target document generated by the user during interaction with the intelligent agent includes: Obtain interactive data generated by the user's actions or information feedback on the target document during interaction with the intelligent agent.

[0050] For example, after retrieving a target document from a knowledge base, the agent can provide the document to the user. Of course, the agent can also provide text, images, and videos along with the target document. After receiving the target document from the agent, the user can interact with the document or provide feedback. For example, the user might click on the target document provided by the agent or browse the document. The user might also provide feedback on the target document provided by the agent, such as stating that the document should also include specific examples of the XX problem.

[0051] To reiterate the aforementioned example of no-reason return, when the intelligent agent provides the user with the target document "7-day no-reason return rules", the user clicks on the target document and browses the contents of "7-day no-reason return rules". After browsing, the user reports: "7-day no-reason return rules" can also record various specific scenario examples of 7-day no-reason return.

[0052] For example, an agent presents a user with a first target document, a second target document, and a third target document. The user clicks on and browses both the first and second target documents, spending more time browsing the first target document or copying and marking parts of it, while quickly browsing the second target document with less time. However, the user doesn't click on or browse the third target document, or may click on it and immediately close it. This indicates that the user has a higher level of engagement and attention with the first target document. This allows for more accurate analysis of user behavior data, facilitating the precise determination of user usage scenarios.

[0053] Interaction data generated by users interacting with the agent on target documents, including actions and feedback, can accurately record documents retrieved by the agent, as well as documents clicked, viewed, or ultimately adopted. This interaction data reflects the user's level of engagement and attention towards the target documents.

[0054] Interaction data generated by users interacting with the intelligent agent and performing operations on the target document or providing information feedback can be stored in the same location. The executing entity that obtains the interaction data can call the data interface that stores the interaction data generated by users interacting with the intelligent agent and performing operations on the target document or providing information feedback, in order to obtain the interaction data generated by users interacting with the intelligent agent and performing operations on the target document or providing information feedback.

[0055] In some embodiments, obtaining the interaction behavior data associated with the target document generated by the user during interaction with the intelligent agent includes: Obtain multi-turn dialogue data between the user and the agent, wherein the multi-turn dialogue data is associated with the target document.

[0056] For example, after the agent provides the user with a target document, the user may engage in further dialogue with the agent regarding the target document, such as asking in-depth follow-up questions about the content of the target document, making summary statements, or providing positive / negative feedback (e.g., "This document is useful" or "This is not what I wanted"). The dialogue may involve multiple rounds, and all rounds of dialogue are related to the target document.

[0057] Multi-turn dialogue data can be stored in a certain location. The multi-turn dialogue data can be obtained by calling the data interface corresponding to the data storage location. This multi-turn dialogue data is key to understanding the specific application scenario of the document.

[0058] In some embodiments, combined with Figure 2 The process of obtaining the first tag content corresponding to the target document based on the interaction behavior data includes the following steps: S21, preprocess the interaction behavior data to remove invalid conversations and noise data between the user and the intelligent agent.

[0059] For example, preprocessing the interaction behavior data may include cleaning the data to remove invalid conversations and noise data between the user and the intelligent agent.

[0060] For example, acquiring a large amount of user-agent interaction behavior data, the raw acquired interaction behavior data includes: User: Are you there? User: kjdflkasjdf; User: How many days of annual leave can I take at most? If a user asks "Are you there?" and immediately closes the conversation, this dialogue can be considered invalid and can be removed. Similarly, the user's input of "kjdflkasjdf" can be interpreted as random input, making this dialogue noisy data, and can also be removed. However, the question "How many days of annual leave can I take at most?" can be considered valid data. After preprocessing the interaction data, the first two invalid and noisy data points are removed, leaving only the valid interaction data "How many days of annual leave can I take at most?" to avoid invalid and noisy data affecting the efficiency and accuracy of subsequent processing of the interaction data.

[0061] S22, set the preprocessed interactive behavior data into an interactive sequence.

[0062] For example, preprocessed interaction data is associated with a session, piecing together a complete retrieval, browsing, and dialogue process into a complete interaction sequence. This can be understood as combining the data of the entire interaction process between the user and the system in chronological order, rather than inputting all the interaction data between the user and the agent in a disordered manner into the target model to be introduced below for analysis. This allows for a more accurate determination of the user's usage scenario, and consequently, a more accurate determination of the first tag content.

[0063] S23, determine the target input information based on the interaction sequence.

[0064] For example, after obtaining the interaction sequence, we can determine which content to input as target input information into the target model described in the next step. Here, target input information can be understood as the key content that is filtered, extracted, and integrated from the complete interaction sequence according to business objectives, specifically for analysis of the large model. Useful core content that needs to be analyzed can be selected from the interaction sequence, such as: a single sentence, multi-turn context, user requests, sensitive content, question intent, etc.

[0065] S24, input the target input information into the target model to obtain the first tag content generated by the target model.

[0066] For example, the target input information is input into the target model for analysis. The target model can be one of the following: a Large Language Model (LLM), a multimodal large model, a model capable of generating text, a model with more than 100 million adjustable parameters, a model that can output corresponding results based on input natural language instructions, a model that can convert input natural language instructions into feature sequences and infer output results based on feature sequences, or a model that can convert input information into feature sequences and infer subsequent features of the feature sequences. The target model can be required to generate concise and accurate label content based on the given target input information. When generating the first label content, a small number of high-quality positive and negative examples can be provided to guide the model in generating labels that meet the requirements. For example, a technical document may obtain different labels than its static content because it is used in different interaction scenarios. After analyzing the target input information, the target model can generate the first label content corresponding to the current target input information. For example, for interaction behavior data related to frequent container restarts, the target model, after analyzing the target input information, generates the first label content as "Practical Guide to Solving Pod Abnormal Restarts".

[0067] Here, the target document includes content related to the target document's static tags (second tags).

[0068] In some embodiments, combined with Figure 3 The step of determining the target input information based on the interaction sequence includes: S231, Perform behavioral analysis on the interaction sequence to determine usage scenario information.

[0069] For example, analysis modules (such as Natural Language Processing (NLP) techniques) can be used to analyze interaction sequences, perform topic extraction, intent recognition, and sentiment analysis on multi-turn dialogues, and abstract user scenarios. When a user queries "API connection timed out," analysis can determine the usage scenario information for each target document under different queries. Here, usage scenario information can be assigned as a identifier to the target document. Scenario information could include "highly frequent use in troubleshooting scenarios," "introductory concept explanations suitable for beginners," or "advanced performance optimization solutions," etc.

[0070] S232, determine the target input information based on the usage scenario information and the target document.

[0071] For example, the target input information can be determined based on the usage scenario information determined in step S231 and the target document retrieved by the agent. For instance, when a user queries "API connection timeout," the agent provides the user with the retrieved target document. Subsequent dialogue between the user and the agent often revolves around "checking network configuration" and "adjusting timeout parameters" related to the target agent. Therefore, based on the scenario information and content related to the target document, the target input information can be obtained from the interaction sequence through methods such as interception, filtering, merging, and extraction.

[0072] For example, the high-frequency dialogue patterns and user intent obtained from the analysis can be used as key background information input into the target model.

[0073] In some embodiments, the first tag of the target document can be updated based on target interaction data between the user and the target document with the first tag.

[0074] For example, the generated first tag content can be stored in a tag knowledge base for reverse reinforcement of the agent (e.g., a RAG system), forming a closed loop with the updating of the target document's tags. This establishes a tag feedback mechanism to continuously adjust the tag generation strategy, such as prioritizing specific scenario tags for highly interactive documents, thereby continuously improving the matching degree between tags and user habits. For instance, tracking the click-through rate, user satisfaction, and dialogue depth of target documents with the first tag. Feedback data such as the target document's click-through rate and user satisfaction can be used to evaluate the quality of the first tag and serve as new training data to continuously optimize the tag generation performance of the tag generation model. For example, if a target document with the first tag "High-Frequency Fault Solutions" is frequently reported by users as primarily providing fault code lookups during subsequent agent-user interactions, or is clicked multiple times when users search for fault codes, then the first tag of that target document can be updated to "Quick Lookup of Equipment Fault Codes."

[0075] Alternatively, the agent provides the user with a target document with a first tag, and the user's feedback on the target document with the first tag is negative, such as the user's feedback "This document is not what I want". In this case, the first tag of the target document can be deleted, and a new first tag can be generated for the target document after the first tag is deleted based on the user interaction data.

[0076] Alternatively, if the agent provides a target document with a first tag to the user, and this target document is also retrieved in other interaction scenarios between the user and the agent, and the user provides positive feedback or adopts the target document with the first tag, then a third tag can be generated for the target document with the first tag. The third tag is adapted to the current interaction scenario between the user and the agent, and the third tag is different from the first tag. In this case, the target document can simultaneously have a first tag, a second tag, and a third tag.

[0077] Furthermore, during the target document retrieval phase, in addition to traditional semantic matching, the agent also matches user queries with personalized tags (first tags). This ensures that even if the user's query terms differ from the wording in the document text, the document can still be retrieved as long as it is highly relevant to the tags.

[0078] For example, a target document might have a first tag and a second tag. The second tag is a static tag, determined based on the document's content (or semantics), and its content remains constant. For instance, the second tag might be "equipment manual" or "electrical equipment," while the first tag might be "high-frequency fault solutions" or "quick reference for equipment fault codes." When performing document retrieval, the agent can refer to both the first and second tag content. During the retrieval process, if the first tag and / or the second tag matches the user's query, target documents with both tags will be retrieved. Conversely, if the second tag content does not match the user's query, but the first tag content does, target documents with the first tag will also be retrieved.

[0079] By referring to the specific examples in the table below, the relationship between static tags, scenarios, and the first tag can be better understood.

[0080]

[0081] When the agent retrieves multiple target documents, the retrieved target documents can be arranged into a result list in order. When arranging the target documents, the matching degree between the first tag of the target document and the user's usage scenario can be used as an important feature to reorder the result list. The target documents that are more in line with the current user scenario are placed at the top, and other target documents are arranged in descending order of matching degree. This makes it easier for users to quickly find the target documents that meet their needs after receiving the target documents from the agent, thus improving the user experience.

[0082] In some embodiments, based on the interaction behavior data, obtaining the first tag content corresponding to the target document includes: Based on the interaction behavior data, the user type is determined.

[0083] For example, based on user interaction data with the intelligent agent, it can be determined whether the user is an expert or a novice. For instance, if the questions raised by the user in the interaction data are routine questions that can be understood by someone in the field, then the user is identified as a novice. Routine questions can be understood as questions that can be answered using basic operating manuals in the field. If the questions raised by the user in the interaction data, or the interaction with the target document provided by the intelligent agent, are more in-depth, then the user is identified as an expert. In-depth interaction can be understood as questions discussed in specialized academic papers.

[0084] Based on the user type, determine the first tag content corresponding to the user type for the target document.

[0085] For example, different first tag content can be determined for target documents for different user types. The first tag content corresponds to the user type, so that the target document corresponding to the first tag content can be retrieved more accurately when the user of the corresponding type interacts with the agent.

[0086] This application also provides a document tag determination system, combined with... Figure 4 The document tag determination system may include: a data acquisition module, a tag content determination module, and a setting module.

[0087] The data acquisition module is configured to obtain interactive behavior data associated with a target document generated by the user during interaction with the intelligent agent, wherein the target document is a document retrieved by the intelligent agent from a knowledge base.

[0088] For example, during user interactions with an agent, interactive behavior data is generated, and this data is associated with a target document. The target document is the document retrieved by the agent from a knowledge base during the user's interaction. A target document can be understood as one or more documents retrieved by the agent from a large pool of documents in response to a user's interaction request.

[0089] A knowledge base can be understood as a repository of information that provides answers to intelligent agents. An intelligent agent can be bound to multiple knowledge bases. Here, a knowledge base can be a knowledge base that stores documents, such as, but not limited to, after-sales knowledge bases, human resources knowledge bases, product instruction bases, platform rule knowledge bases, etc.

[0090] It should be noted that the entity that obtains the interaction behavior data is not the intelligent agent that interacts with the user, nor is it a RAG (Retrieval-Augmented Generation) or a knowledge base. It can be a specific role, which can be an intelligent agent in form.

[0091] For example, when a user interacts with an agent, asking "What should I do if the API connection times out?", the agent performs semantic understanding of the user's interaction, searches and matches it in a knowledge base, and retrieves the target document. The target document could be something like "Solutions for API connection timeout" or "Timeout parameter adjustment schemes." Based on the content of the retrieved target document, the agent responds to the user. The user's subsequent actions related to the retrieved target document can be obtained as interaction behavior data.

[0092] It should be noted that the data acquisition module can collect user interaction data forwarded by the intelligent agent interacting with the user, and carry out the subsequent generation of the first tag based on this data. This does not require additional processing by the original module responsible for dialogue response or retrieval. While not occupying the computing resources of the original core process, it can also independently iterate the tag generation logic according to the user's needs without affecting the stable operation of the original interaction system.

[0093] The tag content determination module is configured to obtain the first tag content corresponding to the target document based on the interaction behavior data.

[0094] For example, interaction behavior data can include data related to user actions such as clicking, saving, forwarding, dwell time, adopting and quoting content from the target document to generate replies, and liking replies generated based on the target document. Analyzing interaction behavior data can reveal the actual usage scenarios of the target document, and the tag content determination module of the large model can be used to generate personalized first tag content for the target document that corresponds to the actual usage scenarios of the target document.

[0095] For example, after the agent responds to a user based on a target document, the user repeatedly asks how to set or adjust API (Application Programming Interface) timeout parameters and repeatedly opens the details of the target document. The user's interaction with the agent revolves around "API timeout" and doesn't deviate from other issues. Therefore, a first tag can be generated for the target document retrieved by the agent. This first tag can characterize the degree to which the target document matches the user's actual needs. For example, the first tag could be "Practical Guide to Solving API Timeouts".

[0096] In this application, the first tag content is generated based on the user's interaction habits with the intelligent agent. By analyzing the user's interaction behavior data with the intelligent agent, the user's actual usage scenario is determined, thereby generating personalized tag content for the target document that matches the user's usage habits.

[0097] The content of the first tag can be stored in the target location, and then retrieved (e.g., read) from the target location when the target document is subsequently tagged.

[0098] The setting module is configured to set a first tag for the target document in the knowledge base based on the content of the first tag, wherein the first tag is used to retrieve the target document.

[0099] For example, after obtaining the first tag content, the setting module can set tags for the target document in the knowledge base, assigning the first tag content to the target document. That is, the target document can obtain a first tag generated based on the user's interaction data with the agent. When the user interacts with the agent regarding API timeouts, the agent can quickly and accurately retrieve the target document with the first tag. Here, the user can be a user who has previously interacted with the agent, or another user.

[0100] This application provides an embodiment of an electronic device, combined with... Figure 5 The electronic device may include a memory and a processor. The memory stores an executable program, and the processor is configured to execute: obtaining interactive behavior data associated with a target document generated by a user during interaction with an intelligent agent, wherein the target document is a document retrieved by the intelligent agent from a knowledge base; obtaining a first tag content corresponding to the target document based on the interactive behavior data; and setting a first tag for the target document in the knowledge base according to the first tag content, wherein the first tag is used for retrieving the target document.

[0101] The processor is also used to perform steps of the other methods described in the above embodiments.

[0102] Furthermore, one embodiment of this application also provides a storage medium storing a computer program that, when executed by a processor, implements the document tag determination method described above. It should be understood that the various solutions in this embodiment have the corresponding technical effects in the above-described method embodiments, and will not be repeated here.

[0103] Furthermore, embodiments of this application also provide a computer program product, which is tangibly stored on a computer-readable medium and includes computer-readable instructions that, when executed, cause at least one processor to perform a document tag determination method such as the one described in the embodiments above.

[0104] It should be noted that the computer storage medium in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. Computer-readable media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access storage media (RAM), read-only storage media (ROM), erasable programmable read-only storage media (EPROM or flash memory), optical fibers, portable compact disk read-only storage media (CD-ROM), optical storage media, magnetic storage media, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program configured for use by or in conjunction with an instruction execution system, system, or device. Program code contained on a computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, antenna, optical fiber, RF, etc., or any suitable combination thereof. Furthermore, those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which are executable by the processor of the computer or other programmable data processing device, produce instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1A system specifying functions within one or more boxes. These computer program instructions may also be stored in a computer-readable storage medium capable of directing a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction system implemented in a process. Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0105] The foregoing has described in detail several embodiments of this application, but this application is not limited to these specific embodiments. Those skilled in the art can make various variations and modifications based on the concept of this application, and all such variations and modifications should fall within the scope of protection claimed in this application.

Claims

1. A method for determining document tags, the method comprising: Obtain interactive behavior data associated with a target document generated by the user during interaction with the intelligent agent, wherein the target document is a document retrieved by the intelligent agent from a knowledge base; Based on the interaction behavior data, obtain the first tag content corresponding to the target document; A first tag is set in the knowledge base based on the content of the first tag, wherein the first tag is used to retrieve the target document.

2. The method according to claim 1, wherein obtaining the interaction behavior data associated with the target document generated by the user during interaction with the intelligent agent includes: The system obtains question data submitted by a user to the agent, wherein the target document is retrieved from the knowledge base based on the question data.

3. The method according to claim 1, wherein obtaining the interaction behavior data associated with the target document generated by the user during interaction with the intelligent agent includes: Obtain interactive data generated by the user's actions or information feedback on the target document during interaction with the intelligent agent.

4. The method according to claim 1, wherein obtaining the interaction behavior data associated with the target document generated by the user during interaction with the intelligent agent includes: Obtain multi-turn dialogue data between the user and the agent, wherein the multi-turn dialogue data is associated with the target document.

5. The method according to claim 1, wherein obtaining the first tag content corresponding to the target document based on the interaction behavior data includes: The interaction data is preprocessed to remove invalid conversations and noise data between the user and the agent; The preprocessed interactive behavior data is set as an interactive sequence; The target input information is determined based on the interaction sequence; The target input information is input into the target model to obtain the first tag content generated by the target model.

6. The method according to claim 5, wherein determining the target input information based on the interaction sequence comprises: Behavioral analysis is performed on the interaction sequence to determine the usage scenario information; The target input information is determined based on the usage scenario information and the target document.

7. The method of claim 1, wherein, The first tag of the target document can be updated based on the target interaction behavior data of the user with the target document bearing the first tag.

8. The method according to claim 1, wherein based on the interaction behavior data, obtaining the first tag content corresponding to the target document, includes: Based on the interaction behavior data, the user type is determined; Based on the user type, determine the first tag content corresponding to the user type for the target document.

9. A document tag determination system, comprising: The data acquisition module is configured to acquire interactive behavior data associated with a target document generated by the user during interaction with the intelligent agent, wherein the target document is a document retrieved by the intelligent agent from a knowledge base; The tag content determination module is configured to obtain the first tag content corresponding to the target document based on the interaction behavior data. A setting module is configured to set a first tag for the target document in the knowledge base based on the content of the first tag, wherein the first tag is used to retrieve the target document.

10. An electronic device comprising a memory and a processor, wherein the memory stores an executable program, and the processor is configured to execute: Obtain interactive behavior data associated with a target document generated by the user during interaction with the intelligent agent, wherein the target document is a document retrieved by the intelligent agent from a knowledge base; Based on the interaction behavior data, obtain the first tag content corresponding to the target document; A first tag is set in the knowledge base based on the content of the first tag, wherein the first tag is used to retrieve the target document.