A method for generating and labeling a topic based on an agent, a related device, and a program product

CN122820197APending Publication Date: 2026-09-25BAIDU COM TIMES TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611007041.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

但海量内容虽极大地丰富了用户的选择空间,却也同时加剧了信息的碎片化与无序性,导致用户在主动获取或被动接收信息时难以快速定位和筛选出真正符合自身兴趣或需求的内容

Benefits of technology

[0013]本公开实施例提供的基于智能体的话题标签生成方法、装置、电子设备、计算机可读存储介质及计算机程序产品,首先,从历史内容集合中抽取出第一历史内容子集合和至少两个第二历史内容子集合;然后,针对第二历史内容子集合,利用话题生成智能体处理第二历史内容子集合中包括的历史内容,生成与第二历史内容子集合各自对应的待评价话题标签集合;最后,基于待评价话题标签集合对于第一历史内容子集合的第一内容覆盖率,从待评价话题标签集合中确定出目标话题标签集合。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820197A_ABST
    Figure CN122820197A_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for generating and labeling topics based on an agent, a related device and a program product, and relates to the technical fields of artificial intelligence such as deep learning, natural language processing, information classification and agent. A specific embodiment of the method for generating topic labels based on an agent comprises: extracting a first historical content sub-set and at least two second historical content sub-sets from a historical content set; for the second historical content sub-set, processing the historical content included in the second historical content sub-set by using a topic generation agent to generate a set of to-be-evaluated topic labels corresponding to each of the second historical content sub-set; and determining a target topic label set from the set of to-be-evaluated topic labels based on a first content coverage rate of the first historical content sub-set for the set of to-be-evaluated topic labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, specifically to the fields of artificial intelligence technology such as deep learning, natural language processing, information classification, and intelligent agents, and particularly to methods for generating and labeling topics based on intelligent agents, devices for generating and labeling topics based on intelligent agents, electronic devices, computer-readable storage media, and computer program products. Background Technology

[0002] With the rapid development of internet technology, the amount of online content has grown exponentially. Users are exposed to an ever-increasing volume of information daily, in increasingly diverse formats, including news, social media updates, audio and video programs, and blogs. While this massive amount of content greatly enriches users' choices, it also exacerbates the fragmentation and disorder of information, making it difficult for users to quickly locate and filter content that truly matches their interests or needs when actively acquiring or passively receiving information.

[0003] Against this backdrop, topic tagging technology has emerged to improve the efficiency of user information acquisition and reduce the cost of information filtering. This technology semantically summarizes and classifies information resources (such as social content, articles, and videos), assigning them structured "topic" tags. This helps users quickly determine the theme and scope of content without having to browse all of it, assisting them in deciding whether to seek out the full content.

[0004] Therefore, how to label topics more effectively is a matter of concern and an urgent need. Summary of the Invention

[0005] This disclosure provides a method for generating and labeling topics based on intelligent agents, an apparatus for generating and labeling topics based on intelligent agents, an electronic device, a computer-readable storage medium, and a computer program product.

[0006] In a first aspect, embodiments of this disclosure propose a topic tag generation method based on an intelligent agent, comprising: extracting a first historical content subset and at least two second historical content subsets from a historical content set; for the second historical content subsets, using a topic generation intelligent agent to process the historical content included in the second historical content subsets, generating a set of topic tags to be evaluated corresponding to each of the second historical content subsets; and determining a target topic tag set from the set of topic tags to be evaluated based on a first content coverage rate of the set of topic tags to be evaluated to the first historical content subsets.

[0007] Secondly, this disclosure proposes a method for labeling topics, including: obtaining content to be labeled and a set of target topic tags, wherein the set of target topic tags is obtained based on the agent-based topic tag generation method described in the first aspect above; determining matching topic tags from the set of target topic tags that match the content to be labeled; and labeling the content to be labeled using the matching topic tags.

[0008] Thirdly, embodiments of this disclosure propose an agent-based topic tag generation apparatus, comprising: a subset extraction unit configured to extract a first historical content subset and at least two second historical content subsets from a historical content set; an agent generation invocation unit configured to process the historical content included in the second historical content subset using a topic generation agent, generating a set of topic tags to be evaluated corresponding to each of the second historical content subsets; and a target set determination unit configured to determine a target topic tag set from the set of topic tags to be evaluated based on a first content coverage rate of the set of topic tags to be evaluated relative to the first historical content subset.

[0009] Fourthly, embodiments of this disclosure propose an apparatus for labeling topics, comprising: a content to be labeled and a topic tag acquisition unit, configured to acquire the content to be labeled and a target topic tag set, wherein the target topic tag set is obtained based on the agent-based topic tag generation apparatus described in the third aspect above; a topic tag matching unit, configured to determine matching topic tags from the target topic tag set that match the content to be labeled; and a topic labeling unit, configured to label the content to be labeled using the matching topic tags.

[0010] Fifthly, embodiments of this disclosure provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the agent-based topic tag generation method as described in any implementation of the first aspect, and / or the topic tagging method as described in any implementation of the second aspect.

[0011] In a sixth aspect, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions that enable a computer, when executed, to implement the agent-based topic tag generation method as described in any implementation of the first aspect, and / or the topic tagging method as described in any implementation of the second aspect.

[0012] In a seventh aspect, embodiments of this disclosure provide a computer program product including a computer program, which, when executed by a processor, can implement the agent-based topic tag generation method as described in any implementation of the first aspect, and / or the topic tagging method as described in any implementation of the second aspect.

[0013] The topic tag generation method, apparatus, electronic device, computer-readable storage medium, and computer program product based on intelligent agents provided in this disclosure firstly extract a first subset of historical content and at least two second subsets of historical content from a historical content set; then, for each of the second subsets of historical content, a topic generation intelligent agent processes the historical content included in the second subset to generate a set of topic tags to be evaluated corresponding to each of the second subsets of historical content; finally, based on a first content coverage rate of the set of topic tags to be evaluated to the first subset of historical content, a target set of topic tags is determined from the set of topic tags to be evaluated.

[0014] This disclosure not only enables the pre-generation of a set of topic tags for subsequent annotation actions based on real historical content, thus effectively managing the topic tags actually used and avoiding unstable annotation quality caused by the divergence of temporarily generated tags, but also ensures the quality of topic tags in the set by leveraging an internal feedback mechanism of generation-evaluation from the same historical content during the generation and determination of topic tags.

[0015] Accordingly, the agent-based topic tag generation method, apparatus, electronic device, computer-readable storage medium, and computer program product provided in this disclosure first obtain the content to be tagged and a target topic tag set, wherein the target topic tag set is obtained based on the above-described agent-based topic tag generation method; then, matching topic tags that match the content to be tagged are determined from the target topic tag set; and finally, the matching topic tags are used to tag the content to be tagged.

[0016] Accordingly, this disclosure can utilize the aforementioned generated set of topic tags for topic labeling, thereby achieving higher quality topic labeling.

[0017] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0018] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1This is an exemplary system architecture to which this disclosure can be applied; Figure 2 A flowchart illustrating an agent-based topic tag generation process provided in this embodiment of the disclosure; Figure 3 A flowchart illustrating a process for determining a set of target topic tags, provided in this embodiment of the disclosure; Figure 4 A flowchart illustrating a topic labeling process provided in this embodiment of the disclosure; Figure 5 A flowchart illustrating the agent-based topic tag generation process in a specific application scenario provided by an embodiment of this disclosure; Figure 6 A structural block diagram of an agent-based topic tag generation device provided in this disclosure embodiment; Figure 7 A structural block diagram of a device for labeling topics provided in an embodiment of this disclosure; Figure 8 This is a schematic diagram of the structure of an electronic device that is suitable for performing a method for generating and labeling topics based on intelligent agents, as provided in an embodiment of this disclosure. Detailed Implementation

[0019] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding; these should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0020] Furthermore, the acquisition, storage, use, processing, transportation, provision, and disclosure of any type of information involved in the technical solutions disclosed herein, such as user personal information (e.g., historical content that may be user content published by users), comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0021] Figure 1 An exemplary system architecture 100 is shown, in which embodiments of the agent-based topic tag generation and topic annotation methods, apparatuses, electronic devices and computer-readable storage media of this disclosure can be applied.

[0022] like Figure 1As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0023] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed. These applications include content publishing applications, content summarizing applications, and instant messaging applications.

[0024] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.

[0025] Server 105 can provide various services through its built-in applications. Taking a content publishing application with "topic tagging" functionality as an example, when running this application, server 105 can achieve the following: First, it can extract a first subset of historical content and at least two second subsets of historical content from the historical content set. Then, server 105 can use a topic generation agent to process the historical content included in the second subsets of historical content, generating a set of topic tags to be evaluated corresponding to each subset. Next, based on the first content coverage of the set of topic tags to be evaluated relative to the first subset of historical content, server 105 determines a target set of topic tags from the set of topic tags to be evaluated. Accordingly, when presenting or providing content to terminal devices 101, 102, and 103, server 105 can first treat this content as content to be tagged. Then, server 105 obtains this content to be tagged and the aforementioned target set of topic tags. Then, server 105 determines matching topic tags from the target set of topic tags that match the content to be tagged. Finally, server 105 uses the matching topic tags to tag the content to be tagged.

[0026] It should be noted that the historical content collection can be obtained from terminal devices 101, 102, and 103 via network 104, or it can be pre-stored locally on server 105 through various means. Therefore, when server 105 detects that this data is already stored locally (e.g., when starting to process previously retained topic tag generation tasks), it can choose to retrieve this data directly from local storage. In this case, the exemplary system architecture 100 may also exclude terminal devices 101, 102, and 103 and network 104.

[0027] Since storing historical content sets and utilizing and calling intelligent agents (e.g., topic generation agents) often require significant computing resources and capabilities, the methods for generating and labeling topics based on intelligent agents provided in the subsequent embodiments of this disclosure are generally executed by a server 105 with strong computing power and abundant computing resources. Correspondingly, the apparatus for generating and labeling topics based on intelligent agents is also generally located in the server 105. However, it should also be noted that when terminal devices 101, 102, and 103 also possess sufficient computing power and resources, they can also complete the aforementioned calculations performed by the server 105 through content publishing applications installed on them, thereby outputting the same results as the server 105. Especially when multiple terminal devices with different computing capabilities exist simultaneously, but the content publishing application determines that the terminal device it is on has strong computing power and a large amount of remaining computing resources, the terminal device can be allowed to perform the above-mentioned calculations, thereby appropriately reducing the computing pressure on server 105. Correspondingly, the device for generating and labeling topics based on intelligent agents can also be set in terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also exclude server 105 and network 104. Alternatively, in different scenarios, the device for generating and labeling topics based on intelligent agents can be deployed on both sides of server 105 and terminal devices 101, 102, and 103, respectively. For example, the device for generating topics based on intelligent agents is deployed in server 105, while the device for labeling topics is deployed in terminal devices 101, 102, and 103.

[0028] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0029] Next, we will explain the process of generating topic tags based on intelligent agents. For ease of understanding, we will also refer to... Figure 2 Let's have a discussion. Figure 2 A flowchart of an agent-based topic tag generation process provided for embodiments of this disclosure includes process 200.

[0030] Process 200 specifically includes the following steps: Step 201: Extract a first subset of historical content and at least two second subsets of historical content from the historical content set; In embodiments of this disclosure, this step is intended to be performed by the agent executing the agent-based topic tag generation method (e.g., Figure 1The server 105 shown obtains the historical content set and then extracts a first historical content subset and at least two second historical content subsets from it.

[0031] Typically, the data and content sources of the historical content collection can be the same as those of the content to be annotated, which will be discussed later. For example, they can correspond to the same data source and content source. For instance, if the data to be annotated is social content (e.g., "posts") published by users in "Baidu Tieba", the historical content could be "historical posts" or "historical comments" published by users in the past.

[0032] In this step, the executing entity can extract from the historical content set a second subset of historical content used to generate "topic tags", a subset of historical content used as material for generating "topic tags", and a subset of historical content used to evaluate the quality of the generated topic tags.

[0033] Accordingly, to avoid instability in the generation results when using the second subset of historical content to generate topic tags, at least two subsets of historical content can be extracted. This allows for multiple rounds of sampling and voting to reduce the instability of a single generation result, ensuring the robustness and comprehensiveness of the generated topic tags. For example, the implementing entity can perform multiple random samplings (to reduce the cost of extracting the second subset of historical content while ensuring that each piece of historical content is extracted with equal and even chance) to extract at least two subsets of second historical content.

[0034] In some embodiments, since the subsequent use of the first historical content subset is to evaluate the generated topic tags (e.g., to assess whether their quality meets requirements), in order to further improve the accuracy of the evaluation, in addition to extracting the first historical content subset in the same way as extracting the second historical content subset, the implementing entity may also choose to extract the first historical content subset according to the content type distribution of the historical content in the historical content set as a reference, so that the first historical content subset is comparable to the historical content set in terms of the content types of the historical content it includes. This allows the extracted first historical content subset to better "represent" and "characterize" the historical content set, thereby reducing evaluation bias caused by observation angles and perspectives during subsequent evaluations.

[0035] In this scenario, when historical content is provided and stored locally by the executing entity, its data format can be multi-field data. These fields may include the identity identifier corresponding to the historical content, text content, and so on. Accordingly, after obtaining the historical data, the executing entity can classify this historical content based on the analysis of the multi-field data to determine the content type distribution corresponding to the historical content set (for ease of description, the content type distribution here can be described as the first content type distribution).

[0036] In different scenarios, this content type distribution can also be described as a "data profile" of historical content, reflecting the content types it includes and their distribution (or proportion). For example, the proportion of "historical content belonging to popular information", the proportion of "historical content belonging to long texts", and the proportion of "historical content belonging to positive emotions".

[0037] Then, the implementing entity can determine the permissible proportion range of each content type based on the first content type distribution, and construct at least one permissible second content type distribution. Accordingly, the content types included in the second content type distribution are comparable to those in the first content type distribution (or, the proportion of the comparable portion is greater than or equal to a proportion threshold), and their proportions in the second content type distribution all conform to the aforementioned proportion range. In other words, the similarity between the second content type distribution and the first content type distribution is greater than or equal to a similarity threshold. This similarity threshold can typically be determined based on the criteria that the two distributions are similar and their composition can be considered "identical" or "highly similar."

[0038] Therefore, this approach allows the first subset of historical content extracted by the implementing entity to have a better "global perspective" on the historical content set, enabling the evaluation and selection process of the subsequent topic tag set to be completed with higher quality by utilizing this global perspective of the historical content set.

[0039] In some embodiments, in order to safeguard the value of historical content, the implementing entity can periodically utilize a collecting agent to collect data from the aforementioned content sources and data sources, so that historical content can be dynamically and continuously updated based on online interactions and transformed into "topic tags" accordingly.

[0040] An intelligent agent is an autonomous system capable of perceiving its environment, making decisions based on that perception, and executing actions to achieve specific goals. An intelligent agent can be a physical entity (such as a robot), a software program (such as intelligent customer service or automated transaction systems), or a combination of both. An intelligent agent typically consists of a perception module, a decision-making module, and an execution module. The perception module is responsible for collecting environmental information; the decision-making module formulates action plans based on the environmental state and built-in strategies; and the execution module translates the decisions into concrete actions. Thus, based on these modules, the intelligent agent can autonomously adapt to dynamic environments, optimize its goal-achieving strategies, and exhibit intelligent behavior in complex tasks.

[0041] Accordingly, the aforementioned collection agent can be configured and trained based on such an agent architecture. It can periodically send corresponding data acquisition requests to the content sources and data sources, and then receive and store the historical content provided by the content sources and data sources. Through this collection agent, the executing entity can conveniently and efficiently complete the acquisition and actual collection of historical content sets by invoking it. This greatly improves the execution efficiency of the executing entity while reducing the cost of collecting this historical content.

[0042] Step 202: For the second subset of historical content, use the topic generation agent to process the historical content included in the second subset of historical content and generate a set of topic tags to be evaluated corresponding to each subset of the second subset of historical content; In the embodiments of this disclosure, after the executing entity extracts a first subset of historical content and at least two second subsets of historical content based on step 201, the executing entity may invoke a topic agent in this step and utilize a topic generation agent to process the historical content included in the second subsets of historical content, generating a set of topic tags corresponding to each of the second subsets of historical content. Since this set of topic tags needs to be evaluated subsequently, for example, through the evaluation process using the first subset of historical content, which will be discussed later, it can be described here as a "set of topic tags to be evaluated" for ease of understanding.

[0043] Similar to the above, the topic generation agent can be an agent with semantic understanding and inductive abilities. For example, it can be an agent built on a Large Language Model (LLM) and deployed with an LLM. An LLM is an artificial intelligence model designed to understand and generate human language. Based on its understanding, the LLM can perform corresponding processing operations to obtain the corresponding results. For example, when a second subset of historical content is obtained, the LLM can, after understanding the instructions (e.g., summarizing the topics involved in the second subset of historical content and generating corresponding topic tags), generate a set of topic tags to be evaluated corresponding to that second subset of historical content.

[0044] For example, the topic generation agent can process the historical content included in the second historical content subset and generate a set of topic tags corresponding to the second historical content subset, such as [topic tag A, topic tag B, topic tag C, topic tag D].

[0045] LLMs can be trained on large amounts of text data and perform a wide range of tasks, including text summarization, translation, sentiment analysis, and more. A key characteristic of LLMs is their large scale; they typically include a large number of parameters to help them learn complex patterns in language data. These models are often based on deep learning architectures, such as transformers, which contributes to their superior performance across various natural language tasks.

[0046] In some embodiments, to improve generation efficiency, the executing agent may utilize a topic-generating agent to process the historical content included in each of the second subsets of historical content in parallel.

[0047] In some embodiments, to avoid the generated set of topic tags to be evaluated being too scattered, and to avoid generating too few topic tags that would make it difficult to cover historical content on a broad scale, the executing entity may also choose to constrain the number of topic tags that can be included in each set of topic tags to be evaluated when calling the topic generation agent, so as to ensure that the topic tags in the set of topic tags are not too scattered or too sparse (i.e., unable to cover historical content on a broad scale).

[0048] Step 203: Based on the first content coverage rate of the set of topic tags to be evaluated to the first historical content subset, determine the target topic tag set from the set of topic tags to be evaluated.

[0049] In the embodiments of this disclosure, after the executing entity generates a set of topic tags to be evaluated corresponding to each of the second historical content subsets based on the above step 202, in this step, for each set of topic tags to be evaluated, the historical content in the first historical content subset that the topic tags included in it can cover is determined.

[0050] Then, for a specific set of topic tags to be evaluated, the executing entity determines its first content coverage rate for the first historical content subset based on the proportion of historical content in all the first historical content subsets that it can cover.

[0051] Then, the executing entity can select the set of topic tags with the highest coverage of the first content to be evaluated as the target set of topic tags to be used. Thus, through multiple rounds of sampling and voting, the instability of the results generated by the single topic generation agent can be reduced, ensuring the robustness and comprehensiveness of the topic tags.

[0052] In some embodiments, the executing entity may also collect a set of topic tags to be evaluated whose first content coverage is greater than or equal to the coverage threshold by comparing the first content coverage with the coverage threshold. Then, the executing entity combines them and removes duplicate topic tags (e.g., identical, semantically similar to or greater than or equal to the similarity threshold) and uses the combined result as the target topic tag set, so that the target topic tag set can cover a larger subset of the first historical content.

[0053] In some embodiments, because the executing entity evaluates the set of topic tags to be evaluated based on a first content coverage rate (e.g., as the first content coverage rate increases, it indicates that the quality of the set of topic tags to be evaluated is higher, and it can cover more of the historical content set), the executing entity can extract the first historical content subset and the second historical content subset using the same quantity standard. That is, the first historical content subset and the second historical content subset extracted by the executing entity include the same amount of historical content.

[0054] This allows for the use of the same quantitative dimensional standards for extraction and evaluation, avoiding situations such as over-evaluation or under-evaluation caused by quantitative dimensional bias, and improving evaluation accuracy.

[0055] It should be understood that, in different embodiments and scenarios, based on different needs, after the target topic tag set is generated, it can be stored and used in a way that replaces the currently valid target topic tag set (that is, the target topic tag set generated in this round is directly and uniquely used as the currently valid target topic tag set), or it can be used as the final "target topic tag set" after merging and deduplicating with the currently valid target topic tag set.

[0056] The topic tag generation method based on intelligent agents provided in this disclosure first extracts a first subset of historical content and at least two second subsets of historical content from a historical content set. Then, for each of the second subsets of historical content, a topic generation intelligent agent processes the historical content included in that subset to generate a set of topic tags to be evaluated corresponding to each of the second subsets. Finally, based on the first content coverage rate of the set of topic tags to be evaluated relative to the first subset of historical content, a target set of topic tags is determined from the set of topic tags to be evaluated. Therefore, this method not only pre-generates a set of topic tags for subsequent annotation actions based on real historical content, effectively managing the topic tags actually used and avoiding instability in annotation quality caused by the dispersion of temporarily generated tags, but also ensures the quality of the topic tags in the set through an internal feedback mechanism of generation-evaluation from the same historical content during the generation and determination process.

[0057] In some embodiments, in order to ensure the granularity of topic tags and avoid lumping a large amount of content under a single higher-level tag, thereby improving the actual quality of topic tags, the executing entity may, in addition to selecting and determining the target topic tag set based on the first content coverage, also detect whether there are any "excessive topic tags" in the topic tag set to be evaluated that are related to or cover too much historical content in the historical content set. If so, the executing entity may choose to refine it by adding more granular subordinate or lower-level topic tags to improve the quality of the final target topic tag set.

[0058] To make it easier to understand, you can combine Figure 3 This will be explained together. Figure 3 A flowchart illustrating a process for determining a set of target topic tags, as provided in this embodiment of the disclosure, includes process 300. For example, process 300 can be an alternative or alternative implementation of step 203 described above.

[0059] Process 300 specifically includes the following steps: Step 301: Based on the first content coverage rate of the set of topic tags to be evaluated to the first historical content subset, extract the set of intermediate topic tags whose first content coverage rate is greater than or equal to the coverage rate threshold; Specifically, this step is actually similar to the discussion above. The executing entity can determine the set of topic tags to be evaluated with a first content coverage rate greater than or equal to the coverage rate threshold as the intermediate topic tag set (instead of directly determining it as the target topic tag set mentioned above) in the manner described in process 200.

[0060] Step 302: Determine the associated historical content in the historical content set and the topic tags in the intermediate topic tag set; Specifically, based on step 301 above, the executing entity can determine the associated historical content of each topic tag in the intermediate topic tag set.

[0061] Then, after determining the associated historical content, the implementing entity can count the number of associated historical content items associated with each topic tag.

[0062] As discussed above, if the number of associated historical content for each topic tag is less than or equal to the threshold (which can usually be set based on different scenarios and hierarchical strategies, according to the standards that require further hierarchical, splitting, and refinement of topic tags), that is, if there are no excessive topic tags as mentioned above, then the executing entity can respond to this by choosing to execute step 303, and as described above, directly use the intermediate topic tag set as the final target topic tag set.

[0063] If there are excessive topic tags, that is, if the number of associated historical content in the intermediate topic tag set exceeds the threshold, the executing entity can respond by selecting to execute step 304.

[0064] Step 303: Use the intermediate topic tag set as the target topic tag set.

[0065] Step 304: Use a topic generation agent to generate sub-topic tags associated with the excessive number of topic tags; Specifically, in this step, the executing entity can choose to continue calling and utilizing the aforementioned topic-generating agent to generate subordinate topic tags associated with the excessive topic tags. For example, the executing entity can use historical content associated with the excessive topic tags from the historical content set as material, provide it to the topic-generating agent, and use "excessive topic tags" as semantic constraints (for example, the generated subordinate topic tags should be subordinate and refined to the excessive topic tags in terms of semantics and textual form) to provide subordinate topic tags associated with the excessive topic tags.

[0066] In some embodiments, if there are excessive topic tags in the intermediate topic tag set during this step, considering that the historical content set may include a large amount of related historical content, in order to improve processing efficiency and avoid information overload, the executing entity may further select to determine a third historical content subset from the historical content associated with the excessive topic tags.

[0067] For example, the executing entity can select a center point using clustering algorithms such as K-Medoids, and determine the third historical content subset based on that center point.

[0068] Then, the executing entity can further choose to use the topic-generating agent to process the historical content included in the third historical content subset and generate subordinate topic tags associated with the excess topic tags.

[0069] Subsequently, in the case of "excessive topic tags", the executing entity can further obtain the target topic tag set based on the intermediate topic tag set and the lower-level topic tags by executing step 305.

[0070] Step 305: By combining the intermediate topic tag set and the subordinate topic tags, the target topic tag set is obtained.

[0071] In some embodiments, if the target topic tag set includes sub-topic tags, the executing entity can store the sub-topic tags in the target topic tag set in association with their associated topic tags, or in other words, their parent topic tags, so that when using these topic tags with sub-branches, the associated topic tags at each level can be provided continuously and hierarchically in a tree structure.

[0072] For example, if the top-level topic tag is "food and cooking," and its next level is associated with "baking tutorials," "home-style recipes," "restaurant reviews," and "ingredient selection," then when the subsequent subject performs labeling and matching, or in other words, when using the target topic tag set, if a piece of content to be labeled can be labeled as "food and cooking," it can also be further and simultaneously labeled as one of "baking tutorials," "home-style recipes," "restaurant reviews," or "ingredient selection."

[0073] It should be understood that in the above process, the "leveling" of the topic tag can be at least two levels or more. For example, in the above example, if the newly generated "baking tutorial" is still an "overloaded topic tag", it can be further used by the topic generation agent to generate lower-level topic tags (e.g., cake making, bread making, cookies and pastries, baking tools), etc.

[0074] That is, after generating new sub-topic tags, the executing entity can further repeat and loop the above process until the currently newly generated level no longer includes excess topic tags, or the currently newly generated level has reached the level constraint limit (for example, reaching the "leaf node" in the above tree structure).

[0075] In some embodiments, during the process of generating subordinate topic tags associated with excess topic tags using a topic generation agent—that is, in response to the existence of excess topic tags in the intermediate topic tag set—if there are at least two excess topic tags, the executing entity can also respond by actually selecting to use the topic generation agent to generate subordinate topic tags associated with each excess topic tag in parallel. This improves the efficiency of subordinate topic tag generation by parallelizing the generation of each topic at different levels.

[0076] In some embodiments, during the process of generating subordinate topic tags associated with the super-quantity topic tag using a topic generation agent, since the topic generation agent is actually based on semantic understanding, analysis capabilities (e.g., the topic generation agent is built based on LLM) and generation capabilities, it can also be required to provide usage constraint information between subordinate topic tags belonging to the same super-quantity topic tag, or in other words, to distinguish boundaries. For example, the topic generation agent can be required to provide constraint information between its generated subordinate topic tags F1, F2, and F3 belonging to the same super-quantity topic tag F. For example, it can be stated that subordinate topic tag F1 should be used in case X1, subordinate topic tag F2 should be used in case X2, and subordinate topic tag F3 should be used in case X3, and so on.

[0077] In response to the existence of an excess of topic tags in the intermediate topic tag set, the executing agent generates subordinate topic tags associated with the excess topic tags using a topic generation agent. Furthermore, in response to the presence of an excess of topic tags in the intermediate topic tag set, the executing agent further selects and utilizes the topic generation agent to generate subordinate topic tags associated with the excess topic tags, as well as the usage constraint information between subordinate topic tags belonging to the same excess topic tag. This allows for the subsequent use of this constraint information to more accurately distinguish the criteria between subordinate topic tags, enabling the utilization of subordinate topic tags at a lower cost and higher efficiency.

[0078] Based on any of the above embodiments, after generating the target topic tag set, in order to ensure the quality of the target topic tag set, the executing entity may also choose to use an evaluation agent to process the target topic tag set in order to evaluate the quality of the target topic tag set.

[0079] As discussed above, an evaluation agent is an agent trained to process and evaluate a set of target topic tags based on a predetermined evaluation dimension, and to generate evaluation metrics corresponding to that dimension. Evaluation metrics typically include at least one of the following: the second content coverage of the target topic tag set relative to the historical content set; the uniformity of the distribution of the third content type in the third historical content subset; and the semantic similarity between lower-level topic tags.

[0080] This will establish a complete quality closed loop from generation and screening to evaluation, thereby improving the intelligence level of the label generation process.

[0081] In some optional implementations of this embodiment, corresponding threshold values ​​can be pre-determined for each evaluation metric (e.g., coverage threshold corresponding to the second content coverage, uniformity threshold corresponding to distribution uniformity, similarity threshold for semantic similarity, etc.) so that the quality of the target topic tag set can be measured using the threshold values. If the quality is poor or needs to be improved or updated, the topic generation agent can be used to further adjust it to achieve a complete closed loop from generation to evaluation to optimization.

[0082] In this scenario, an analytical agent can be configured. This agent can be pre-trained to generate adjustment information to reduce the target evaluation metric based on historical content and / or lower-level topic tags associated with the target evaluation metric. For example, if the second content coverage rate is less than a certain coverage threshold, the analytical agent can identify historical content in the historical content set that is not covered by the target topic tag set and generate "adjustment information" instructing the generating agent to generate new topic tags for this uncovered historical content to reduce the second content coverage rate.

[0083] For example, if the uniformity of the distribution of the third content type in the aforementioned third historical content subset is less than the corresponding uniformity threshold, the analysis agent can generate adjustment information based on the content categories that were not hit or covered, instructing the generation agent to generate new sub-topic tags for these content categories that were not hit or covered.

[0084] For example, if the semantic similarity between lower-level topic tags is greater than the similarity threshold, the analysis agent can generate information that considers the two to be "the same or similar topic tags" and generate adjustment information that requires adjusting at least one of them to avoid duplicate topic tags.

[0085] Accordingly, the topic-generating agent can avoid duplicate topic tags by adjusting one of them, or by choosing to delete one of them (either by itself or by instructing the executing agent to delete it).

[0086] Therefore, the implementing entity can respond when there is a target evaluation indicator whose value is greater than or equal to the threshold value of the corresponding evaluation indicator. The analytical agent can generate adjustment information to reduce the target evaluation indicator based on the historical content and / or lower-level topic tags associated with the target evaluation indicator.

[0087] Then, the implementing entity can use the topic-generating agent to update the target topic tag set based on this adjustment information. This improves the quality of the topic tags in the target topic tag set.

[0088] It should be understood that, in different embodiments, the aforementioned intelligent agents can be independent of each other or integrated into a complete intelligent agent system (e.g., they are intelligent agents within a system, or different functional modules within an intelligent agent). Furthermore, the aforementioned execution entity can also be understood in some scenarios as an "intelligent agent" (e.g., a scheduling intelligent agent, a data processing intelligent agent, etc.), or a module within the aforementioned intelligent agent system. For example, the aforementioned execution entity can actually be such an "intelligent agent system" deployed within server 105.

[0089] Next, we will discuss the process of using the aforementioned target topic tag set for topic labeling. For ease of understanding, we will also refer to... Figure 4 Let's have a discussion. Figure 4 A flowchart of a topic labeling process provided for an embodiment of this disclosure includes process 400.

[0090] Process 400 specifically includes: Step 401: Obtain the set of content to be annotated and target topic tags; In embodiments of this disclosure, this step is intended to be performed by the entity executing the method for labeling topics (e.g., Figure 1 The server 105 shown above obtains the content to be annotated (e.g., posts, audio, images, comments, etc., that belong to the same data source and content source as the historical content set) and the set of target topic tags currently available for use, generated by the process discussed above. For example, the set of target topic tags generated in process 200 above.

[0091] Step 402: Identify matching topic tags from the target topic tag set that match the content to be tagged; In the embodiments of this disclosure, based on the steps described above, the executing entity can determine matching topic tags from the target topic tag set that match the content to be labeled. For example, the executing entity can use semantic matching to match the topic tag with the highest semantic similarity as the matching topic tag that matches the content to be labeled.

[0092] In some embodiments, in this step, if the target topic tag set includes at least two levels of topic tags, the executing entity can choose to use a step-by-step matching method to completely determine the multiple levels of matching topic tags that match the content to be labeled. That is, in this step, if the target topic tag set includes at least two levels of topic tags, the executing entity can respond by first determining the first-level topic tags that match the content to be labeled from the first-level and top-level topic tags.

[0093] Then, if the first topic tag is actually associated with a subordinate topic tag, the executing entity can continue to select a second topic tag from its associated subordinate topic tags through methods such as semantic matching. Accordingly, if the topic tag in this round and at this level still has subordinates, the executing entity can repeat this process until the last level is reached.

[0094] Accordingly, if the second topic tag (in the current round) is the last-level topic tag in the target topic tag set, the executing entity can respond by identifying the first and second topic tags as matching topic tags that match the content to be annotated.

[0095] Therefore, the implementing entity can use the aforementioned tree-structured set of target topic tags to perform precise and hierarchical labeling, thereby improving the quality of topic labeling.

[0096] Step 403: Use matching topic tags to annotate the content to be annotated.

[0097] In the embodiments of this disclosure, the executing entity can label the content to be labeled with the matching topic tags determined in step 402, so as to provide users with topic tags and topic labeling results for summarizing the topics of the content to be labeled.

[0098] To enhance understanding, this disclosure also provides an implementation scheme for a specific application scenario. For ease of explanation, this will be referenced in conjunction with the provided scheme. Figure 5 Please provide an explanation. Figure 5 This is a flowchart illustrating the agent-based topic tag generation process in a specific application scenario provided by an embodiment of the present disclosure, including process 500.

[0099] For example, the server 105 described above can also be used as the "executing entity" in this process 500 for discussion and explanation.

[0100] In process 500, firstly, the executing entity can extract a first historical content subset and at least two second historical content subsets from the historical content set 510 by executing S501. For example, the first historical content subset 515, and second historical content subsets 521, 522, ..., 52N (where N is a positive integer) are extracted.

[0101] Then, the executing entity can continue to execute S502 to utilize the topic generating agent 530 to process the historical content included in the second historical content subset, and generate a set of topic tags to be evaluated corresponding to each of the second historical content subsets. For example, the set of topic tags to be evaluated 541 corresponding to the second historical content subset 521, the set of topic tags to be evaluated 542 corresponding to the second historical content subset 522, ..., the set of topic tags to be evaluated 54N corresponding to the second historical content subset 52N.

[0102] Then, the executing entity can continue to execute S503 to determine the target topic tag set from the topic tag set 541, topic tag set 542, ..., topic tag set 54N to be evaluated based on the first content coverage rate of the topic tag set 541, topic tag set 542, ..., topic tag set 54N to be evaluated for the first historical content subset 515.

[0103] For ease of explanation, it can be exemplarily assumed that the "first content coverage rate" corresponding to the second historical content subset 521 is the highest in this application scenario. Accordingly, after executing S503, the executing entity can use and determine the topic tag set 541 to be evaluated as the "target topic tag set".

[0104] Next, the executing entity can continue to execute S504 to use the evaluation agent 531 to process the set of topic tags to be evaluated 541 (i.e., the "target topic tag set"), and generate evaluation metrics 550 and 551 for the set of topic tags to be evaluated 541. It should be understood that the number of evaluation metrics in this step is merely an exemplary choice made for the convenience of discussion, and is not intended to impose any restrictions on the actual quantity of evaluation quality (i.e., it is not limited to generating "2 evaluation metrics").

[0105] For ease of discussion, in process 500, "evaluation index 550" can be exemplified as the target evaluation index whose index value is greater than or equal to the threshold value corresponding to evaluation index 550.

[0106] Accordingly, in such a situation, the executing entity can respond by choosing to continue executing S505 to use the analytical agent 532 to generate adjustment information 560 for reducing the evaluation 550 indicator based on the historical content and / or lower-level topic tags associated with the evaluation indicator 550.

[0107] It should be understood that, for ease of explanation, process 500 simplifies the process of detecting the existence of excessive topic tags and, if so, generating sub-topic tags. However, in practice, if the process of detecting the existence of excessive topic tags and generating "sub-topic tags" exists, the analysis agent 532 can, based on the specific circumstances, choose to generate adjustment information 560 for historical content and / or sub-topic tags.

[0108] Finally, the executing entity can execute S506 and use the topic generating agent 530 to update the previous target topic set, i.e. the set of topic tags to be evaluated 541, based on the adjustment information 560, to obtain the final updated target topic tag set 570 that will be used.

[0109] Further reference Figure 6 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a topic tag generation device based on intelligent agents. This device embodiment is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0110] like Figure 6 As shown, the agent-based topic tag generation device 600 of this embodiment may include: a subset extraction unit 601, a generator agent invocation unit 602, and a target set determination unit 603. The subset extraction unit 601 is configured to extract a first historical content subset and at least two second historical content subsets from the historical content set; the generator agent invocation unit 602 is configured to use a topic generation agent to process the historical content included in the second historical content subsets, generating a set of topic tags to be evaluated corresponding to each of the second historical content subsets; the target set determination unit 603 is configured to determine a target topic tag set from the set of topic tags to be evaluated based on a first content coverage rate of the set of topic tags to be evaluated relative to the first historical content subsets.

[0111] In this embodiment, the specific processing of the subset extraction unit 601, the agent generation invocation unit 602, and the target set determination unit 603 in the agent-based topic tag generation device 600, and the resulting technical effects, can be found in the following references: Figure 2The relevant descriptions of steps 201-203 in the corresponding embodiments will not be repeated here.

[0112] In some optional implementations of this embodiment, the first historical content subset is extracted based on the first content type distribution of the historical content set, and the similarity between the second content type distribution of the historical content in the first historical subset and the first content type distribution is greater than or equal to the similarity threshold.

[0113] In some optional implementations of this embodiment, the target set determination unit 603 includes: an intermediate set determination subunit, configured to extract an intermediate topic tag set whose first content coverage is greater than or equal to a coverage threshold based on the first content coverage rate of the topic tag set to be evaluated to the first historical content subset; an associated historical content determination subunit, configured to determine the associated historical content in the historical content set associated with each topic tag in the intermediate topic tag set; and a target set determination subunit, configured to use the intermediate topic tag set as the target topic tag set in response to the fact that the number of associated historical content associated with each topic tag in the intermediate topic tag set is less than or equal to a number threshold.

[0114] In some optional implementations of this embodiment, the target set determination unit 603 further includes: a lower-level tag generation subunit, configured to generate lower-level topic tags associated with the excess topic tags using a topic generation agent in response to the existence of excess topic tags in the intermediate topic tag set; wherein the number of associated historical content associated with the excess topic tags is greater than a quantity threshold; and a topic tag combination subunit, configured to obtain the target topic tag set by combining the intermediate topic tag set and the lower-level topic tags.

[0115] In some optional implementations of this embodiment, the lower-level tag generation subunit includes: an associated historical content extraction module, configured to determine a third historical content subset from the historical content associated with the excess topic tags in response to the existence of excess topic tags in the intermediate topic tag set; and a lower-level tag generation module, configured to use a topic generation agent to process the historical content included in the third historical content subset and generate lower-level topic tags associated with the excess topic tags.

[0116] In some optional implementations of this embodiment, the lower-level tag generation subunit is further configured to generate lower-level topic tags associated with each of the excess topic tags in parallel using a topic generation agent in response to the existence of at least two excess topic tags in the intermediate topic tag set.

[0117] In some optional implementations of this embodiment, the lower-level tag generation subunit is further configured to, in response to the existence of excess topic tags in the intermediate topic tag set, use a topic generation agent to generate lower-level topic tags associated with the excess topic tags, as well as usage constraint information between lower-level topic tags belonging to the same excess topic tag.

[0118] In some optional implementations of this embodiment, the apparatus 600 further includes: an evaluation agent invocation unit, configured to use the evaluation agent to process the target topic tag set and generate an evaluation index of the target topic tag set, wherein the evaluation index includes at least one of the following: the second content coverage of the target topic tag set to the historical content set, the distribution uniformity of the third content type distribution in the third historical content subset, and the semantic similarity between lower-level topic tags.

[0119] In some optional implementations of this embodiment, the apparatus 600 further includes: an analysis agent invocation unit, configured to, in response to a target evaluation indicator in which the indicator value is greater than or equal to the indicator value threshold corresponding to the evaluation indicator, generate adjustment information for reducing the target evaluation indicator using the analysis agent based on historical content and / or lower-level topic tags associated with the target evaluation indicator; and a topic generation agent update invocation unit, configured to update the target topic tag set using the topic generation agent based on the adjustment information.

[0120] In some optional implementations of this embodiment, the historical content set is periodically collected by a data collection agent.

[0121] In some optional implementations of this embodiment, the number of historical contents included in the first historical content subset and the second historical content subset is the same.

[0122] This embodiment exists as a device embodiment corresponding to the above method embodiment. The topic tag generation device based on intelligent agents provided in this embodiment can not only pre-generate a set of topic tags for subsequent annotation actions based on real historical content, so as to effectively manage the topic tags actually used and avoid the instability of annotation quality caused by the divergence of temporarily generated tags, but also ensure the quality of topic tags in the topic tag set by using an internal loop feedback mechanism of generation-evaluation from the same historical content during the generation and determination of topic tags.

[0123] Further reference Figure 7 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a device for labeling topics, which is similar to... Figure 4 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0124] like Figure 7 As shown, the topic labeling device 700 in this embodiment may include: a content to be labeled and topic tag acquisition unit 701, a topic tag matching unit 702, and a topic labeling unit 703. The content to be labeled and topic tag acquisition unit 701 is configured to acquire the content to be labeled and a target topic tag set, wherein the target topic tag set is obtained based on the agent's topic tag generation device 600; the topic tag matching unit 702 is configured to determine matching topic tags from the target topic tag set that match the content to be labeled; and the topic labeling unit 703 is configured to label the content to be labeled using the matching topic tags.

[0125] In this embodiment, the specific processing and technical effects of the topic labeling device 700, including the content to be labeled and topic tag acquisition unit 701, topic tag matching unit 702, and topic labeling unit 703, can be found in reference to [reference needed]. Figure 4 The relevant descriptions of steps 401-403 in the corresponding embodiments will not be repeated here.

[0126] In some optional implementations of this embodiment, the topic tag matching unit 702 includes: a first-level tag matching subunit, configured to match a first-level topic tag that matches the content to be labeled from the first-level topic tags in response to the target topic tag set including at least two levels of topic tags; a lower-level tag matching subunit, configured to match a second topic tag from the lower-level topic tags in response to the first topic tag being associated with a lower-level topic tag; and a multi-level topic tag matching subunit, configured to determine the first topic tag and the second topic tag as matching topic tags that match the content to be labeled in response to the second topic tag being the last-level topic tag in the target topic tag set.

[0127] This embodiment exists as a device embodiment corresponding to the above method embodiment. The device for labeling topics provided in this embodiment can use the generated set of topic tags to perform topic labeling, thereby achieving higher quality topic labeling.

[0128] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0129] Figure 8A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0130] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0131] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0132] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as methods for generating and labeling topics based on intelligent agents. For example, in some embodiments, the methods for generating and labeling topics based on intelligent agents can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the methods for generating and labeling topics based on intelligent agents described above can be performed. Alternatively, in other embodiments, computing unit 801 may be configured in any other suitable manner (e.g., by means of firmware) to perform a method for generating and labeling topics based on agents.

[0133] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0134] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0135] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0136] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0137] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0138] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, also known as cloud computing servers or cloud hosts, which are hosting products within the cloud computing service ecosystem to address the management difficulties and weak business scalability inherent in traditional physical hosts and Virtual Private Servers (VPS) services. Servers can also be categorized as distributed system servers or servers incorporating blockchain technology.

[0139] According to the technical solution of this disclosure, not only can a set of topic tags for subsequent annotation actions be pre-generated based on real historical content to effectively manage the topic tags actually used and avoid unstable annotation quality caused by the dispersion of temporarily generated tags, but also, during the generation and determination of topic tags, the quality of topic tags in the set can be guaranteed by using an internal feedback mechanism of generation-evaluation from the same historical content. Accordingly, the generated set of topic tags can also be used for topic annotation, achieving higher-quality topic annotation.

[0140] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution provided in this disclosure can be achieved, and this is not limited herein.

[0141] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for generating topic tags based on intelligent agents, comprising: Extract a first subset of historical content and at least two second subsets of historical content from the set of historical content; For the second subset of historical content, a topic generation agent is used to process the historical content included in the second subset of historical content and generate a set of topic tags to be evaluated corresponding to each of the second subset of historical content. Based on the first content coverage rate of the set of topic tags to be evaluated to the first historical content subset, the target topic tag set is determined from the set of topic tags to be evaluated.

2. The method according to claim 1, wherein, The first historical content subset is extracted based on the first content type distribution of the historical content set, and the similarity between the second content type distribution of the historical content in the first historical subset and the first content type distribution is greater than or equal to a similarity threshold.

3. The method according to claim 1, wherein, The step of determining the target topic tag set from the topic tag set to be evaluated based on the first content coverage rate of the topic tag set to be evaluated relative to the first historical content subset includes: Based on the first content coverage rate of the set of topic tags to be evaluated to the first historical content subset, extract the intermediate set of topic tags whose first content coverage rate is greater than or equal to the coverage rate threshold. Determine the associated historical content in the historical content set and the topic tags in the intermediate topic tag set; In response to the fact that the number of associated historical contents associated with each topic tag in the intermediate topic tag set is less than or equal to the number threshold, the intermediate topic tag set is taken as the target topic tag set.

4. The method according to claim 3, further comprising: In response to the existence of an excess of topic tags in the intermediate topic tag set, the topic generation agent generates a lower-level topic tag associated with the excess topic tags; wherein the number of associated historical content associated with the excess topic tags is greater than a quantity threshold. The target topic tag set is obtained by combining the intermediate topic tag set and the lower-level topic tags.

5. The method according to claim 4, wherein, The response to the existence of an excess of topic tags in the intermediate topic tag set, utilizing the topic generation agent to generate lower-level topic tags associated with the excess topic tags, includes: In response to the existence of an excess of topic tags in the intermediate topic tag set, a third subset of historical content is determined from the historical content associated with the excess of topic tags; The topic-generating agent processes the historical content included in the third historical content subset and generates sub-topic tags associated with the super-quantity topic tags.

6. The method according to claim 4, wherein, The response to the existence of an excess of topic tags in the intermediate topic tag set, utilizing the topic generation agent to generate lower-level topic tags associated with the excess topic tags, includes: In response to the existence of at least two excess topic tags in the intermediate topic tag set, the topic generation agent generates subordinate topic tags associated with each of the excess topic tags in parallel.

7. The method according to claim 4, wherein, The response to the existence of an excess of topic tags in the intermediate topic tag set, utilizing the topic generation agent to generate lower-level topic tags associated with the excess topic tags, includes: In response to the existence of an excess of topic tags in the intermediate topic tag set, the topic generation agent generates subordinate topic tags associated with the excess topic tags, as well as usage constraint information between the subordinate topic tags belonging to the same excess topic tag.

8. The method according to any one of claims 1-7, wherein, The method further includes: An evaluation agent is used to process a target topic tag set and generate an evaluation metric for the target topic tag set. The evaluation metric includes at least one of the following: the second content coverage of the target topic tag set to the historical content set, the distribution uniformity of the third content type distribution in the third historical content subset, and the semantic similarity between the lower-level topic tags.

9. The method according to claim 8, further comprising: In response to the existence of a target evaluation indicator whose value is greater than or equal to the threshold value corresponding to the evaluation indicator, the analysis agent generates adjustment information for reducing the target evaluation indicator based on the historical content and / or the lower-level topic tags associated with the target evaluation indicator. The topic-generating agent updates the target topic tag set based on the adjustment information.

10. The method according to claim 1, wherein, The historical content set is collected periodically by a data collection agent.

11. The method according to claim 1, wherein, The first and second subsets of historical content contain the same amount of historical content.

12. A method for labeling topics, comprising: Obtain the content to be labeled and the target topic tag set, wherein the target topic tag set is obtained based on the agent-based topic tag generation method according to any one of claims 1-11; From the target topic tag set, determine the matching topic tags that match the content to be tagged; The content to be annotated is labeled using the matching topic tags.

13. The method according to claim 12, wherein, The step of determining matching topic tags from the target topic tag set that match the content to be tagged includes: In response to the fact that the target topic tag set includes at least two levels of topic tags, the first-level topic tags that match the content to be labeled are matched from the first-level topic tags; In response to the fact that the first topic tag is associated with a subordinate topic tag, a second topic tag is matched from the subordinate topic tag; In response to the second topic tag being the last-level topic tag in the target topic tag set, the first topic tag and the second topic tag are determined as matching topic tags that match the content to be labeled.

14. A topic tag generation device based on intelligent agents, comprising: The subset extraction unit is configured to extract a first subset of historical content and at least two second subsets of historical content from the historical content set; The intelligent agent invocation unit is configured to process the historical content included in the second historical content subset using the topic generation intelligent agent, and generate a set of topic tags to be evaluated corresponding to each of the second historical content subsets. The target set determination unit is configured to determine the target topic tag set from the topic tag set to be evaluated based on the first content coverage rate of the topic tag set to be evaluated relative to the first historical content subset.

15. A device for labeling topics, comprising: The unit for obtaining content to be annotated and topic tags is configured to obtain the content to be annotated and a set of target topic tags, wherein the set of target topic tags is obtained based on the agent-based topic tag generation device as described in claim 14; The topic tag matching unit is configured to determine from the target topic tag set a matching topic tag that matches the content to be tagged; The topic labeling unit is configured to label the content to be labeled using the matching topic tags.

16. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which enable the at least one processor to perform the agent-based topic tag generation method of any one of claims 1-11, and / or the topic tagging method of any one of claims 12-13.

17. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the agent-based topic tag generation method of any one of claims 1-11, and / or the topic tagging method of any one of claims 12-13.

18. A computer program product comprising a computer program that, when executed by a processor, implements the agent-based topic tag generation method according to any one of claims 1-11, and / or the topic tagging method according to any one of claims 12-13.