Method and system for automatic generation of hierarchical semantic processing tags for web content
By constructing hierarchical semantic tag sets and entity links with domain knowledge graphs, the problem of low efficiency in semantic processing of network content in existing technologies is solved, achieving efficient and accurate semantic tag generation and analysis, and improving the targeting of content creation and operation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG GUANYU NETWORK TECHNOLOGY CO LTD
- Filing Date
- 2025-10-27
- Publication Date
- 2026-06-02
Smart Images

Figure CN121412433B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of semantic processing technology, specifically to a method and system for automatically generating hierarchical semantic processing tags for web content. Background Technology
[0002] With the rapid development of the internet, online content has become massive and diverse, encompassing various forms such as social media posts, short video comments, online articles, and forum discussions, exhibiting rich variety and depth. However, current semantic processing of online content generally relies on static methods such as keyword matching, rule-based classification, or simple sentiment analysis. This results in semantic tags that are often one-sided and fragmented, failing to accurately reveal the core theme and extended concepts of a work or content. Furthermore, cross-work or competitor comparative analysis requires extensive manual processing and experience-based judgment, leading to low efficiency and susceptibility to subjective biases. Consequently, the generation of hierarchical semantic tags is inefficient, and the analytical accuracy is limited, making it difficult to provide content creators and operators with comprehensive, effective, and reliable content analysis and decision support.
[0003] In summary, existing technologies suffer from technical problems such as low efficiency in generating hierarchical semantic tags and limited analysis accuracy due to the static and rule-dependent semantic processing mode, which makes it difficult to fully and effectively perceive network content. Summary of the Invention
[0004] The purpose of this application is to provide a method and system for automatically generating hierarchical semantic processing tags for web content, in order to solve the technical problems in the prior art where the semantic processing mode is static and rule-dependent, making it difficult to fully and effectively perceive web content, resulting in low efficiency in generating hierarchical semantic tags and limited analysis accuracy.
[0005] To achieve the above objectives, this application provides a method and system for automatically generating hierarchical semantic processing tags for web content.
[0006] Firstly, this application provides a method for automatically generating hierarchical semantic processing tags for online content. The method includes: receiving a semantic analysis request from a target user and parsing to obtain a first work and a second work; extracting core entities and performing entity derivation analysis on the first work and the second work respectively, establishing retrieval vectors, and collecting a first online content set corresponding to the first work and a second online content set corresponding to the second work from online data sources; extracting key semantic concepts and associated sentiment polarities from the first online content set and the second online content set respectively, linking them with pre-constructed domain knowledge graphs to construct a first hierarchical semantic tag set and a second hierarchical semantic tag set; comparing the first hierarchical semantic tag set and the second hierarchical semantic tag set, performing tag alignment classification and semantic aggregation based on knowledge domain and sentiment polarity, and generating a hierarchical comparison report.
[0007] Optionally, the first work is a work operated by the target user on the network, and the second work is a competing work of the first work.
[0008] Optionally, for the first work and the second work, the work name, author, and main character name are extracted from the metadata as the first core entity and the second core entity, respectively; using the first core entity and the second core entity as query terms, the first derivative text content set and the second derivative text content set are retrieved and crawled from a preset online community; word frequency co-occurrence statistics are performed on the first derivative text content set and the second derivative text content set to extract the first derivative words that frequently co-occur with the first core entity and the second derivative words that frequently co-occur with the second core entity; a first retrieval vector is established based on the first core entity and the first derivative words, and a second retrieval vector is established based on the second core entity and the second derivative words.
[0009] Optionally, the derivative vocabulary may include at least derivative terms related to characters, skills, unique items, and core plot points.
[0010] Optionally, Step 1: Perform sentence-level segmentation and semantic vectorization on each content fragment in the first network content set to generate several first fragment semantic concepts; Step 2: Link the several first fragment semantic concepts to the corresponding entity nodes of the domain knowledge graph; Step 3: Based on the tree structure of the domain knowledge graph, organize the linked semantic concepts according to their categories and associate them with the corresponding sentiment polarities in the original content fragments to generate a first hierarchical semantic tag set with knowledge domains as branches and semantic concepts as leaves; Perform Steps 1 to 3 on the second network content set to generate a second hierarchical semantic tag set.
[0011] Optionally, the vector similarity between the plurality of first fragment semantic concepts and entity nodes in the domain knowledge graph is calculated; based on the vector similarity, for first fragment semantic concepts with a similarity exceeding a preset threshold, they are directly linked to the corresponding entity nodes in the domain knowledge graph.
[0012] Optionally, if there exists a first fragment semantic concept whose vector similarity to any entity node in the domain knowledge graph is less than the preset threshold, a link failure is identified; the identified semantic concept with the link failure is extracted, input into a pre-trained semantic classification model, and the knowledge domain category to which it belongs is output; the identified semantic concept is temporarily attached as a new leaf node under the knowledge domain category to which it belongs.
[0013] Optionally, the first hierarchical semantic tag set and the second hierarchical semantic tag set are traversed, and matching is performed according to the knowledge domain nodes to which they belong, so that tags with similar attributes are grouped into the same comparison group; within each comparison group, the number of semantic concepts and the distribution of sentiment polarity belonging to the first work and the second work are counted respectively, the positive voice volume ratio and discussion popularity index of the first work in the same knowledge domain are calculated, and a differential comparison is made with the corresponding indicators of the second work; based on the differential comparison results, the advantageous and disadvantageous areas of the first work relative to the second work are automatically identified, and the hierarchical comparison report is constructed.
[0014] Optionally, the newly attached semantic concepts and their associations with knowledge domain nodes are submitted to the review queue; after receiving the administrator's approval instruction, the new semantic concepts and their associations are updated in the domain knowledge graph.
[0015] Secondly, this application also provides an automatic generation system for hierarchical semantic processing tags for online content. The system includes: a request parsing module, used to receive semantic analysis requests from target users and parse to obtain a first work and a second work; a data acquisition module, used to extract core entities and perform entity derivation analysis on the first work and the second work respectively, establish retrieval vectors, and collect a first online content set corresponding to the first work and a second online content set corresponding to the second work from online data sources; a semantic tag construction module, used to extract key semantic concepts and associated sentiment polarities from the first online content set and the second online content set respectively, link them with pre-built domain knowledge graphs, and construct a first hierarchical semantic tag set and a second hierarchical semantic tag set; and a report generation module, used to compare the first hierarchical semantic tag set and the second hierarchical semantic tag set, perform tag alignment classification and semantic aggregation on knowledge domain and sentiment polarity, and generate a hierarchical comparison report.
[0016] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0017] By receiving semantic analysis requests from target users, the system parses and retrieves a first and a second work. It then extracts core entities and performs entity derivation analysis on both works, establishing retrieval vectors. The system collects a first set of online content corresponding to the first work and a second set of online content corresponding to the second work from online data sources. Key semantic concepts and associated sentiment polarities are extracted from both sets and linked to a pre-constructed domain knowledge graph to build a first-level and a second-level semantic tag set. By comparing these two sets, the system performs tag alignment, classification, and semantic aggregation based on knowledge domain and sentiment polarity, generating a hierarchical comparison report. This achieves comprehensive and effective perception of online content, efficient and accurate generation of precise semantic tags, improved semantic processing efficiency and analysis accuracy, thereby better meeting user needs and enhancing the relevance and application value of content creation and operation.
[0018] The above description is merely an overview of the technical solution of this application. To enable a clearer understanding of the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0020] Figure 1 This is a flowchart illustrating the hierarchical semantic processing tag automatic generation method for web content according to this application.
[0021] Figure 2 This is a schematic diagram of the hierarchical semantic processing tag automatic generation system for web content in this application;
[0022] Figure labeling: Request parsing module 11, data acquisition module 12, semantic tag construction module 13, report generation module 14. Detailed Implementation
[0023] This application provides a method and system for automatically generating hierarchical semantic processing tags for web content. It addresses the technical problems in existing technologies where the static and rule-dependent semantic processing models make it difficult to comprehensively and effectively perceive web content, leading to low efficiency in hierarchical semantic tag generation and limited analysis accuracy. The method achieves the technical effect of comprehensively and effectively perceiving web content, efficiently and accurately generating precise semantic tags, and improving semantic processing efficiency and analysis accuracy.
[0024] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be understood that this application is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. It should also be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all of them.
[0025] Example 1, please refer to the appendix. Figure 1 This application provides a method for automatically generating hierarchical semantic processing tags for web content, wherein the method specifically includes:
[0026] Receive semantic analysis requests from target users, and parse them to obtain the first and second works.
[0027] Furthermore, the first work is a work operated online by the target user, and the second work is a competing work of the first work.
[0028] Specifically, target users upload semantic analysis requests to the server via natural language input boxes or forms. Target users refer to content creators or operators of works running on online platforms, such as internet works operating on short video platforms, social media, and online cultural and creative platforms, including but not limited to operators or authors of novels, comics, and games. The semantic analysis request refers to a set of analysis instructions submitted by the target user through an interactive interface, open API, or platform access service. This includes analysis objectives, analysis object information, analysis scope, and analysis preferences. Analysis objectives may include, for example, analyzing the semantic features of the work's content and differences from competitors; analysis object information may include, for example, the work's name, author, type tags, and content; the analysis scope may include, for example, the selected content source, time period, and network platform type; and output preferences may include, for example, a knowledge domain layered report.
[0029] After receiving a semantic analysis request from a target user, the server uses natural language processing (NLP) technology to perform semantic recognition and parameter extraction on the user's input request. This process identifies and retrieves the first and second works. The first work is an original work operated, published, or managed by the target user on the online platform, including any type of digital work such as short videos, novels, games, comics, animations, and film / television content, which exhibits continuous user interaction and content dissemination behavior within the online environment. The second work is a competing work that rivals the first work in terms of content theme, target audience, or market positioning. It can be automatically generated by the server or actively specified by the user. For example, when the user inputs "Please help me analyze the differences between my work A and its competitors" in natural language, the server uses NLP technology for intent recognition and named entity recognition (NER) to extract "A" as the name of the first work. Then, using built-in semantic retrieval technology, it searches databases or online content libraries for works with high similarity in theme, type, tags, or audience, automatically identifying the corresponding competing work, i.e., the second work B. If a user enters "First work: A, Second work: B, please specify the advantages and disadvantages of work A compared to work B" via a form, the field content will be read directly to obtain the first and second works.
[0030] By parsing user requests and identifying the subject of the work, the system achieves accurate acquisition of the semantic analysis object from the user's intent, thereby improving the targeting and reliability of hierarchical semantic processing.
[0031] The core entities of the first work and the second work are extracted and entity derivation analysis is performed respectively to establish retrieval vectors. The first network content set corresponding to the first work and the second network content set corresponding to the second work are collected from network data sources.
[0032] Furthermore, core entities are extracted and entity derivation analysis is performed on the first work and the second work respectively to establish retrieval vectors, including: extracting the work name, author, and main character name from the metadata of the first work and the second work respectively as the first core entity and the second core entity; using the first core entity and the second core entity as query terms, retrieving and crawling the first and second derivative text content sets from a preset online community; performing word frequency co-occurrence statistics on the first and second derivative text content sets to extract the first derivative words that frequently co-occur with the first core entity and the second derivative words that frequently co-occur with the second core entity; establishing a first retrieval vector based on the first core entity and the first derivative words, and establishing a second retrieval vector based on the second core entity and the second derivative words.
[0033] Specifically, after acquiring the first and second works, their metadata is parsed to extract representative semantic elements, including the work's title, author, and main character names. These extracted semantic elements are then integrated to form the first core entity of the first work and the second core entity of the second work. Metadata refers to basic information and descriptive data related to the work, containing its core information. Core entities are the basic units reflecting the work's main semantic features and narrative subject. For example, for novels, core entities include the title, author, and characters; for film and television works, core entities also include the lead actors and iconic scene names.
[0034] The first and second core entities are used as search terms and input into a preset online community data source for semantic retrieval and content collection. These online communities are the main venues for user discussions of works, containing a large amount of text content related to the works, such as social platforms and content communities like Weibo, Baidu Tieba, Bilibili, Douban, and Zhihu. For example, using work A and the names of its related authors and main characters A1 and A2 as keywords, a search on Weibo yields a large amount of user discussion content about work A and its protagonist. Using web crawling technology, text content related to the first and second core entities is crawled from the online communities and integrated to form a first and second derivative text content set. The text content in the derivative text content set includes user comments, discussion posts, short video descriptions, or creative secondary interpretation texts, reflecting the dissemination characteristics and audience feedback of the works in the online environment.
[0035] Preprocess the first derivative text content set and the second derivative text content set obtained by crawling, including removing stop words, punctuation marks and special characters. Stop words such as "de", "shi", "zai", "ni", "wo" and other common but meaningless words. Perform word frequency co-occurrence statistics on the preprocessed text, count the number of times each word and the core entity co-occur in the same window, such as a sentence or a paragraph in the text, and combine with the Pointwise Mutual Information (PMI) or TF-IDF algorithm to identify the high-frequency words with significant semantic associations with the core entity, obtaining the first derivative words and the second derivative words, which can further expand the semantic boundaries of the role. After extracting the first derivative words and the second derivative words, through a semantic embedding model, such as BERT or SimCSE, vectorize and identify the first core entity with the corresponding first derivative words to construct the first retrieval vector. Similarly, vectorize and identify the second core entity with the corresponding second derivative words to construct the second retrieval vector. Among them, the retrieval vector is a mathematical expression for information retrieval, which converts the text content into a vector form for convenient information retrieval. Perform semantic retrieval and content aggregation from the network data source according to the constructed first retrieval vector to collect the first network content set corresponding to the first work. At the same time, perform semantic retrieval and content aggregation from the network data source according to the constructed second retrieval vector to collect the second network content set corresponding to the second work. The first network content set and the second network content set respectively represent the most relevant text information sets of the first work and the second work in the network semantic space, avoiding the interference of irrelevant content.
[0036] Through the automatic expansion and text collection from work metadata to the network semantic space, it is possible to comprehensively and effectively perceive network content, not only improving the efficiency and coverage of text collection, but also improving the accuracy of semantic association, thereby enhancing the precision and reliability of hierarchical semantic analysis.
[0037] Furthermore, the derivative words at least include derivative terms regarding characters, skills, exclusive props, and core plots.
[0038] Specifically, the derivative words at least include derivative semantic contents such as work characters, skill settings, exclusive props, and core plots. For example, the character names and their aliases in the work, or the important plots or events in the work. By collecting the discussions related to the plot, it is possible to analyze the user's evaluation and emotional response to the plot. By setting the derivative words for retrieval, the retrieval vector can collect various discussion contents related to the work more comprehensively and accurately, avoiding information omission and improving the comprehensiveness and accuracy of information.
[0039] Extract the key semantic concepts and associated emotional polarities from the first network content set and the second network content set respectively, and perform entity linking with the pre-constructed domain knowledge graph to construct the first hierarchical semantic label set and the second hierarchical semantic label set.
[0040] Furthermore, key semantic concepts and associated sentiment polarities are extracted from the first and second network content sets respectively, and entity links are established with the pre-constructed domain knowledge graph to construct a first-level semantic tag set and a second-level semantic tag set. This includes: Step 1: Sentence-level segmentation and semantic vectorization are performed on each content fragment in the first network content set to generate several first fragment semantic concepts; Step 2: The several first fragment semantic concepts are linked to the corresponding entity nodes in the domain knowledge graph; Step 3: Based on the tree structure of the domain knowledge graph, the linked semantic concepts are organized according to their categories and associated with the corresponding sentiment polarities in the original content fragments to generate a first-level semantic tag set with knowledge domains as branches and semantic concepts as leaves; Steps 1 to 3 are performed on the second network content set to generate the second-level semantic tag set.
[0041] Specifically, sentence-level segmentation is performed on each content fragment in the first network content set. For example, two methods are used: rule-based sentence segmentation and machine learning model-based sentence segmentation. Rule-based sentence segmentation initially segments each content fragment based on punctuation marks and conjunctions, such as periods, question marks, exclamation marks, and commas, and conjunctions such as "because," "but," and "at the same time." Machine learning-based sentence segmentation uses the contextual attention mechanism of pre-trained language models, such as oBERTa or MacBERT, to identify semantic breakpoints and further segment semantically independent segments that are not explicitly separated by punctuation. By combining these methods, each content fragment in the first network content set is accurately segmented into several sentences with complete semantic expression, avoiding information fragmentation or semantic aliasing. A BERT, SimCSE, or Sentence-BERT semantic encoding model is used to perform semantic vectorization on each segmented sentence. Each sentence is input into the semantic encoding model, and through attention host and context adaptive weighting, a semantic vector is output. This semantic vector simultaneously preserves semantic theme, sentiment features, and contextual relevance. Semantic clustering and deduplication are performed on all obtained semantic vectors to remove semantically repetitive segments. Semantic clustering uses K-Means or DBSCAN methods to cluster based on the cosine similarity between semantic vectors, generating several first fragment semantic concepts. Each fragment semantic concept corresponds to the core idea or theme of a content fragment, such as character growth, plot twist, audience resonance, etc.
[0042] Several semantic concepts are linked to a pre-constructed domain knowledge graph, which is a structured knowledge network composed of a large number of entity nodes and semantic relationship edges. This graph is used to express the knowledge hierarchy and conceptual associations within a specific domain. Specifically, knowledge information is collected from multi-source data related to the domain, including encyclopedia entries, industry literature, user comments, and expert-annotated corpora, forming an initial knowledge corpus. Key entities and semantic relationships are identified from the corpus using named entity recognition and relation extraction algorithms. Key entities include, for example, roles, events, scenes, and attributes; semantic relationships include, for example, belonging, parameters, influence, and comparison. Entity relationship pairs are constructed based on the identified key entities and semantic relationships. Ontology modeling methods are used to classify and layer entities according to domain logic, establishing a multi-layered structure including a topic layer, entity layer, and instance layer, giving the knowledge a clear hierarchical logic. Finally, a graph database, such as Neo4j, is used to fuse and store the graphs, forming the domain knowledge graph.
[0043] The matching relationship between semantic concepts and graph nodes is determined by calculating the similarity between semantic concept vectors and entity node vectors in the knowledge graph. When the similarity exceeds a preset threshold, a link is established between the semantic concept and the corresponding entity node. Based on the tree structure of the domain knowledge graph, the successfully linked semantic concepts are hierarchically organized according to categories such as character traits, plot settings, and audience evaluations, forming semantic tags with the knowledge domain as the backbone and semantic concepts as the leaves. Simultaneously, a sentiment analysis model, such as the BiLSTM-Attention model, is used to identify the emotional polarity expressed in each content fragment, including positive, neutral, or negative. For example, through BiLSTM-Attention model analysis, if a content fragment contains praise for the plot and appreciation for skills (e.g., the plot is excellent), the overall emotion is positive, and the emotional polarity is identified as positive. If a content fragment expresses a feeling that the plot is okay, uses relatively neutral language, and does not have a clear positive or negative emotional tendency, the emotional polarity is neutral. If the content fragments contain criticism and dissatisfaction with the work, such as a dragging plot or an overall negative emotional expression, then the emotional polarity is identified as negative. The emotional polarity is then associated with the corresponding semantic concept nodes to generate a first-level semantic tag set. Each tag includes a corresponding semantic category and emotional attribute, achieving a fusion of knowledge semantics and emotional dimensions. For the second set of network content, the same process is executed to generate a corresponding second-level semantic tag set.
[0044] By using sentence-level segmentation and semantic vectorization, key semantic concepts of each content fragment can be accurately extracted and classified and organized according to the tree structure of the domain knowledge graph. Through sentiment polarity association, it can not only understand the content theme but also perceive the user's emotional tendency, thereby automatically identifying the strengths and weaknesses of the work in different knowledge domains. This provides content creators and operators with reliable and targeted content optimization directions, effectively improving the content quality, audience relevance, and market dissemination effect of the work.
[0045] By comparing the first hierarchical semantic tag set with the second hierarchical semantic tag set, tag alignment classification and semantic aggregation based on knowledge domain and sentiment polarity are performed to generate a hierarchical comparison report.
[0046] Furthermore, by comparing the first hierarchical semantic tag set and the second hierarchical semantic tag set, tag alignment classification and semantic aggregation are performed based on knowledge domain and sentiment polarity to generate a hierarchical comparison report. This includes: traversing the first hierarchical semantic tag set and the second hierarchical semantic tag set, matching them according to their respective knowledge domain nodes, and grouping tags with similar attributes into the same comparison group; within each comparison group, counting the number of semantic concepts and sentiment polarity distribution belonging to the first work and the second work respectively, calculating the positive voice ratio and discussion popularity index of the first work in the same knowledge domain, and comparing them with the corresponding indicators of the second work; based on the comparison results, automatically identifying the advantageous and disadvantageous areas of the first work relative to the second work, and constructing the hierarchical comparison report.
[0047] Specifically, the process iterates through the first-level and second-level semantic tag sets, matching tags based on their respective knowledge domain nodes. These knowledge domain nodes refer to the various levels of classification in the tree structure defined in the domain knowledge graph, such as character settings, plot development, props and equipment, and audience feedback. Tags belonging to the same knowledge domain are grouped into the same comparison group, forming a semantic set for comparison. For example, in the character comparison group, the number of positive, neutral, and negative emotional concepts for each character in the first and second works is counted.
[0048] Within each comparison group, the number of semantic concepts and the emotional polarity distribution belonging to the first and second works are statistically analyzed. The number of semantic concepts refers to the number of semantic tags contained in each work within the stated domain, reflecting the breadth of the work's content coverage. The emotional polarity distribution includes the proportion of positive, neutral, and negative tags, reflecting audience feedback or textual emotional tendencies. Based on the statistical data, the positive voice share of the first work in the same knowledge domain is calculated, i.e., the proportion of positive semantic concepts in that domain relative to the total number of semantic concepts. Simultaneously, a discussion heat index is calculated based on the number of semantic concepts to quantify audience attention and discussion heat. For example, if the first work has 100 positive emotional concepts in the character domain, and the total number of emotional concepts is 200, then the positive voice share is 50%. The discussion heat index can be calculated by taking the logarithm of the number of semantic concepts and multiplying it by a time decay coefficient. Similarly, the positive voice share and discussion heat index of the second work in the same knowledge domain are calculated, allowing for a differentiated comparison of the indicators of the first and second works in each comparison group. By comparing the proportion of positive voice and the discussion popularity index, the system can automatically identify the strengths and weaknesses of the first work compared to the second work. Strengths refer to areas where the work performs exceptionally well and receives a lot of positive audience feedback, while weaknesses refer to areas where coverage is limited and emotional feedback is predominantly negative. A hierarchical comparison report is generated based on these strengths and weaknesses, visually demonstrating the semantic differences between the first work and its competitor, the second work, across various knowledge domains.
[0049] By aligning knowledge domains and aggregating sentiment polarity, the strengths and weaknesses of a work are quantified, revealing the specific semantic and audience feedback differences between the first and second works. This improves the accuracy and efficiency of the entire online content processing at the semantic level, providing targeted creative optimization directions for content creators and operators. It also provides a scientific basis for competitor analysis, market positioning, and work improvement, enhancing the online performance and user satisfaction of works, thereby increasing their market competitiveness and user stickiness.
[0050] Furthermore, linking the plurality of first fragment semantic concepts to the corresponding entity nodes of the domain knowledge graph includes: calculating the vector similarity between the plurality of first fragment semantic concepts and the entity nodes in the domain knowledge graph; and, based on the vector similarity, directly linking the first fragment semantic concepts whose similarity exceeds a preset threshold to the corresponding entity nodes of the domain knowledge graph.
[0051] Specifically, the similarity between vectors of several first-fragment semantic concepts and entity nodes in the domain knowledge graph is calculated. For example, cosine similarity can be used to determine the similarity between the first-fragment semantic concept vector and the entity node in the domain knowledge graph. The closer the calculated cosine value is to 1, the more semantically similar the two are. To improve computational efficiency, vector indexing structures such as FAISS or Annoy can be used to quickly retrieve and calculate similarity, obtaining multiple vector similarities. By comprehensively considering practical needs, a preset threshold is set. For example, by statistically analyzing the correctly linked entity nodes and semantic concepts in the domain knowledge graph, the average and median similarity values are calculated and used as the preset threshold. If the average similarity between correctly linked semantic concepts and entity nodes is 0.85, the preset threshold is set to 0.85. If the actual application scenario requires high link accuracy and emphasizes link accuracy, hoping that the semantic concepts of the linked entity nodes are as correct as possible, the threshold is set relatively higher, such as 0.9. Conversely, if it is desired that as many relevant semantic concepts as possible are linked to entity nodes, the threshold can be set relatively lower, such as 0.8. Multiple vector similarities are compared with the preset threshold. When the calculated similarity exceeds, i.e., is greater than or equal to, the preset threshold, it indicates that the semantic concept is highly similar to a certain entity node in the knowledge graph. The first fragment semantic concept is then directly linked to the corresponding entity node in the domain knowledge graph, realizing the structured association between semantics and knowledge. By calculating and establishing links based on vector similarity, a high-precision link between semantic concepts and domain knowledge is achieved, transforming textual information into traceable semantic tags. This enables semantic concepts to obtain clear domain positioning and hierarchical affiliation in the knowledge graph, thereby improving the accuracy, effectiveness, and reliability of hierarchical semantic tags and analysis.
[0052] Furthermore, after calculating the vector similarity between the several first fragment semantic concepts and entity nodes in the domain knowledge graph, the method further includes: if there is a first fragment semantic concept whose vector similarity with any entity node in the domain knowledge graph is less than the preset threshold, a link failure flag is set; the flag semantic concept with the link failure flag is extracted, input into a pre-trained semantic classification model, and the knowledge domain category to which it belongs is output; the flag semantic concept is temporarily attached as a new leaf node under the knowledge domain category to which it belongs.
[0053] Specifically, after calculating the vector similarity, for each fragment semantic concept, the vector similarity with all entity nodes in the knowledge graph is checked against a preset threshold. If the vector similarity of any entity node in the domain knowledge graph is less than the preset threshold, the first fragment semantic concept is marked as a link failure, indicating that it has no directly or highly similar matching entity nodes in the existing knowledge graph, and may be an emerging term, a user-created expression, or an uncovered domain. The link failure flag is used to distinguish between matched and unmatched fragment semantic concepts. Semantic concepts with link failure flags are extracted and input into a pre-trained semantic classification model for knowledge domain category analysis. The semantic classification model is a deep learning-based text classification architecture, such as the BiLSTM-Attention model. By encoding the semantic concept vector with context and training it with labeled domain corpora, it outputs the knowledge domain category to which the semantic probability belongs, achieving intelligent classification of new concepts. Even if there is no corresponding node in the knowledge graph, the domain to which the semantic probability belongs can be clearly identified at the semantic level.
[0054] After obtaining the knowledge domain category, the identifier semantic concept is temporarily attached as a new leaf node under the knowledge domain category. The temporary attachment refers to establishing a temporary hierarchical relationship between the identifier semantic probabilities in the tree structure of the knowledge graph. This enables intelligent identification and inclusion of new semantic concepts or uncovered semantic concepts, ensuring the integrity of semantic coverage, preventing information loss, and improving the comprehensiveness and accuracy of hierarchical tags.
[0055] Furthermore, after temporarily attaching the identified semantic concept as a new leaf node under its knowledge domain category, the process further includes: submitting the temporarily attached new semantic concept and its association with the knowledge domain node to the review queue; and updating the new semantic concept and its association to the domain knowledge graph after receiving the administrator's approval instruction.
[0056] Specifically, when a new semantic concept is identified and temporarily attached to the knowledge graph as a new leaf node, the association between the new leaf node and the knowledge domain node is automatically submitted to the review queue. The review queue is a list of tasks awaiting manual review, used to verify the semantic accuracy, domain affiliation, and association of the new concept. Each record includes the concept name, source text, semantic vector representation, the knowledge domain category, and the temporary attachment association, providing administrators with comprehensive reference. For example, identifying the new concept "Super Fireball" and temporarily attaching it to the skill category, this information is submitted to the review queue. Administrators view the new semantic probabilities and their associations in the queue, and based on professional knowledge and contextual information, determine whether the new semantic concept should be formally included in the knowledge graph. Furthermore, administrators can check whether "Super Fireball" is indeed a valid skill name and whether it belongs to a skill category. When the administrator confirms that the new semantic concept and its associations are accurate, a review approval instruction is issued; for example, after confirming that "Super Fireball" is a valid skill name and that the term is categorized as a skill, an approval instruction is issued. Upon receiving the administrator's approval instruction, the system automatically updates the new semantic concepts and their relationships into the domain knowledge graph, updating the tree structure and node indexes, and simultaneously updating the semantic vector indexes for subsequent retrieval and tag generation. For example, "Super Fireball" is officially added to the skill category of the domain knowledge graph, and related indexes and links are updated.
[0057] Through review and formal update mechanisms, a safe and reliable process has been achieved for new semantic concepts to be permanently attached to the knowledge graph, improving the comprehensiveness and accuracy of the knowledge graph, and thus enhancing the integrity and reliability of hierarchical semantic tag generation. This provides stable and effective content tags for content analysis, competitor comparison, and creative optimization.
[0058] Example 2: Based on the same inventive concept as the hierarchical semantic processing tag automatic generation method for web content in Example 1, this application also provides a hierarchical semantic processing tag automatic generation system for web content. Please refer to the appendix. Figure 2 The hierarchical semantic processing tag automatic generation system for web content includes:
[0059] The request parsing module 11 is used to receive semantic analysis requests from target users and parse them to obtain the first work and the second work; the data acquisition module 12 is used to extract core entities and perform entity derivation analysis on the first work and the second work respectively, establish retrieval vectors, and collect the first network content set corresponding to the first work and the second network content set corresponding to the second work from network data sources; the semantic tag construction module 13 is used to extract key semantic concepts and associated sentiment polarities from the first network content set and the second network content set respectively, link them with pre-built domain knowledge graphs, and construct a first hierarchical semantic tag set and a second hierarchical semantic tag set; the report generation module 14 is used to compare the first hierarchical semantic tag set and the second hierarchical semantic tag set, perform tag alignment classification and semantic aggregation on knowledge domain and sentiment polarity, and generate a hierarchical comparison report.
[0060] Furthermore, the request parsing module 11 is also used to: the first work is a work operated by the target user on the network, and the second work is a competing work of the first work.
[0061] Furthermore, the data acquisition module 12 is also used to: extract the work name, author, and main character name from the metadata of the first work and the second work respectively as the first core entity and the second core entity; use the first core entity and the second core entity as query terms to retrieve and crawl the first derivative text content set and the second derivative text content set from a preset online community; perform word frequency co-occurrence statistics on the first derivative text content set and the second derivative text content set, and extract the first derivative words that frequently co-occur with the first core entity and the second derivative words that frequently co-occur with the second core entity; establish a first retrieval vector based on the first core entity and the first derivative words, and establish a second retrieval vector based on the second core entity and the second derivative words.
[0062] Furthermore, the data acquisition module 12 is also used to: include at least derivative terms related to characters, skills, exclusive items, and core plots in the derivative vocabulary.
[0063] Furthermore, the semantic tag construction module 13 is also used for: Step 1: performing sentence-level segmentation and semantic vectorization representation on each content fragment in the first network content set to generate several first fragment semantic concepts; Step 2: linking the several first fragment semantic concepts to the corresponding entity nodes of the domain knowledge graph; Step 3: organizing the linked semantic concepts according to their categories based on the tree structure of the domain knowledge graph, and associating them with the corresponding sentiment polarity in the original content fragments to generate a first hierarchical semantic tag set with knowledge domains as branches and semantic concepts as leaves; and performing steps 1 to 3 on the second network content set to generate a second hierarchical semantic tag set.
[0064] Furthermore, the semantic tag construction module 13 is also used to: calculate the vector similarity between the plurality of first fragment semantic concepts and entity nodes in the domain knowledge graph; and, based on the vector similarity, directly link the first fragment semantic concepts whose similarity exceeds a preset threshold to the corresponding entity nodes in the domain knowledge graph.
[0065] Furthermore, the semantic label construction module 13 is also used to: if there is a first fragment semantic concept whose vector similarity to any entity node in the domain knowledge graph is less than the preset threshold, mark the link failure; extract the label semantic concept with the link failure mark, input it into the pre-trained semantic classification model, and output the knowledge domain category to which it belongs; and temporarily attach the label semantic concept as a new leaf node under the knowledge domain category to which it belongs.
[0066] Furthermore, the report generation module 14 is also used to: traverse the first hierarchical semantic tag set and the second hierarchical semantic tag set, match them according to the knowledge domain nodes they belong to, and group tags with similar attributes into the same comparison group; within each comparison group, count the number of semantic concepts and the distribution of emotional polarity belonging to the first work and the second work respectively, calculate the positive voice volume ratio and discussion popularity index of the first work in the same knowledge domain, and make a differential comparison with the corresponding indicators of the second work; based on the differential comparison results, automatically identify the advantageous and disadvantageous areas of the first work relative to the second work, and construct the hierarchical comparison report.
[0067] Furthermore, the semantic tag construction module 13 is also used to: submit the temporarily attached new semantic concepts and their associations with knowledge domain nodes to the review queue; and after receiving the administrator's approval instruction, update the new semantic concepts and their associations to the domain knowledge graph.
[0068] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Figure 1The method and specific examples for automatically generating hierarchical semantic processing tags for web content in Embodiment 1 are also applicable to the automatic generation system for hierarchical semantic processing tags for web content in this embodiment. Through the foregoing detailed description of the method for automatically generating hierarchical semantic processing tags for web content, those skilled in the art can clearly understand the automatic generation system for hierarchical semantic processing tags for web content in this embodiment. Therefore, for the sake of brevity, it will not be described in detail here. As for the automatic generation system for hierarchical semantic processing tags for web content disclosed in the embodiments, since it corresponds to the method for automatically generating hierarchical semantic processing tags for web content disclosed in the embodiments, the description is relatively simple; relevant details can be found in the method section.
[0069] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0070] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for automatic generation of hierarchical semantic processing tags for network content, characterized in that, include: Receive semantic analysis requests from target users, and parse them to obtain the first and second works; The core entities of the first work and the second work are extracted and entity derivation analysis is performed respectively to establish retrieval vectors. The first network content set corresponding to the first work and the second network content set corresponding to the second work are collected from network data sources. Key semantic concepts and associated sentiment polarities are extracted from the first network content set and the second network content set, respectively, and entity links are made with the pre-constructed domain knowledge graph to construct a first-level semantic tag set and a second-level semantic tag set. By comparing the first hierarchical semantic tag set with the second hierarchical semantic tag set, tag alignment classification and semantic aggregation are performed on knowledge domain and sentiment polarity to generate a hierarchical comparison report; Key semantic concepts and associated sentiment polarities are extracted from the first and second network content sets, respectively, and entity links are established with pre-constructed domain knowledge graphs to construct a first-level semantic tag set and a second-level semantic tag set, including: Step 1: Perform sentence-level segmentation and semantic vectorization representation on each content fragment in the first network content set to generate several first fragment semantic concepts; Step 2: Link the several first fragment semantic concepts to the corresponding entity nodes of the domain knowledge graph; Step 3: Based on the tree structure of the domain knowledge graph, organize the linked semantic concepts according to their categories and associate them with the corresponding sentiment polarities in the original content fragments to generate the first hierarchical semantic tag set with the knowledge domain as the branches and the semantic concepts as the leaves. Steps one through three are performed on the second network content set to generate the second hierarchical semantic tag set.
2. The method of claim 1, wherein the network content-oriented hierarchical semantic processing tag auto-generation method is characterized by, The first work is a work operated online by the target user, and the second work is a competing work of the first work.
3. The method of claim 1, wherein the network content-oriented hierarchical semantic processing tag auto-generation method is characterized by, The first and second works are subjected to core entity extraction and entity derivation analysis respectively, and retrieval vectors are established, including: For the first work and the second work, the work name, author, and main character name are extracted from the metadata as the first core entity and the second core entity, respectively. Using the first core entity and the second core entity as query terms, retrieve and crawl the first and second derivative text content sets from the preset online communities; Perform word frequency co-occurrence statistics on the first and second derived text content sets to extract the first derived words that frequently co-occur with the first core entity and the second derived words that frequently co-occur with the second core entity. A first retrieval vector is established based on the first core entity and the first derived vocabulary, and a second retrieval vector is established based on the second core entity and the second derived vocabulary.
4. The method of claim 3, wherein the network content-oriented hierarchical semantic processing tag auto-generation method is characterized by, Derivative vocabulary includes at least terms related to characters, skills, unique items, and core plot points.
5. The method of claim 1, wherein the method further comprises: generating a hierarchical semantic processing tag for each of the plurality of network content items based on the hierarchical semantic processing tag generation rule set. Linking the aforementioned first fragment semantic concepts to the corresponding entity nodes of the domain knowledge graph includes: Calculate the vector similarity between the semantic concepts of the aforementioned first fragments and the entity nodes in the domain knowledge graph; Based on vector similarity, for the first fragment semantic concept whose similarity exceeds a preset threshold, it is directly linked to the corresponding entity node of the domain knowledge graph.
6. The method for automatically generating hierarchical semantic processing tags for web content as described in claim 5, characterized in that, After calculating the vector similarity between the semantic concepts of the several first fragments and the entity nodes in the domain knowledge graph, the method further includes: If there exists a first fragment semantic concept whose vector similarity to any entity node in the domain knowledge graph is less than the preset threshold, a link failure is identified. Extract the semantic concepts of the identifiers with link failure indicators, input them into a pre-trained semantic classification model, and output the knowledge domain category to which they belong; The semantic concept of the identifier is temporarily attached as a new leaf node under the knowledge domain category.
7. The method for automatically generating hierarchical semantic processing tags for web content as described in claim 1, characterized in that, By comparing the first hierarchical semantic tag set with the second hierarchical semantic tag set, tag alignment classification and semantic aggregation based on knowledge domain and sentiment polarity are performed to generate a hierarchical comparison report, including: Traverse the first hierarchical semantic tag set and the second hierarchical semantic tag set, match them according to the knowledge domain nodes, and group tags with the same attributes into the same comparison group; Within each comparison group, the number of semantic concepts and the distribution of emotional polarity belonging to the first and second works were counted respectively. The positive voice volume ratio and discussion popularity index of the first work in the same knowledge field were calculated and compared with the corresponding indicators of the second work. Based on the results of the differential comparison, the advantages and disadvantages of the first work compared to the second work are automatically identified, and the hierarchical comparison report is constructed.
8. The hierarchical semantic processing tag automatic generation method for web content as described in claim 6, characterized in that, After temporarily attaching the semantic concept of the identifier as a new leaf node under its knowledge domain category, the process also includes: Submit the newly attached semantic concepts and their associations with knowledge domain nodes to the review queue; After receiving the administrator's approval instruction, the new semantic concepts and their related relationships are updated in the domain knowledge graph.
9. A hierarchical semantic processing tag automatic generation system for web content, characterized in that, The steps for implementing the hierarchical semantic processing tag automatic generation method for web content according to any one of claims 1 to 8, wherein the hierarchical semantic processing tag automatic generation system for web content comprises: The request parsing module is used to receive semantic analysis requests from target users and parse them to obtain the first and second works. The data acquisition module is used to extract the core entities and perform entity derivation analysis on the first work and the second work respectively, establish retrieval vectors, and collect the first network content set corresponding to the first work and the second network content set corresponding to the second work from network data sources. The semantic tag construction module is used to extract key semantic concepts and associated sentiment polarities from the first network content set and the second network content set respectively, link them with the pre-built domain knowledge graph, and construct a first-level semantic tag set and a second-level semantic tag set. The report generation module is used to compare the first hierarchical semantic tag set with the second hierarchical semantic tag set, perform tag alignment classification and semantic aggregation based on knowledge domain and sentiment polarity, and generate a hierarchical comparison report.