Intelligent advertisement material optimization system and method based on knowledge extraction and knowledge fusion

By using multimodal semantic parsing and knowledge fusion technology, the problem of insufficient semantic relationship modeling in advertising creative optimization is solved, achieving semantic consistency and style adaptability in advertising content optimization, thereby improving advertising effectiveness.

CN121860702AInactive Publication Date: 2026-04-14JIANGSU SHENJIANG BOHUI TECHNOLOGY DEVELOPMENT CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing advertising creative optimization methods struggle to model deep semantic relationships and cross-modal association structures in multimodal content, lacking context awareness capabilities. This results in rigid optimization results that cannot adapt to the advertising needs of different styles and scenarios.

Method used

Employing multimodal semantic parsing, coreference resolution, semantic structure modeling, graph embedding fusion, and sentence reconstruction technologies, and through joint semantic parsing, syntactic dependency analysis, entity recognition, and knowledge fusion, an advertising creative optimization system with strong semantic consistency and natural style expression is constructed.

Benefits of technology

It significantly improves the semantic consistency and information completeness of advertising content, enhances the logical rationality and accuracy of advertising dissemination, and possesses high controllability and style adaptability, making it suitable for multi-channel and multi-style advertising.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121860702A_ABST
    Figure CN121860702A_ABST
Patent Text Reader

Abstract

The invention discloses an advertisement material intelligent optimization system and method based on knowledge extraction and knowledge fusion, and the method comprises the following steps: S1, obtaining and preprocessing a to-be-optimized advertisement material, and collecting an external knowledge source; s2, combined semantic analysis is executed, and cross-modal semantic alignment is carried out; s3, extracting entity units and corresponding semantic attributes, performing type labeling on the entity units, and establishing a semantic relation structure; s4, performing semantic fusion with an external knowledge source, calculating similarity between semantic vectors, performing entity mapping, performing normalization, and correcting conflict boundaries and attribute value divergence; s5, extracting a target expression fragment, and performing optimization according to a preset optimization rule set; and S6, verifying and detecting the optimized advertisement material, and outputting an optimization result. According to the method, deep understanding and cross-modal fusion optimization of advertisement materials can be realized, and the accuracy, attraction and propagation effect of advertisement semantic expression are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent optimization technology for advertising materials, and in particular to an intelligent optimization system and method for advertising materials based on knowledge extraction and knowledge fusion. Background Technology

[0002] With the increasing sophistication and personalization of internet advertising, the automatic optimization of ad creatives has become a crucial step in improving campaign effectiveness and click-through rates. Existing ad creatives primarily include multimodal content such as text, images, and videos. Information within different types of creatives often exists in unstructured or weakly structured forms, lacking a unified semantic expression and content organization. This means that existing ad optimization methods often only perform superficial content replacement or style changes for a single modality, making it difficult to achieve precise control over the overall semantic intent and content reconstruction.

[0003] Some studies have attempted to incorporate natural language processing and computer vision technologies to identify elements and extract preliminary structures from advertising materials based on keyword extraction and image object detection. However, most of these studies remain at the stage of shallow feature analysis, failing to fully utilize the deep semantic relationships and cross-modal association structures behind the advertising slogans. Furthermore, some methods use rule bases or simple classification models to achieve templated content replacement, lacking context awareness and understanding capabilities. This results in rigid optimization results, a poor user experience, and an inability to adapt to the advertising needs of different styles and scenarios.

[0004] In the field of multimodal fusion, existing technologies mainly rely on attention mechanisms for semantic alignment. However, these methods suffer from high model complexity and strong dependence on training data, making it difficult to generalize to real-world advertising corpora. Furthermore, the inability to effectively model referential consistency and semantic coherence among multimodal expressions often leads to semantic conflicts and style misalignments during content rearrangement and optimization.

[0005] In terms of knowledge acquisition, current ad optimization has not fully explored the potential of external knowledge sources to improve the accuracy and richness of expression. Traditional methods mostly employ static tags or keyword mapping, lacking a deep integration mechanism with structured knowledge resources such as knowledge graphs, making it difficult to effectively enhance ad content at the semantic level. Furthermore, there is a lack of systematic modeling of semantic relationships and contextual dependencies between entities, resulting in a lack of logical consistency and semantic integrity in the final generated content.

[0006] Especially in terms of advertising slogan optimization, existing sentence transformation methods still mainly rely on template replacement, lacking fine-grained control over syntactic structure, semantic attachment, and style mapping, and are unable to make personalized optimization adjustments for different audience groups and advertising scenarios.

[0007] Therefore, how to provide an intelligent optimization system and method for advertising materials based on knowledge extraction and knowledge fusion is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0008] One objective of this invention is to propose an intelligent optimization system and method for advertising materials based on knowledge extraction and knowledge fusion. This invention fully integrates technologies such as multimodal semantic parsing, coreference resolution, semantic structure modeling, graph embedding fusion, and sentence reconstruction, and systematically realizes the collaborative processing of advertising content in each stage of semantic understanding, knowledge completion, and expression optimization. It has the advantages of strong semantic consistency of optimization results, natural style expression, and strong cross-modal understanding ability.

[0009] The intelligent optimization method for advertising creatives based on knowledge extraction and knowledge fusion according to embodiments of the present invention includes the following steps: S1. Obtain the advertising materials to be optimized and preprocess them to construct a multimodal feature set, while collecting external knowledge sources; S2. Perform joint semantic parsing on the multimodal feature set and use coreference resolution to perform cross-modal semantic alignment to form a semantic structure sequence; S3. Perform syntactic dependency analysis and semantic role labeling on the semantic structure sequence, extract entity units and corresponding semantic attributes, use entity recognition methods to label the entity units by type, and establish the semantic relationship structure between entities based on contextual dependency relationships. S4. Semantically fuse the semantic relation structure with the entity nodes and relation edges in the external knowledge source, calculate the similarity between semantic vectors based on graph embedding representation and perform entity mapping, normalize the node pairs with large semantic deviations, and correct conflict boundaries and attribute value discrepancies. S5. Based on the fused semantic structure, extract the target expression fragments from the advertising materials, and perform semantic content replacement, expression structure adjustment and style mapping according to the preset optimization rule set; S6. Perform semantic consistency verification and expression integrity detection on the optimized advertising materials, and output the optimization results.

[0010] Optionally, the advertising material includes text content, image content, and video frame sequence. The preprocessing includes performing language standardization, special symbol removal, and paragraph segmentation on the text content; performing resolution normalization, redundant background removal, and main area annotation on the image content; and performing keyframe extraction and time segmentation on the video frame sequence.

[0011] Optionally, the external knowledge sources include advertising industry knowledge graphs, user interest tag libraries, and historical campaign feedback datasets, specifically including: The advertising industry knowledge graph is constructed based on structured advertising data and industry public corpus. It uses triples to represent entities and corresponding semantic relationships. The entities include brands, product categories, marketing scenarios and performance styles. The semantic relationships include attribute relationships, subordinate relationships and style association relationships. The user interest tag library consists of a set of user browsing behavior, click behavior and conversion behavior. User behavior is tagged and modeled using a clustering algorithm, and a many-to-many mapping relationship between user identifiers and interest tags is established. The historical campaign feedback dataset is extracted from the records of each ad creative campaign, recording the creative identifier, campaign channel, target audience characteristics, and performance metrics, including click-through rate, conversion rate, and dwell time.

[0012] Optionally, S2 specifically includes: S21. Perform joint semantic parsing on the multimodal feature set, specifically including: For the text content, a multi-layer semantic encoding method based on the BERT model is used to extract semantic representations that include lexical structure, contextual dependencies and inter-sentence relations; For the image content, the ResNet-50 network is used to perform object detection, and the detected area is segmented and feature embedded to obtain local visual semantic vectors; For video frames, the I3D network is used to perform 3D convolutional action modeling to extract behavioral semantic vectors and event label representations from temporally continuous frames; Semantic vector sets for text, image, and video modalities are obtained respectively; S22. Based on the timestamps, location information, bounding box indexes and semantic labels contained in the semantic representation, perform cross-modal semantic alignment processing, establish multi-dimensional mapping relationships for semantic segments with referential relationships, and mark potential core-referenced segments; S23. A coreference resolution method based on central entity constraints is adopted to perform semantic discrimination and reference classification on the marked fragments, fuse semantic units with coreference relations, perform vector merging and subject label replacement operations, and eliminate semantic conflicts and redundant expressions. S24. Rearrange the processed semantic units according to modal independence order and semantic dependency strength, and combine them to generate a semantic structure sequence.

[0013] Optionally, S3 specifically includes: S31. Perform syntactic dependency analysis on text segments in semantic structure sequence, extract subject-predicate, modification and parallel dependency paths between words, control semantic grouping with dependency edge weights, and perform structural segmentation on long sentences with semantic span exceeding a set threshold. S32. Perform semantic role labeling within the structural segmentation unit, establish semantic mapping relationships with action verbs as the core, and extract the role correspondence between verbs and their semantic objects through a limited context window, and label semantic roles and structural position indices. S33. Using the semantic role labeling results and syntactic boundaries as joint inputs, perform entity recognition on candidate phrases, exclude boundary overlaps and semantically repetitive segments, use BiLSTM-CRF to label entity types, and form a non-overlapping entity sequence. S34. Perform dependency path backtracking on entity pairs in the entity sequence, calculate the difference between path depth and word order, filter entity pairs that satisfy contextual dependency constraints, and construct semantic connections with related directions based on semantic role relationships, specifically including: Syntactic dependency path backtracking is performed on entity pairs in non-overlapping entity sequences. A limited breadth-first traversal algorithm is used to extract the shortest syntactic path between entity pairs and record the verb, conjunction and modifier nodes in the path. The word nodes in the backtracking path are weighted by path depth and word order distance, and the context relationship score between entity pairs is calculated. Semantic connection candidate pairs are set according to the score. Perform semantic role consistency judgment on candidate entity pairs, filter out entity pairs with semantic category conflicts or duplicate semantic roles, and set the direction of the connection edge according to semantic directionality; Weighted directional edges are generated from the filtering results, and the semantic relationships between entity pairs are represented as structured connection information including role type, context label and connection confidence, which is then added to the semantic relationship structure graph.

[0014] Optionally, S33 specifically includes: S331. Align the semantic role labeling results with the syntactic boundary, construct a feature tensor containing role labels, part-of-speech tags, boundary positions and context window indices, and composite input entity recognition model; S332. Perform boundary overlap detection on the candidate phrase set, and adopt a dual threshold strategy of maximum overlap rate and label consistency to remove redundant segments with overlap rate higher than the set threshold or label conflict to form a preliminary entity set. S333. In the BiLSTM encoding process, a multi-channel input structure of word vector embedding, position embedding and role tag embedding is introduced to output a context-aware representation vector sequence. S334. Input the representation vector sequence into the CRF layer, calculate the optimal label path based on the state transition probability and observation probability, and generate the label sequence. S335. Recover entity intervals based on entity boundary markers in the label sequence, filter preferred entities by word order, and establish a non-overlapping entity index sequence.

[0015] Optionally, S4 specifically includes: S41. Retrieve the semantic counterpart nodes of entity nodes in the semantic relation structure in the external knowledge source, extract the label, attribute fields and related edge information, calculate the structural similarity between nodes using SimRank, and filter node pairs with consistent semantic labels or embedding vector similarity higher than the threshold. S42. Perform structural-level fusion processing on the selected node pairs, perform field alignment and missing data completion operations, set conflict markers for fields with inconsistent attribute values, and mark the original relation weights of conflicting edges. S43. The fused structure is vector-encoded using a graph embedding method based on GraphSAGE, and the contextual feature representation of each node is extracted. The semantic similarity between entity nodes is then recalculated in a unified vector space. S44. Perform attribute vector normalization and weight reset on entity nodes with similarity below the set threshold, adjust the main value source of conflicting labels according to the preset fusion rules, and correct the attribute divergence of relation edges caused by semantic conflicts.

[0016] Optionally, S5 specifically includes: S51. Locate the core entity nodes and behavior predicate nodes on the main semantic path in the fused semantic structure, and filter the start and end boundaries of the target expression fragment based on path depth and semantic association weight. S52. Perform semantic consistency screening on the selected segments, remove candidate segments with structural ambiguity, semantic breaks or expression conflicts, and generate a sequence of target expression segments according to word order. S53. Perform semantic mapping on the content words in the target expression segment, perform equivalent word replacement based on WordNet vocabulary hierarchy and sentiment dictionary rules, and call the style vocabulary to adjust the sentiment tendency and tone expression; S54. Based on the sentence structure templates in the preset rule set, the expression structure of the expression fragment is rearranged, and the word order, modification, dependency and syntax patterns are adjusted to enhance the content appeal and style consistency while maintaining the semantics. S55. Embed the replaced expression fragment into the corresponding context position in the original material to construct the optimized semantic expression content of the advertisement.

[0017] Optionally, the preset rule set includes sentence structure constraint rules, word order rearrangement pattern library and style adaptation parameter set. The rule set is used to constrain the range of structural transformation of expression fragments and the conditions for rearrangement order selection. The sentence structure template contains multiple structural styles built based on semantic scenarios, specifically including declarative, emphatic, interrogative, inverted, and rhetorical sentence structures. Each template defines fixed word order rules and variable component adjustment ranges corresponding to semantic roles, and uses a structural index mapping method to mark the positional boundaries of the subject, predicate, object, and modifier. The expression structure rearrangement process specifically includes: Based on the semantic tags of entities and verbs in the current segment, match the corresponding sentence structure templates and identify the syntactic nodes in the segment that need to be adjusted; The positions of modifiers and predicates are sorted and rearranged according to template rules, and the structural similarity scores of the candidate permutation paths are labeled. The Viterbi algorithm is used to decode and compute all candidate paths, and the word order arrangement that maintains semantic consistency and has the best structural score is selected. After the rearrangement is completed, modal particles, conjunctions, and adverbs are inserted or replaced according to the style adaptation parameters.

[0018] An intelligent optimization system for advertising creatives based on knowledge extraction and knowledge fusion, according to an embodiment of the present invention, includes: The data preprocessing module is used to acquire and preprocess the advertising materials to be optimized, construct a multimodal feature set, and collect external knowledge sources; The joint semantic parsing module performs joint semantic parsing on the multimodal feature set and uses coreference resolution to perform cross-modal semantic alignment, forming a semantic structure sequence; The knowledge extraction module is used to perform syntactic dependency analysis and semantic role labeling on semantic structure sequences, extract entity units and corresponding semantic attributes, use entity recognition methods to label entity units by type, and establish semantic relationship structure between entities based on contextual dependency relationships. The knowledge fusion module performs semantic fusion with the semantic relation structure and entity nodes and relation edges in the external knowledge source. It calculates the similarity between semantic vectors based on graph embedding representation and performs entity mapping. It normalizes node pairs with large semantic deviations and corrects conflict boundaries and attribute value discrepancies. The expression optimization module is used to extract target expression fragments from advertising materials based on the fused semantic structure, and perform semantic content replacement, expression structure adjustment and style mapping according to the preset optimization rule set; The expression verification module performs semantic consistency verification and expression integrity detection on the optimized advertising creatives and outputs the optimization results.

[0019] The beneficial effects of this invention are: First, by performing joint semantic parsing and cross-modal coreference resolution on multimodal information such as text, images, and video frames in advertising materials, this invention effectively breaks down semantic barriers between different modalities, significantly improves the semantic consistency and information completeness among various elements in advertising content, and solves problems such as modality separation, missing context, and semantic mismatch in traditional methods, providing a high-quality semantic structure foundation for subsequent optimization.

[0020] Secondly, this invention introduces structured knowledge graphs as external knowledge sources and combines them with structure and representation learning methods such as SimRank and GraphSAGE to achieve deep integration of semantic relation structures and knowledge graph entities. This enables accurate identification of key entities and semantic gaps in content, and automatic supplementation of missing attribute information, contextual elements, and scene knowledge in advertising copy, thereby significantly improving the logical rationality and accuracy of advertising content.

[0021] Finally, this invention constructs a multi-level expression optimization mechanism. Based on semantic path filtering, style mapping, word meaning replacement and sentence reconstruction, it performs language optimization processing on advertising expression fragments while preserving semantics. This not only enhances the emotional appeal and user acceptance of advertising slogans, but also has high controllability and style adaptability. It can be widely applied to advertising needs across multiple channels and styles, and has strong practical value and promotion prospects. Attached Figure Description

[0022] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0023] Figure 1 This is an overall flowchart of the intelligent optimization method for advertising materials based on knowledge extraction and knowledge fusion proposed in this invention; Figure 2 This is a schematic diagram illustrating the multimodal semantic parsing and coreference resolution of the intelligent optimization method for advertising materials based on knowledge extraction and knowledge fusion proposed in this invention; Figure 3 This is a module structure diagram of the intelligent optimization system for advertising materials based on knowledge extraction and knowledge fusion proposed in this invention. Detailed Implementation

[0024] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0025] refer to Figure 1-2 The intelligent optimization method for advertising creatives based on knowledge extraction and knowledge fusion includes the following steps: S1. Obtain the advertising materials to be optimized and preprocess them to construct a multimodal feature set, while collecting external knowledge sources; S2. Perform joint semantic parsing on the multimodal feature set and use coreference resolution to perform cross-modal semantic alignment to form a semantic structure sequence; S3. Perform syntactic dependency analysis and semantic role labeling on the semantic structure sequence, extract entity units and corresponding semantic attributes, use entity recognition methods to label the entity units by type, and establish the semantic relationship structure between entities based on contextual dependency relationships. S4. Semantically fuse the semantic relation structure with the entity nodes and relation edges in the external knowledge source, calculate the similarity between semantic vectors based on graph embedding representation and perform entity mapping, normalize the node pairs with large semantic deviations, and correct conflict boundaries and attribute value discrepancies. S5. Based on the fused semantic structure, extract the target expression fragments from the advertising materials, and perform semantic content replacement, expression structure adjustment and style mapping according to the preset optimization rule set; S6. Perform semantic consistency verification and expression integrity detection on the optimized advertising materials, and output the optimization results.

[0026] In this embodiment, the advertising material includes text content, image content, and video frame sequence. The preprocessing includes performing language standardization, special symbol removal, and paragraph segmentation on the text content; performing resolution normalization, redundant background removal, and main area annotation on the image content; and performing keyframe extraction and time segmentation on the video frame sequence.

[0027] In this embodiment, the external knowledge sources include advertising industry knowledge graphs, user interest tag libraries, and historical campaign feedback datasets, specifically including: The advertising industry knowledge graph is constructed based on structured advertising data and industry public corpus. It uses triples to represent entities and corresponding semantic relationships. The entities include brands, product categories, marketing scenarios and performance styles. The semantic relationships include attribute relationships, subordinate relationships and style association relationships. The user interest tag library consists of a set of user browsing behavior, click behavior and conversion behavior. User behavior is tagged and modeled using a clustering algorithm, and a many-to-many mapping relationship between user identifiers and interest tags is established. The historical campaign feedback dataset is extracted from the records of each ad creative campaign, recording the creative identifier, campaign channel, target audience characteristics, and performance metrics, including click-through rate, conversion rate, and dwell time.

[0028] In this embodiment, S2 specifically includes: S21. Perform joint semantic parsing on the multimodal feature set, specifically including: For the text content, a multi-layer semantic encoding method based on the BERT model is used to extract semantic representations that include lexical structure, contextual dependencies and inter-sentence relations; For the image content, the ResNet-50 network is used to perform object detection, and the detected area is segmented and feature embedded to obtain local visual semantic vectors; For video frames, the I3D network is used to perform 3D convolutional action modeling to extract behavioral semantic vectors and event label representations from temporally continuous frames; Semantic vector sets for text, image, and video modalities are obtained respectively; S22. Based on the timestamps, location information, bounding box indexes and semantic labels contained in the semantic representation, perform cross-modal semantic alignment processing, establish multi-dimensional mapping relationships for semantic segments with referential relationships, and mark potential core-referenced segments; S23. A coreference resolution method based on central entity constraints is adopted to perform semantic discrimination and reference classification on the marked fragments, fuse semantic units with coreference relations, perform vector merging and subject label replacement operations, and eliminate semantic conflicts and redundant expressions. The coreference resolution method adopts a strategy based on central entity constraints, identifies the core semantic entity in the semantic structure sequence, calculates the centrality score based on the frequency of occurrence, semantic weight and contextual reference density of the entity in the multimodal corpus, and determines the central reference object. Using the central entity as the classification anchor point, cosine similarity calculation is performed on the semantic representation between the coreference candidate unit and the central entity, and segments with similarity higher than the threshold are selected as the belonging targets. For semantic units belonging to the same central entity, vector-level semantic fusion processing is performed, and their contextual semantic representations are merged by weighted averaging. The subject label is selected as the referential identifier based on information entropy. Repeated references in the original position are replaced with tags, and reference indices in the structural sequence are updated; To address semantic conflicts that arise during the attribution process, a referential consistency matrix is ​​constructed to calculate the degree of semantic deviation. When the conflict intensity exceeds a set threshold, a conflict divergence correction operation is performed, which calls a fuzzy matching dictionary to replace the conflicting subject. S24. Rearrange the processed semantic units according to modal independence order and semantic dependency strength, and combine them to generate a semantic structure sequence.

[0029] In this embodiment, S3 specifically includes: S31. Perform syntactic dependency analysis on text segments in semantic structure sequence, extract subject-predicate, modification and parallel dependency paths between words, control semantic grouping with dependency edge weights, and perform structural segmentation on long sentences with semantic span exceeding a set threshold. S32. Perform semantic role labeling within the structural segmentation unit, establish semantic mapping relationships with action verbs as the core, and extract the role correspondence between verbs and their semantic objects through a limited context window, and label semantic roles and structural position indices. S33. Using the semantic role labeling results and syntactic boundaries as joint inputs, perform entity recognition on candidate phrases, exclude boundary overlaps and semantically repetitive segments, use BiLSTM-CRF to label entity types, and form a non-overlapping entity sequence. S34. Perform dependency path backtracking on entity pairs in the entity sequence, calculate the difference between path depth and word order, filter entity pairs that satisfy contextual dependency constraints, and construct semantic connections with related directions based on semantic role relationships, specifically including: Syntactic dependency path backtracking is performed on entity pairs in non-overlapping entity sequences. A limited breadth-first traversal algorithm is used to extract the shortest syntactic path between entity pairs and record the verb, conjunction and modifier nodes in the path. The word nodes in the backtracking path are weighted by path depth and word order distance, and the context relationship score between entity pairs is calculated. Semantic connection candidate pairs are set according to the score. Perform semantic role consistency judgment on candidate entity pairs, filter out entity pairs with semantic category conflicts or duplicate semantic roles, and set the direction of the connection edge according to semantic directionality; Weighted directional edges are generated from the filtering results, and the semantic relationships between entity pairs are represented as structured connection information including role type, context label and connection confidence, which is then added to the semantic relationship structure graph.

[0030] In this embodiment, S33 specifically includes: S331. Align the semantic role labeling results with the syntactic boundary, construct a feature tensor containing role labels, part-of-speech tags, boundary positions and context window indices, and composite input entity recognition model; S332. Perform boundary overlap detection on the candidate phrase set, using a dual threshold strategy of maximum overlap rate and label consistency to remove redundant segments with overlap rates exceeding the set threshold or with conflicting labels, forming a preliminary entity set. The boundary overlap detection specifically includes: For any two segments in the candidate phrase set, perform cross-detection of start and end boundary indices, calculate the ratio between their overlap length and the shortest length of the two segments, and obtain the boundary overlap rate. When the boundary overlap rate is higher than the first set threshold, a label consistency determination is performed, comparing the entity type labels and semantic role annotations of the two segments; When entity type labels are different, or semantic roles have conflicting relationships, the fragment pair is marked as a label conflict type redundant fragment; If the boundary overlap rate is higher than the second set threshold and the two fragment entities are of the same type, then the fragment pair is marked as a structurally overlapping redundant fragment. For phrase fragment pairs marked as redundant, perform confidence ranking, prioritize retaining fragments with higher entity recognition confidence or higher semantic weight, and delete redundant fragments; S333. In the BiLSTM encoding process, a multi-channel input structure of word vector embedding, position embedding and role tag embedding is introduced to output a context-aware representation vector sequence. S334. Input the representation vector sequence into the CRF layer, calculate the optimal label path based on the state transition probability and observation probability, and generate the label sequence. S335. Recover entity intervals based on entity boundary markers in the label sequence, filter preferred entities by word order, and establish a non-overlapping entity index sequence.

[0031] In this embodiment, S4 specifically includes: S41. Retrieve the semantically corresponding nodes of entity nodes in the semantic relation structure in the external knowledge source, extract label, attribute fields and associated edge information, calculate the structural similarity between nodes using SimRank, and filter node pairs with consistent semantic labels or embedding vector similarity higher than a threshold, specifically including: For each entity node in the semantic relationship structure, perform semantic matching calls from external knowledge sources to extract a set of candidate entity nodes with the same name or synonyms, and extract their tag type, attribute field set, and established entity relationship edge information. For both entity nodes and candidate nodes, construct adjacency subgraphs centered on the nodes, and calculate their structural similarity in the graph structure based on the SimRank method. When the SimRank value is higher than the structural similarity threshold, a label consistency judgment is performed to compare whether there is a complete match or semantic equivalence relationship between the entity label fields. When the label fields are consistent, or the cosine similarity between the embedded vectors is higher than the vector similarity threshold, the candidate node is included in the set of matching node pairs. If any two of the three criteria of structural similarity, label consistency and embedding similarity are met, the corresponding node pair will be selected as a semantic fusion candidate pair and enter the subsequent structural fusion process. S42. Perform structural-level fusion processing on the selected node pairs, perform field alignment and missing data completion operations, set conflict markers for fields with inconsistent attribute values, and mark the original relation weights of conflicting edges. S43. The fused structure is vector-encoded using a graph embedding method based on GraphSAGE, and the contextual feature representation of each node is extracted. The semantic similarity between entity nodes is then recalculated in a unified vector space. S44. Perform attribute vector normalization and weight reset on entity nodes with similarity below the set threshold, adjust the main value source of conflicting labels according to the preset fusion rules, and correct the attribute divergence of relation edges caused by semantic conflicts.

[0032] In this embodiment, S5 specifically includes: S51. Locate the core entity nodes and behavior predicate nodes on the main semantic path in the fused semantic structure, and filter the start and end boundaries of the target expression fragment based on path depth and semantic association weight. S52. Perform semantic consistency screening on the selected segments, combine the importance score of the LexRank algorithm, remove candidate segments with structural ambiguity, semantic breaks or expression conflicts, and generate the target expression segment sequence according to the word order. S53. Perform semantic mapping on the content words in the target expression segment, perform equivalent word replacement based on WordNet vocabulary hierarchy and sentiment dictionary rules, and call the style vocabulary to adjust the sentiment tendency and tone expression; S54. Based on the sentence structure templates in the preset rule set, perform expression structure rearrangement on the expression fragments, and use the Viterbi algorithm to optimize the word order arrangement path, adjust the word order structure, modification dependency and syntactic pattern. S55. Embed the replaced expression fragment into the corresponding context position in the original material to construct the optimized semantic expression content of the advertisement.

[0033] In this embodiment, the preset rule set includes sentence structure constraint rules, word order rearrangement pattern library and style adaptation parameter set. The rule set is used to constrain the range of structural transformation of expression fragments and the conditions for rearrangement order selection. The sentence structure template contains multiple structural styles built based on semantic scenarios, specifically including declarative, emphatic, interrogative, inverted, and rhetorical sentence structures. Each template defines fixed word order rules and variable component adjustment ranges corresponding to semantic roles, and uses a structural index mapping method to mark the positional boundaries of the subject, predicate, object, and modifier. The expression structure rearrangement process specifically includes: Based on the semantic tags of entities and verbs in the current segment, match the corresponding sentence structure templates and identify the syntactic nodes in the segment that need to be adjusted; The positions of modifiers and predicates are sorted and rearranged according to template rules, and the structural similarity scores of the candidate permutation paths are labeled. The Viterbi algorithm is used to decode and compute all candidate paths, and the word order arrangement that maintains semantic consistency and has the best structural score is selected. After the rearrangement is completed, modal particles, conjunctions, and adverbs are inserted or replaced according to the style adaptation parameters.

[0034] refer to Figure 3 An intelligent advertising creative optimization system based on knowledge extraction and knowledge fusion includes: The data preprocessing module is used to acquire and preprocess the advertising materials to be optimized, construct a multimodal feature set, and collect external knowledge sources; The joint semantic parsing module performs joint semantic parsing on the multimodal feature set and uses coreference resolution to perform cross-modal semantic alignment, forming a semantic structure sequence; The knowledge extraction module is used to perform syntactic dependency analysis and semantic role labeling on semantic structure sequences, extract entity units and corresponding semantic attributes, use entity recognition methods to label entity units by type, and establish semantic relationship structure between entities based on contextual dependency relationships. The knowledge fusion module performs semantic fusion with the semantic relation structure and entity nodes and relation edges in the external knowledge source. It calculates the similarity between semantic vectors based on graph embedding representation and performs entity mapping. It normalizes node pairs with large semantic deviations and corrects conflict boundaries and attribute value discrepancies. The expression optimization module is used to extract target expression fragments from advertising materials based on the fused semantic structure, and perform semantic content replacement, expression structure adjustment and style mapping according to the preset optimization rule set; The expression verification module performs semantic consistency verification and expression integrity detection on the optimized advertising creatives and outputs the optimization results.

[0035] Example 1: To verify the feasibility of this invention in practice, it was applied to an advertising delivery system to test and evaluate its performance in the automatic optimization of multimodal advertising materials. The system includes an advertising operation platform, a content optimization module, an external knowledge fusion module, and an automatic generation interface. The materials cover industries such as food and beverage, digital electronics, and beauty and personal care, and include product description text, scene images, and video clips of people.

[0036] In practice, advertising operators upload a set of ad creatives to be optimized, including product descriptions, static images of user scenarios, and short video clips expressing the brand's philosophy. The original creatives have several issues before deployment, such as inconsistent text description styles, missing content elements, and a disconnect between text and images, affecting the overall performance of the ads.

[0037] The system first performs unified preprocessing on text, images, and video frames to construct a multimodal feature set. Then, it extracts lexical and semantic vectors from the text using a multi-layer BERT model, performs object detection on the images using ResNet-50 and extracts local semantic features, and combines the I3D network to model user actions in the video, extracting event labels and temporal behavioral information. The three modalities are then input into a joint semantic parsing module, where cross-modal alignment is achieved under coreference resolution constraints. This eliminates ambiguous expressions such as "it" and "this scene" across different modalities, generating a semantic structure sequence.

[0038] Next, the system extracts high-value advertising semantic elements such as "iced cola," "young people," and "drink after exercise" based on syntactic dependency paths and semantic role labeling results. The entity recognition module uses BiLSTM-CRF to perform type labeling and deduplication on all candidate entities, and combined with contextual dependency analysis, establishes a semantic relationship structure of "beverage—behavior—context." Subsequently, the system retrieves relevant node information from the food and beverage industry knowledge graph, performs entity-level fusion through SimRank and GraphSAGE, performs high similarity matching between the "beverage" node and the "carbonated beverage" category entity in the knowledge graph, and supplements its corresponding semantic attributes such as usage scenarios, emotional descriptions, and target audiences, correcting contextual information not explicitly stated in the original material.

[0039] During the content optimization phase, the system identified text fragments with unclear expressions or inconsistent language styles, such as vague expressions like "suitable for everyone" or "can quench thirst." Based on WordNet semantic relationships and a sentiment dictionary, it performed synonym replacements, optimizing these fragments into more attractive and expressive phrases like "targeting young, energetic groups" and "provides a rapid cooling sensation." Subsequently, the Viterbi algorithm was used to optimize the word order of the expression fragments, and preset sentence structure templates were applied to create diverse reconstructed sentence structures. The final generated new advertising slogans have a consistent style, clear information, and are automatically embedded into the context of the original image and video clips, achieving overall optimization of multimodal advertising creatives.

[0040] In a comparative experiment, 100 sets of original advertising creatives and 100 sets of creatives optimized using the method of this invention were selected and A / B tested separately. The main evaluation metrics included click-through rate (CTR), user dwell time, and ad content quality score. The testing platform targeted a randomly sampled ad audience and conducted comparative analysis within the same time period.

[0041] The data results are shown in Table 1. The optimized group outperformed the original group in all core metrics. The click-through rate increased by an average of 26.7%, content rating increased by 21.3%, and the average user dwell time increased by 3.4 seconds. This fully demonstrates the significant advantages of this invention in terms of the semantic richness, expressive accuracy, and style adaptability of advertising creative content.

[0042] Table 1. Comparison of Ad Creative Performance Before and After Optimization

[0043] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for intelligent optimization of advertising creatives based on knowledge extraction and knowledge fusion, characterized in that, Includes the following steps: S1. Obtain the advertising materials to be optimized and preprocess them to construct a multimodal feature set, while collecting external knowledge sources; S2. Perform joint semantic parsing on the multimodal feature set and use coreference resolution to perform cross-modal semantic alignment to form a semantic structure sequence; S3. Perform syntactic dependency analysis and semantic role labeling on the semantic structure sequence, extract entity units and corresponding semantic attributes, use entity recognition methods to label the entity units by type, and establish the semantic relationship structure between entities based on contextual dependency relationships. S4. Semantically fuse the semantic relation structure with the entity nodes and relation edges in the external knowledge source, calculate the similarity between semantic vectors based on graph embedding representation and perform entity mapping, normalize the node pairs with large semantic deviations, and correct conflict boundaries and attribute value discrepancies. S5. Based on the fused semantic structure, extract the target expression fragments from the advertising materials, and perform semantic content replacement, expression structure adjustment and style mapping according to the preset optimization rule set; S6. Perform semantic consistency verification and expression integrity detection on the optimized advertising materials, and output the optimization results.

2. The intelligent optimization method for advertising materials based on knowledge extraction and knowledge fusion according to claim 1, characterized in that, The advertising materials include text content, image content, and video frame sequences. The preprocessing includes performing language standardization, special symbol removal, and paragraph segmentation on the text content; performing resolution normalization, redundant background removal, and main area annotation on the image content; and performing keyframe extraction and time segmentation on the video frame sequences.

3. The intelligent optimization method for advertising materials based on knowledge extraction and knowledge fusion according to claim 1, characterized in that, The external knowledge sources include advertising industry knowledge graphs, user interest tag libraries, and historical campaign feedback datasets, specifically including: The advertising industry knowledge graph is constructed based on structured advertising data and industry public corpus. It uses triples to represent entities and corresponding semantic relationships. The entities include brands, product categories, marketing scenarios and performance styles. The semantic relationships include attribute relationships, subordinate relationships and style association relationships. The user interest tag library consists of a set of user browsing behavior, click behavior and conversion behavior. User behavior is tagged and modeled using a clustering algorithm, and a many-to-many mapping relationship between user identifiers and interest tags is established. The historical campaign feedback dataset is extracted from the records of each ad creative campaign, recording the creative identifier, campaign channel, target audience characteristics, and performance metrics, including click-through rate, conversion rate, and dwell time.

4. The intelligent optimization method for advertising materials based on knowledge extraction and knowledge fusion according to claim 1, characterized in that, S2 specifically includes: S21. Perform joint semantic parsing on the multimodal feature set, specifically including: For the text content, a multi-layer semantic encoding method based on the BERT model is used to extract semantic representations that include lexical structure, contextual dependencies and inter-sentence relations; For the image content, the ResNet-50 network is used to perform object detection, and the detected area is segmented and feature embedded to obtain local visual semantic vectors; For video frames, the I3D network is used to perform 3D convolutional action modeling to extract behavioral semantic vectors and event label representations from temporally continuous frames; Semantic vector sets for text, image, and video modalities are obtained respectively; S22. Based on the timestamps, location information, bounding box indexes and semantic labels contained in the semantic representation, perform cross-modal semantic alignment processing, establish multi-dimensional mapping relationships for semantic segments with referential relationships, and mark potential core-referenced segments; S23. A coreference resolution method based on central entity constraints is adopted to perform semantic discrimination and reference classification on the marked fragments, fuse semantic units with coreference relations, perform vector merging and subject label replacement operations, and eliminate semantic conflicts and redundant expressions. S24. Rearrange the processed semantic units according to modal independence order and semantic dependency strength, and combine them to generate a semantic structure sequence.

5. The intelligent optimization method for advertising materials based on knowledge extraction and knowledge fusion according to claim 1, characterized in that, S3 specifically includes: S31. Perform syntactic dependency analysis on text segments in semantic structure sequence, extract subject-predicate, modification and parallel dependency paths between words, control semantic grouping with dependency edge weights, and perform structural segmentation on long sentences with semantic span exceeding a set threshold. S32. Perform semantic role labeling within the structural segmentation unit, establish semantic mapping relationships with action verbs as the core, and extract the role correspondence between verbs and their semantic objects through a limited context window, and label semantic roles and structural position indices. S33. Using the semantic role labeling results and syntactic boundaries as joint inputs, perform entity recognition on candidate phrases, exclude boundary overlaps and semantically repetitive segments, use BiLSTM-CRF to label entity types, and form a non-overlapping entity sequence. S34. Perform dependency path backtracking on entity pairs in the entity sequence, calculate the difference between path depth and word order, filter entity pairs that satisfy contextual dependency constraints, and construct semantic connections with related directions based on semantic role relationships, specifically including: Syntactic dependency path backtracking is performed on entity pairs in non-overlapping entity sequences. A limited breadth-first traversal algorithm is used to extract the shortest syntactic path between entity pairs and record the verb, conjunction and modifier nodes in the path. The word nodes in the backtracking path are weighted by path depth and word order distance, and the context relationship score between entity pairs is calculated. Semantic connection candidate pairs are set according to the score. Perform semantic role consistency judgment on candidate entity pairs, filter out entity pairs with semantic category conflicts or duplicate semantic roles, and set the direction of the connection edge according to semantic directionality; Weighted directional edges are generated from the filtering results, and the semantic relationships between entity pairs are represented as structured connection information including role type, context label and connection confidence, which is then added to the semantic relationship structure graph.

6. The intelligent optimization method for advertising materials based on knowledge extraction and knowledge fusion according to claim 5, characterized in that, S33 specifically includes: S331. Align the semantic role labeling results with the syntactic boundary, construct a feature tensor containing role labels, part-of-speech tags, boundary positions and context window indices, and composite input entity recognition model; S332. Perform boundary overlap detection on the candidate phrase set, and adopt a dual threshold strategy of maximum overlap rate and label consistency to remove redundant segments with overlap rate higher than the set threshold or label conflict to form a preliminary entity set. S333. In the BiLSTM encoding process, a multi-channel input structure of word vector embedding, position embedding and role tag embedding is introduced to output a context-aware representation vector sequence. S334. Input the representation vector sequence into the CRF layer, calculate the optimal label path based on the state transition probability and observation probability, and generate the label sequence. S335. Recover entity intervals based on entity boundary markers in the label sequence, filter preferred entities by word order, and establish a non-overlapping entity index sequence.

7. The intelligent optimization method for advertising materials based on knowledge extraction and knowledge fusion according to claim 1, characterized in that, S4 specifically includes: S41. Retrieve the semantic counterpart nodes of entity nodes in the semantic relation structure in the external knowledge source, extract the label, attribute fields and related edge information, calculate the structural similarity between nodes using SimRank, and filter node pairs with consistent semantic labels or embedding vector similarity higher than the threshold. S42. Perform structural-level fusion processing on the selected node pairs, perform field alignment and missing data completion operations, set conflict markers for fields with inconsistent attribute values, and mark the original relation weights of conflicting edges. S43. The fused structure is vector-encoded using a graph embedding method based on GraphSAGE, and the contextual feature representation of each node is extracted. The semantic similarity between entity nodes is then recalculated in a unified vector space. S44. Perform attribute vector normalization and weight reset on entity nodes with similarity below the set threshold, adjust the main value source of conflicting labels according to the preset fusion rules, and correct the attribute divergence of relation edges caused by semantic conflicts.

8. The intelligent optimization method for advertising materials based on knowledge extraction and knowledge fusion according to claim 1, characterized in that, S5 specifically includes: S51. Locate the core entity nodes and behavior predicate nodes on the main semantic path in the fused semantic structure, and filter the start and end boundaries of the target expression fragment based on path depth and semantic association weight. S52. Perform semantic consistency screening on the selected segments, combine the importance score of the LexRank algorithm, remove candidate segments with structural ambiguity, semantic breaks or expression conflicts, and generate the target expression segment sequence according to the word order. S53. Perform semantic mapping on the content words in the target expression segment, perform equivalent word replacement based on WordNet vocabulary hierarchy and sentiment dictionary rules, and call the style vocabulary to adjust the sentiment tendency and tone expression; S54. Based on the sentence structure templates in the preset rule set, perform expression structure rearrangement on the expression fragments, and use the Viterbi algorithm to optimize the word order arrangement path, adjust the word order structure, modification dependency and syntactic pattern. S55. Embed the replaced expression fragment into the corresponding context position in the original material to construct the optimized semantic expression content of the advertisement.

9. The intelligent optimization method for advertising materials based on knowledge extraction and knowledge fusion according to claim 8, characterized in that, The preset rule set includes sentence structure constraint rules, word order rearrangement pattern library and style adaptation parameter set. The rule set is used to constrain the range of structural transformation of expression fragments and the conditions for rearrangement order selection. The sentence structure template contains multiple structural styles built based on semantic scenarios, specifically including declarative, emphatic, interrogative, inverted, and rhetorical sentence structures. Each template defines fixed word order rules and variable component adjustment ranges corresponding to semantic roles, and uses a structural index mapping method to mark the positional boundaries of the subject, predicate, object, and modifier. The expression structure rearrangement process specifically includes: Based on the semantic tags of entities and verbs in the current segment, match the corresponding sentence structure templates and identify the syntactic nodes in the segment that need to be adjusted; The positions of modifiers and predicates are sorted and rearranged according to template rules, and the structural similarity scores of the candidate permutation paths are labeled. The Viterbi algorithm is used to decode and compute all candidate paths, and the word order arrangement that maintains semantic consistency and has the best structural score is selected. After the rearrangement is completed, modal particles, conjunctions, and adverbs are inserted or replaced according to the style adaptation parameters.

10. An intelligent optimization system for advertising creatives based on knowledge extraction and knowledge fusion, comprising executing the intelligent optimization method for advertising creatives based on knowledge extraction and knowledge fusion as described in any one of claims 1 to 9, characterized in that, include: The data preprocessing module is used to acquire and preprocess the advertising materials to be optimized, construct a multimodal feature set, and collect external knowledge sources; The joint semantic parsing module performs joint semantic parsing on the multimodal feature set and uses coreference resolution to perform cross-modal semantic alignment, forming a semantic structure sequence; The knowledge extraction module is used to perform syntactic dependency analysis and semantic role labeling on semantic structure sequences, extract entity units and corresponding semantic attributes, use entity recognition methods to label entity units by type, and establish semantic relationship structure between entities based on contextual dependency relationships. The knowledge fusion module performs semantic fusion with the semantic relation structure and entity nodes and relation edges in the external knowledge source. It calculates the similarity between semantic vectors based on graph embedding representation and performs entity mapping. It normalizes node pairs with large semantic deviations and corrects conflict boundaries and attribute value discrepancies. The expression optimization module is used to extract target expression fragments from advertising materials based on the fused semantic structure, and perform semantic content replacement, expression structure adjustment and style mapping according to the preset optimization rule set; The expression verification module performs semantic consistency verification and expression integrity detection on the optimized advertising creatives and outputs the optimization results.

Citation Information

Cited By

  • Intelligent data processing method and system based on knowledge graph

    CN121051252A