Advertisement creativity matching method based on multi-modal content generation
By establishing cross-modal time anchoring bands and semantic boundary lists, combined with cultural fingerprint databases and semantic guardrail sets, the problem of semantic overlap in multimodal advertising content generation was solved, achieving clear semantic hierarchy and cultural consistency in advertising creatives, and improving the coordination and dissemination effect of advertising creatives.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-07
AI Technical Summary
In the current multimodal advertising content generation, the semantic distribution boundaries between modalities are easily compressed or overlapped, leading to semantic overlap, misuse of cultural imagery or deviation in value orientation, which affects the cultural consistency and communication orientation of advertising creativity.
By establishing cross-modal time anchoring bands and semantic boundary lists, constructing a cultural fingerprint database and semantic guardrail set, synchronous registration of textual and image information and breathing phase traction control are achieved, ensuring semantic convergence and stable brand symbolic expression.
Maintaining a clear semantic hierarchy during the creative process of advertising is crucial to avoiding semantic overlap and cultural mismatch. It is essential to ensure the coherence and adaptability of the generated content in terms of emotional rhythm, semantic imagery, and cultural context, thereby enhancing the coordination and communication effectiveness of advertising creatives.
Smart Images

Figure CN121808075A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital media content generation technology, and more specifically to an advertising creative matching method based on multimodal content generation. Background Technology
[0002] Multimodal content-generated ad creative matching refers to the use of artificial intelligence technology in the ad design and delivery process to fuse and generate ad creative elements based on various types of content data—including text, images, audio, video, user behavior patterns, and emotional characteristics. By understanding the target audience's interests, contextual semantics, and the characteristics of the communication medium, it automatically generates context-appropriate ad creative elements, achieving precise matching between ad content and communication channels, audience emotions, and display timing. This method overcomes the limitations of traditional advertising relying on single materials or manual planning, enabling ad creatives to form an adaptive and adjustable generation loop. This allows for a dynamic alignment between ad presentation and audience psychological resonance through multidimensional information interaction.
[0003] The existing technology has the following shortcomings: In the existing multimodal advertising content generation process, joint training is often used to semantically map and fuse multi-source data such as text, images, and audio. However, due to differences in the abstraction levels, contextual dependencies, and cultural metaphors of different modalities, existing technologies are prone to potential semantic space collapse during the generation stage. This means that the semantic distribution boundaries between modalities are incorrectly compressed or overlapped, causing the model to be unable to correctly distinguish the semantic hierarchy between textual imagery and visual symbols. Consequently, image elements and textual metaphors overlap semantically, resulting in mismatched generated content. Furthermore, when semantic overlap involves regional cultural symbols or brand symbolic elements, it can easily lead to the misuse of cultural imagery or deviations in value orientation, resulting in the misinterpretation of the brand symbol system and severely impacting the cultural consistency and communication orientation of advertising creative output.
[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide an advertising creative matching method based on multimodal content generation to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: an advertising creative matching method based on multimodal content generation, comprising the following steps: Establish a cross-modal temporal anchoring band for advertising creative matching, perform metaphorical hierarchical decomposition on the input text information, and annotate the symbol axis of the image information to obtain an initial semantic boundary list containing multimodal semantic mapping relationships; A cultural fingerprint database is constructed based on the initial semantic boundary list, and regional taboo information and brand symbol information are mapped to constraint labels to obtain a set of semantic guardrails containing semantic boundary constraint information. Based on the semantic guardrail set, a content gate is established at the content generation entry point to receive text and image information, synchronize text intent and visual elements, and generate an alignment instruction sequence for multimodal alignment control. Breathing phase traction control is executed according to the alignment instruction sequence, and the rhythm and emotional curve of audio and image content are adjusted according to the time interval to obtain phase-compressed multimodal content results. The semantic boundary list is updated based on the phase-compressed multimodal content results, the cultural fingerprint weights are recalculated, new semantic boundary data is generated, and the content generation model is controlled to maintain semantic convergence and stable brand symbolic expression in subsequent generation processes.
[0007] Preferably, the steps for obtaining the initial semantic boundary list are as follows: Collect complete text information used in the advertisement, identify semantic units in the text information and divide them into narrative semantic units, emotional semantic units, symbolic semantic units and logical semantic units, mark the time position according to the order of appearance and semantic intensity of the semantic units, and form a time anchoring band composed of multiple time anchor points; Based on the time anchoring band, the text information is decomposed into metaphorical levels, and the text semantics are divided into basic, intermediate and deep layers according to the degree of abstraction of expression. Each metaphorical level is then associated with a time anchor point. Based on the time anchor band, semantic extraction and symbol axis labeling of image information are performed to identify visual symbols in the image and draw the trajectory of symbol change, so that the symbol axis nodes correspond to the time anchor points of the text metaphor level. By integrating temporal anchoring data from textual metaphor hierarchy and image symbol axis, semantic boundary markers are established and semantic interaction relationships are integrated to generate an initial semantic boundary list containing multimodal semantic mapping relationships.
[0008] Preferably, the steps for generating the semantic guardrail set are as follows: After obtaining the initial list of semantic boundaries, cultural features are identified for the semantic boundary nodes. Semantic elements related to regional cultural symbols, traditional color habits, religious symbolic patterns, festival phrases and brand spirit are extracted to form a list of cultural semantic elements that includes semantic category, cultural symbolic attributes, visual features, time location and semantic weight. Based on the list of cultural semantic elements, a cultural fingerprint database structure is established, and a cultural semantic fingerprint composed of cultural category, symbolic semantics, visual features, emotional attributes and temporal features is constructed, which corresponds to the boundary nodes in the initial semantic boundary list. Regional taboo information and brand symbol information are mapped into semantic constraint tags and associated with cultural semantic fingerprints to form constraint data that includes constraint category, constraint strength, cultural origin, brand orientation and semantic influence level. Integrate cultural semantic fingerprints and semantic constraint tags to generate a set of semantic guardrails containing semantic boundary constraint information in chronological order.
[0009] Preferably, when generating the semantic guardrail set, semantic constraint boundary values are assigned to semantic boundary nodes according to the constraint category and constraint strength in the semantic constraint tags. By limiting the expression range of textual metaphors and image symbols, a hierarchical semantic guardrail structure is constructed, so that the advertising content can maintain stable expression within the cultural semantic range, and strengthen the consistency of brand color, tone and visual composition within the brand symbol area.
[0010] Preferably, the steps for generating the alignment instruction sequence are as follows: Based on the boundary constraint information in the semantic guardrail set, a semantic input filtering layer is set at the content generation entry point to perform semantic screening on the input text and image information, blocking metaphorical words and culturally sensitive elements that exceed the allowed range of the semantic guardrail set; It receives semantically filtered text and image information, identifies textual intent from text information, extracts visual elements from image information, and performs corresponding matching based on cultural and brand constraints in the semantic guardrail set. Based on the time anchoring information in the semantic guardrail set, the text intent and visual elements are time-aligned and semantically registered to form semantic matching units. Transform semantic matching units into a sequence of alignment instructions that includes textual intent content, visual element attributes, cultural category codes, brand constraint markers, and sentiment bias values.
[0011] Preferably, the generated alignment instruction sequence is arranged in chronological order and cultural category order, and connection information is set between adjacent instructions to record the semantic change direction and intensity trend of the time period before and after, so that the content generation process maintains semantic continuity and forms a stable alignment relationship between language expression, visual symbols and cultural semantics.
[0012] Preferably, the steps for obtaining multimodal content results are as follows: A breathing phase traction framework is established based on the time anchors in the alignment instruction sequence. Each time anchor is divided into a rhythm control zone and an emotion control zone, so that the audio content and the image content form a breathing rhythm structure in the time dimension. The audio content is rhythmically adjusted based on the semantic intensity and emotional tendency in the alignment instruction sequence. The audio rhythm, volume and timbre are adjusted in different time intervals to make the audio content correspond to the semantic rhythm. Based on the same time interval structure, the rhythm and emotional curve of the image content are adjusted to keep the rhythm of the picture, color brightness and changes in visual center of gravity synchronized with the rhythm of the audio. Phase compression is performed based on a breathing-style phase traction framework to align the rhythmic peaks and valleys of the audio and image content in time, resulting in phase-compressed multimodal content.
[0013] Preferably, the semantic boundary list is updated based on the phase-compressed multimodal content results, the cultural fingerprint weights are recalculated, new semantic boundary data is generated, and the semantic distribution and brand expression of the content generation model are controlled as follows: Based on the phase-compressed multimodal content results, semantic association information is extracted, and textual semantic cues, visual symbolic expressions, audio emotional features and cultural tendencies are identified to generate multimodal semantic association information; The original semantic boundary list is updated based on multimodal semantic association information, the semantic boundary nodes are repositioned and multimodal weight coefficients are added to keep the semantic space synchronized in terms of time and emotion. Based on the updated semantic boundary list, the cultural fingerprint weights are recalculated, and a cultural balance vector is formed by combining cultural category statistics and semantic strength calculation. By combining the updated semantic boundary list with the cultural balance vector, new semantic boundary data is generated, which includes semantic hierarchy identifiers, cultural attribute tags, brand weight coefficients, and sentiment curve parameters, in order to maintain semantic convergence and the stability of brand symbolic expression.
[0014] The technical effects and advantages provided by the present invention in the above technical solution are as follows: This invention establishes cross-modal temporal anchoring bands and semantic boundary lists before content generation, enabling precise mapping between textual and image information in both temporal and semantic dimensions. This maintains clarity of semantic hierarchy and consistency of expression during content generation. Through the synergistic effect of semantic guardrail sets and content gates, textual intent and visual elements are synchronized during the input stage, avoiding semantic overlap and cultural mismatch. This ensures that the generated content maintains coherence and adaptability across emotional rhythm, semantic imagery, and cultural context.
[0015] This invention utilizes breathing-style phase traction control and dynamic updates of cultural fingerprint weights to create a continuous convergence mechanism for multimodal content during rhythmic and emotional changes, maintaining consistency in brand symbolic expression at both the visual and auditory levels. The generated results achieve synchronous stability in semantics, emotion, and cultural expression, strengthening the brand identity of the advertising content while ensuring the accuracy and coherence of cultural semantic transmission, thereby enhancing the overall coordination and communication effectiveness of advertising creative generation. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0017] Figure 1 This is a flowchart of the advertising creative matching method based on multimodal content generation according to the present invention. Detailed Implementation
[0018] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.
[0019] This invention provides, for example Figure 1 The ad creative matching method based on multimodal content generation shown includes the following steps: Establish a cross-modal temporal anchoring band for advertising creative matching, perform metaphorical hierarchical decomposition on the input text information, and annotate the symbol axis of the image information to obtain an initial semantic boundary list containing multimodal semantic mapping relationships; To accurately establish the semantic correspondence between textual and image information during the ad creative generation process, and to ensure the consistency and coordination of multimodal semantic expression over time, the following method is used to establish a cross-modal time anchoring band for ad creative matching: metaphorical hierarchical decomposition of the input textual information and symbolic axis annotation of the image information, thereby obtaining an initial semantic boundary list containing multimodal semantic mapping relationships. The specific steps are as follows: When inputting advertising creative content, the complete text information used in the advertisement is first collected. This text information typically consists of brand slogans, product function descriptions, emotional guidance, scene descriptions, and cultural symbolism. To establish a cross-modal time anchor band, it is necessary to first identify different types of semantic units in the text. Each semantic unit is categorized into narrative, emotional, symbolic, or logical semantic units based on its content attributes. For each semantic unit, its temporal position is marked according to its order of appearance, duration, and semantic emphasis within the advertising content. A preliminary time framework for the time anchor band is formed with chronological order as the horizontal axis and semantic intensity as the vertical axis. The time anchor band consists of multiple time anchors, each representing the temporal position and semantic intensity of a semantic unit in the text. The time interval between adjacent time anchors reflects the rhythmic changes in the text's semantics. The time anchor band established in this way provides a temporal reference basis for subsequent textual metaphorical hierarchy decomposition.
[0020] After obtaining the time anchor bands, the advertising text information is decomposed into metaphorical levels. This decomposition process divides the text semantics into multiple semantic levels based on the degree of abstraction and emotional intensity. First, core metaphorical words appearing in the text are identified, such as expressions related to brand spirit, cultural symbolism, emotional guidance, or product value. Then, based on the context in which the metaphors are attached, these words are semantically categorized. For example, warmth, protection, freshness, and exploration belong to emotional expression metaphors; East, future, nature, and city belong to cultural symbolic metaphors; speed, persistence, precision, and purity belong to functional metaphors. Each type of metaphor is further divided into basic, intermediate, and deep layers based on its semantic depth. The basic layer reflects directly perceptible semantic features, the intermediate layer embodies emotional guidance and image shaping, and the deep layer reflects brand culture and core values. Within the time anchor bands, each time anchor point corresponds to one or more metaphorical levels, forming a multi-layered correspondence between metaphorical levels and time anchor bands. In this way, the semantic content of the text is simultaneously unfolded in both the temporal and abstract dimensions, giving the advertising text information a multi-layered, alignable structure.
[0021] After completing the hierarchical decomposition of textual metaphors, semantic extraction and symbolic axis annotation are performed on image information based on the same time anchor band. Advertising image information generally includes elements such as people, background, products, environment, color, and lighting. To maintain a correspondence between image semantics and textual metaphor levels, image information needs to be time-annotated according to the time anchor points in the time anchor band. For each frame or scene, the main visual symbols are identified, such as the posture of figures representing the brand image, the color scheme reflecting the emotional atmosphere, changes in facial expressions conveying emotional guidance, the compositional center highlighting product features, and background elements representing cultural attributes. These visual symbols are semantically annotated, and their change trajectories are drawn along the timeline based on their appearance time, duration, and visual center of gravity in the image. Multiple symbolic change trajectories are summarized into a symbolic axis, which describes the temporal extension direction and semantic development path of image semantics. Through the symbolic axis, the evolutionary patterns of emotion, symbolism, and visual center of gravity in the image can be clearly understood. Then, the nodes of the symbol axis are mapped to the time anchors in the text metaphor hierarchy to form a cross-modal semantic correspondence, so that the text semantics and image semantics match each other in terms of time and content.
[0022] After establishing the correspondence between textual metaphor levels and image symbol axes, an initial semantic boundary list is generated by integrating the temporal anchoring data of both. This initial semantic boundary list clarifies the distribution boundaries and interaction ranges between different modalities of semantics. First, the correspondence between textual temporal anchors and image symbol nodes is traversed. When a significant change in the semantic intensity of the textual metaphor level is detected within a certain time interval, while the visual center of gravity of the image symbol axis shifts, a semantic boundary marker is established for that time interval. Each semantic boundary marker includes its temporal location, textual semantic level number, image symbol number, semantic association weight, and boundary extension range. The distance between adjacent semantic boundary markers reflects the changing trend of the semantic coupling between text and image. By continuously recording the semantic boundary markers across all time intervals, a complete semantic boundary trajectory is formed. Each node in the semantic boundary trajectory records specific semantic interaction relationships, including textual metaphor category, emotional tendency, image symbol attributes, and semantic mapping direction. Finally, all semantic boundary markers and trajectory information are integrated according to temporal order, semantic level order, and visual symbol order to obtain the initial semantic boundary list. This list comprehensively reflects the multimodal semantic mapping relationships between textual metaphor hierarchy, image symbol axis, and temporal anchoring band. Through this list, the distribution range of advertising content in the semantic space can be precisely defined, providing a clear semantic reference baseline for advertising creative generation. The initial semantic boundary list not only serves as the input foundation for the subsequent construction of the cultural fingerprint database but also provides a basis for semantic consistency control and cultural constraints during the advertising content generation process. This ensures that multimodal semantic boundaries are not compressed or overlapped in subsequent content generation, thereby maintaining a clear semantic hierarchy, accurate imagery expression, and stable brand recognition in the advertising creative output.
[0023] A cultural fingerprint database is constructed based on the initial semantic boundary list, and regional taboo information and brand symbol information are mapped to constraint labels to obtain a set of semantic guardrails containing semantic boundary constraint information. To ensure semantic consistency and brand identity stability across different cultural contexts, and to prevent cultural misuse or value deviations during ad generation, a cultural fingerprint database is constructed based on an initial semantic boundary list. This database maps regional taboo information and brand symbolic information to constraint tags, generating a semantic guardrail set containing semantic boundary constraint information. The specific steps are as follows: After obtaining the initial semantic boundary list, cultural features are identified for all semantic elements in the list. The initial semantic boundary list includes textual metaphor hierarchy information, image symbol axis information, semantic boundary nodes, and time anchoring band data. Each semantic boundary node corresponds to a specific textual semantic fragment and image symbol unit. To construct a cultural fingerprint database, it is necessary to first identify which semantic elements have cultural features or brand-related attributes. Cultural feature elements include content related to regional cultural symbols, traditional color habits, religious symbolic patterns, festival symbolic phrases, and social customs and emotions. For example, visual symbols such as red, dragon, lantern, rising sun, and lotus, as well as textual words such as reunion, auspiciousness, protection, prayer, and nobility. Brand-related elements include content directly related to brand spirit, brand vision, brand history, brand identification graphics, or brand tone language, such as brand slogans, logo colors, tone style, and core value semantics. Each semantic boundary node is analyzed item by item, extracting the cultural features and brand-related semantic units that appear within it, and recording their appearance position on the time anchoring band, duration, and corresponding area on the image symbol axis, forming a list of cultural semantic elements. Each record in the list includes a semantic category, cultural symbolic attributes, visual feature description, temporal location, and semantic weight. Through this process, cultural semantic elements are structured and organized, providing a clear and traceable semantic input foundation for the subsequent construction of a cultural fingerprint database.
[0024] After obtaining the list of cultural semantic elements, the structural framework of the cultural fingerprint database was established. The cultural fingerprint database is used to represent the stable correspondence between different cultural characteristics and semantic boundaries, enabling advertising content to possess cultural constraints during the generation process. The cultural fingerprint database consists of a series of cultural semantic fingerprints, each containing five dimensions of information: cultural category, symbolic semantics, visual features, emotional attributes, and temporal features. The cultural category distinguishes the cultural region to which the semantic element belongs, such as Eastern culture, Western culture, Middle Eastern culture, Latin American culture, and Nordic culture; the symbolic semantics describes the ideological content expressed by the cultural element, such as peace, purity, authority, mystery, romance, or striving; the visual features record how the semantics are expressed in the image, including the main compositional form, color scheme, material characteristics, and light and shadow levels; the emotional attributes record the psychological reaction of the audience when perceiving the cultural symbol, such as warmth, awe, calmness, excitement, or solemnity; and the temporal features describe the interval and duration of the cultural semantic element's appearance within the advertising time anchor zone. A correspondence table is created for all cultural semantic elements according to the above five dimensions, forming a cultural fingerprint index sequence. Each cultural fingerprint entry is associated with a specific boundary node in the initial semantic boundary list, ensuring that the cultural fingerprint can be accurately located in both temporal and semantic dimensions. In this way, the cultural fingerprint database establishes a full-link mapping from semantic boundaries to cultural semantics and then to visual expression, enabling multimodal content generation to have a foundation for cultural feature recognition and alignment.
[0025] After the cultural fingerprint database structure is completed, to prevent advertising content from triggering culturally sensitive issues or misaligning brand symbols, it is necessary to map regional taboo information and brand symbol information into semantic constraint tags. Regional taboo information refers to symbols, colors, image elements, or expressions considered inappropriate or having negative connotations in different cultural regions. For example, in East Asia, the number "four" is generally considered unlucky; in the Middle East, certain animal images, human postures, or religious symbols are strictly restricted; in some parts of Europe and America, black may symbolize solemnity or mourning and is not suitable for use with celebratory themes; and in Southeast Asia, the use of certain plants, insects, or deities requires caution. Brand symbol information includes brand-specific visual elements, brand values, core brand colors, brand tone, and brand narrative direction. To transform these cultural and brand attributes into constraint tags, regional taboos and brand symbols must first be summarized in a structured form. Each taboo information or brand symbol record includes content type, symbol characteristics, semantic meaning, cultural origin, applicable time frame, and expression boundaries. These records are then compared item by item with the cultural semantic fingerprints in the cultural fingerprint database. When the symbolic meaning, visual features, or emotional attributes of a cultural fingerprint are associated with a taboo or brand symbolic information, a corresponding semantic constraint label is generated. Each semantic constraint label consists of constraint category, constraint strength, constraint period, cultural origin, brand orientation, and semantic impact level. For example, when a cultural fingerprint involves the element of red lanterns, the constraint label can specify that it is suitable for festive themes and prohibited from use in mourning contexts; when a cultural fingerprint involves the brand's blue tones, the constraint label can define that it should maintain a calm and rational style and should not be transformed into warm or passionate semantics. Through this mapping process, each semantic fingerprint in the cultural fingerprint database is given corresponding cultural and brand constraints, giving the semantic content controllable expression boundaries.
[0026] After completing the semantic constraint label mapping, all cultural semantic fingerprints containing constraint labels in the cultural fingerprint database are integrated to generate a semantic guardrail set containing semantic boundary constraint information. The semantic guardrail set is used to define the expression range, constraints, and cultural restrictions of each semantic interval during the advertising content generation process. The process of generating the semantic guardrail set includes three specific operations. First, based on the order of the time anchor band, the cultural semantic fingerprint entries are arranged according to their temporal characteristics, so that the cultural fingerprint and its constraint label corresponding to each semantic boundary node are accurately located on the timeline. Second, based on the constraint category and constraint strength in the constraint label, a semantic constraint boundary value is assigned to each semantic boundary node to limit the usable expression area of textual metaphors and image symbols within that time period. For example, when a certain time period involves culturally taboo symbols, the semantic guardrail set sets prohibited visual features and textual semantics within that area; when a certain time period involves brand symbol expression, the semantic guardrail set strengthens the brand's exclusive colors, tone, and compositional elements within that area to maintain a unified style. Third, all semantic boundary constraint values are integrated according to chronological order, semantic hierarchy order, and cultural category order to form a hierarchical semantic guardrail set. Each layer of the semantic guardrail set corresponds to a specific cultural category, and each layer is arranged chronologically, forming a top-down cultural semantic control structure. Through this structure, advertising content generation is confined within a safe semantic range at the cultural level, while brand expression maintains continuity and consistency at both the visual and linguistic levels. The resulting semantic guardrail set provides a deterministic semantic boundary reference for subsequent content gate establishment, semantic registration, and phase-guided control, ensuring stability in multimodal content generation regarding cultural expression, brand symbolism, and semantic consistency.
[0027] Based on the semantic guardrail set, a content gate is established at the content generation entry point to receive text and image information, synchronize text intent and visual elements, and generate an alignment instruction sequence for multimodal alignment control. To ensure semantic consistency and rhythmic coordination between text and image information during the advertising creative generation stage, and to guarantee that the generated content conforms to the cultural boundaries and brand constraints defined by the semantic guardrail set, a content gate is established at the content generation entry point based on the semantic guardrail set. This gate receives text and image information, synchronously registers textual intent and visual elements, and generates an alignment instruction sequence for multimodal alignment control. The specific steps are as follows: After generating the semantic guardrail set, the content generation process needs to establish an entry control mechanism to ensure that the semantic information input into the content generation stage is always within a controlled range. The content gate is established based on the boundary constraint information in the semantic guardrail set. The semantic guardrail set includes semantic boundary values, cultural categories, brand constraint tags, and allowed expression ranges for each time period. First, a semantic input filtering layer is set at the content generation entry point. This layer semantically filters text and image information about to enter the generation stage based on the boundary constraint information recorded in the semantic guardrail set. When the input text information contains metaphorical words, culturally sensitive content, or descriptions that do not conform to the brand's semantic constraints and exceed the allowed range of the semantic guardrail set, the information is temporarily blocked or delayed before entering the generation stage. Similarly, when image information contains visual elements that violate cultural taboos or deviate from the brand's symbolic characteristics, such as inappropriate color combinations, culturally conflicting patterns, or discordant scene arrangements, the input of that image element is adjusted to a pending review state. In this way, the content gate forms a semantic filtering channel at the content generation entry point, ensuring that all subsequent input multimodal content conforms to the constraints of the semantic guardrail set. This step establishes a semantically safe starting point for content generation, providing a limited input space for synchronous registration.
[0028] After establishing the content gate, the system continuously receives semantically filtered text and image information. Text information comes from advertising creative descriptions, brand slogans, and scene-based narratives, while image information comes from advertising visual materials, brand visual templates, and environmental background images. The receiving phase is not only an information input process but also a preliminary synchronization phase of multimodal semantic features. Therefore, it is necessary to identify textual intent within the text information after it has passed through the content gate, that is, to identify the core semantic objective expressed by each piece of text. Textual intent is typically manifested as action intent, emotional intent, brand intent, and value intent. For example, "warm companionship" embodies emotional intent, "technology empowerment" embodies brand intent, and "pure protection" embodies product value intent. Simultaneously, semantic analysis is performed on the image information to extract corresponding visual elements. Visual elements include the main subject's shape, character posture, scene composition, background lighting, main color distribution, and symbolic objects. Each visual element is identified as a visual expression unit corresponding to the textual intent and marked according to the cultural and brand constraints defined in the semantic guardrail set. When the textual intent involves emotional expression, visual elements with the same emotional attributes are prioritized for matching; when the textual intent contains cultural symbols, visual elements with the same cultural attributes are prioritized for matching. This matching method between textual intent and visual elements provides a basic semantic pairing structure for subsequent synchronous registration.
[0029] After obtaining the textual intent and corresponding visual elements, a synchronization registration relationship needs to be established along the time dimension to ensure that the textual semantics and visual expression are consistent across the timeline. The synchronization registration process is based on the time anchoring information, cultural categories, and semantic boundary constraints recorded in the semantic guardrail set. First, the semantic time slices in the textual intent are time-aligned with the visual change periods in the image information, ensuring that each semantic segment and its corresponding visual scene appear within the same time interval. Second, within each time interval, the semantic intensity of the textual intent is proportionally correlated with the expressive intensity of the visual elements, synchronizing the emotional fluctuations of the semantic expression with the rhythmic changes in the visuals. When a time interval is defined as a brand reinforcement area in the semantic guardrail set, the part of the textual intent involving the core brand expression is placed at the center of the time anchor point, and the corresponding visual elements highlight the brand logo or brand symbol graphics. When a time interval is defined as a cultural constraint area in the semantic guardrail set, both the textual intent and visual elements are adjusted to conform to the norms of that cultural expression. Through this time alignment and semantic intensity registration method, textual semantics and visual semantics form a synchronous correspondence across the timeline, semantic hierarchy, and cultural constraint dimensions. After registration is completed, each time segment corresponds to a semantic matching unit. This unit contains text intention identifiers, visual element descriptions, cultural category identifiers, and brand constraint tags, providing structured input for generating alignment instruction sequences.
[0030] After synchronizing and registering textual intent and visual elements, these registration results need to be transformed into an alignment instruction sequence usable in the content generation stage. The alignment instruction sequence is a structured set of instructions describing how multimodal content maintains semantic consistency during generation. When generating the alignment instruction sequence, firstly, the content of each semantic matching unit is arranged sequentially according to the time anchor band. Each semantic matching unit is transformed into a set of alignment instructions, including textual intent content, visual element attributes, cultural category codes, brand constraint markers, semantic intensity parameters, and sentiment values. Secondly, boundary constraint information from the semantic guardrail set is embedded into the instruction sequence, ensuring that each set of alignment instructions is controlled by semantic boundary constraints. When the generation stage processes the corresponding time segment, the content generation process automatically maintains semantic consistency between text and images based on this alignment instruction. Thirdly, to ensure the continuity of content generation, connection information is set between adjacent instructions in the instruction sequence, recording the direction and intensity trends of semantic changes between preceding and following time periods, making the semantic connection of the generation process natural. Finally, all instruction sequences are summarized in chronological and cultural category order to form a complete multimodal alignment control structure. This structure plays a semantic synchronization guiding role in the subsequent advertising creative generation stage, ensuring that the language expression, visual symbols and cultural semantics in the content generation process are unified in the time dimension, coordinated in emotional expression, and coherent in brand expression.
[0031] Breathing phase traction control is executed according to the alignment instruction sequence, and the rhythm and emotional curve of audio and image content are adjusted according to the time interval to obtain phase-compressed multimodal content results. To ensure consistent emotional rhythm in the audio and visual presentation of advertising creative content, and to achieve semantic, rhythmic, and emotional coordination during multimodal fusion generation, breathing-style phase traction control is executed according to the alignment instruction sequence. The rhythm and emotional curves of the audio and image content are adjusted according to time intervals to obtain phase-compressed multimodal content results. The specific steps are as follows: The generated alignment instruction sequence already contains textual intent, visual elements, cultural categories, brand constraints, and time anchoring information. To maintain semantic consistency during the multimodal content generation stage, a breathing-style phase traction framework needs to be established based on this instruction sequence. Breathing-style phase traction is a semantic synchronization method based on time-gap rhythm control. Its core is to create a dynamic correspondence between audio and image content in terms of rhythm and emotion through the periodic adjustment of time intervals. First, the time anchors in the alignment instruction sequence are extracted as the core reference nodes for the breathing rhythm. Each time anchor represents the start and end time of a semantic segment. These time anchors are arranged in sequence to form a time-gap division structure. Then, a rhythm control interval and an emotion control interval are assigned to each time gap. The rhythm control interval is used to limit the time range of rhythmic changes in the audio content, and the emotion control interval is used to limit the duration of emotional transitions in the image content. At this point, the breathing-style phase traction framework forms a periodic breathing rhythm structure on the timeline. The inhalation segment of the breathing rhythm represents the stage of semantic convergence and visual cohesion, while the exhalation segment represents the stage of semantic expansion and emotional release. This temporal structure provides a foundation for phase adjustment of audio and image content, enabling multimodal content to have a sense of rhythm in the temporal dimension.
[0032] After establishing the breathing-style phase traction framework, the audio content is first adjusted at the rhythm level to correspond to the semantic rhythm in the alignment instruction sequence. The audio content includes background music, ambient sounds, voiceover, and brand sound effects. For each time interval, the rhythmic pattern of the audio content is determined based on the semantic intensity and emotional tendency marked in the alignment instruction sequence. When the semantic intensity is rising, the audio rhythm gradually speeds up, the volume curve rises, and the timbre becomes brighter; when the semantic intensity is moderate, the audio rhythm slows down, the volume curve is stable, and the timbre becomes softer; when the semantic intensity is at its peak, the audio rhythm maintains a stable high-frequency rhythm and extends the duration of the sound tail to create a psychological extension effect. For time periods defined as brand expression zones in the semantic guardrail set, the audio content needs to strengthen brand identification sounds, such as brand logo sounds, rhythmic slogans, or melodic symbols; for time periods in cultural constraint zones, the audio content should adjust the instrument type and rhythmic pattern according to cultural semantic characteristics, such as Eastern culture tending to use soft linear rhythms, and Western culture tending to use layered beat structures. Through this audio-level rhythm control, the sound portion of the entire advertisement content is matched with the semantic rhythm, providing a synchronous basis for adjusting the image rhythm.
[0033] After the audio content's rhythm is adjusted, the image content's rhythm and emotional curve are adjusted according to the same time-slot structure. The image content includes the main character, product scene, background environment, lighting changes, and color atmosphere. First, the visual rhythm of each time slot is determined according to the time slots in the breathing-style phase traction framework. When the breathing rhythm is in the inhalation phase, the image content's rhythm tends to contract, the frequency of camera cuts decreases, the main object's composition is concentrated, color brightness decreases, and the visual center of gravity focuses towards the center, creating a semantic convergence effect. When the breathing rhythm is in the exhalation phase, the image content's rhythm tends to expand, the frequency of camera cuts increases, the visual space stretches, color brightness increases, and the visual center of gravity spreads outward, creating an emotional release effect. Second, within each time slot, an image emotional curve is drawn based on the emotional tendency information in the alignment instruction sequence. The emotional curve describes the trend of changes in image color, lighting, and character expressions over time. For example, when the semantic emotional tendency is "warm protection," the emotional curve rises slowly, with orange-yellow as the main hue and smooth lighting; when the semantic emotional tendency is "power breakthrough," the emotional curve fluctuates significantly, with obvious hue contrast and strong lighting. In this way, the visual content maintains a consistent breathing rhythm with the audio content over time, allowing viewers to experience synchronized emotional changes. The rhythmic pulses of the audio and the visual rhythm of the images alternate, creating a multimodal sensory fusion effect.
[0034] After adjusting the rhythm and emotional curves of the audio and image content, phase compression is performed to achieve maximum synchronization of the multimodal content in the temporal dimension. The core of phase compression is to gradually shorten the time difference between the audio and image rhythms, causing the peaks and troughs of both modalities to overlap within the same breathing cycle. First, based on a breathing-style phase traction framework, the phase difference between audio and image rhythm changes in all time intervals is identified. Then, during the closing phase of the breathing cycle, the lead of the audio rhythm peaks is adjusted so that the rhythm peaks of the audio and image appear synchronously; during the releasing phase of the breathing cycle, the tail of the image rhythm is extended so that the visual expression is consistent with the tail of the audio. In this way, the rhythm peaks and troughs of the audio and image are perfectly aligned in time. Second, the fluctuations in the emotional curve are synchronized with the changes in the intensity of the audio rhythm, ensuring consistency in the intensity of emotional expression across both modalities. When the audio shows an upward trend in emotion, the image brightness increases synchronously, and the character's expression becomes more positive; when the audio shows a downward trend in emotion, the image tone gradually cools, and the lighting becomes softer. Through this series of phase adjustments, the audio and image content resonate across the dimensions of time, rhythm, and emotion, resulting in a phase-compressed multimodal content. The final multimodal content exhibits temporal consistency, rhythmic coherence, and emotional harmony, creating a unified auditory and visual expression for the advertising creative. This result not only maintains the cultural and brand constraints set by the semantic guardrail set but also presents a natural breathing rhythm in the multimodal output, enhancing the overall expressiveness and cultural adaptability of the advertising content.
[0035] The semantic boundary list is updated based on the phase-compressed multimodal content results, the cultural fingerprint weights are recalculated, new semantic boundary data is generated, and the content generation model is controlled to maintain semantic convergence and stable brand symbolic expression in subsequent generation processes. To ensure semantic consistency, cultural coherence, and brand symbol stability in advertising creative content generation, the semantic boundary list is updated based on the phase-compressed multimodal content results. Cultural fingerprint weights are recalculated, and new semantic boundary data is generated to control semantic convergence and brand symbol expression stability during content generation. The specific steps are as follows: After obtaining the phase-compressed multimodal content results, the first step is to extract the semantic association information contained within. The multimodal content results include audio and image content that have undergone breathing-style phase traction control, achieving temporal consistency in rhythm, emotion, and semantics. To transform this result into a data foundation for updating semantic boundaries, a comprehensive analysis of the semantic structure within the multimodal content is necessary. First, identify the semantic units appearing in the multimodal content, including textual semantic cues, visual symbolic expressions, audio emotional features, and emotional rhythm transition nodes. Each semantic unit corresponds to a complete semantic expression segment, which simultaneously contains temporal information, emotional attributes, and cultural tendencies. Second, in the temporal dimension, match the rhythmic peaks and troughs of the audio content with the inflection points of the emotional curve of the image content one by one, identifying semantically synchronized segments and semantically reversed segments. Semantically synchronized segments represent parts where audio and image semantics converge, while semantically reversed segments represent parts where audio and image semantics exhibit slight deviations. Third, assign cultural attributes to each semantic unit, and combine this with the semantic guardrail set from the previous stage to determine whether it is located in a cultural constraint zone or a brand reinforcement zone. In this way, the extracted semantic association information not only records the temporal relationship between semantics and emotion, but also preserves the distribution of cultural and brand semantics within multimodal content. This semantic association information will serve as the basic input data for updating the semantic boundary list.
[0036] After obtaining multimodal semantic association information, the original semantic boundary list is updated to reflect the semantic changes after phase compression. The semantic boundary list initially originates from cross-modal time anchoring bands, describing the distribution of textual metaphor hierarchy, image symbol axes, and semantic boundary nodes. As the multimodal content undergoes breathing-like phase traction control, the temporal position, emotional intensity, and semantic density of some semantic boundaries are adjusted, necessitating the repositioning of these boundary nodes. First, the time synchronization segments in the semantic association information are compared with the time anchors in the semantic boundary list. When the audio-visual rhythms in the multimodal content completely overlap, the original semantic boundary positions are retained; when there is a slight time shift in the multimodal content, the semantic boundary nodes are moved forward or backward according to the direction of rhythmic change, aligning them with the new semantic synchronization segments. Second, the boundary nodes of the semantic reverse segments are merged or segmented. When the audio-visual semantics are not completely consistent within a certain time period, transition nodes need to be added to the original semantic boundaries to smooth the semantic connection between different modalities. In this way, the updated semantic boundary list can accurately describe the multimodal semantic space structure after phase compression. Finally, multimodal weight coefficients are added to each semantic boundary node to record the semantic contribution ratio of audio and image at that node, making subsequent cultural fingerprint weight adjustments more accurate. Through the above operations, the semantic boundary list completes the transformation from a static structure to a dynamic structure, realizing the synchronous updating of the semantic space in terms of time and emotion.
[0037] After updating the semantic boundary list, the cultural fingerprint weights need to be recalculated based on the semantic distribution of the new boundaries. Cultural fingerprint weights describe the semantic proportion and influence intensity of different cultural elements in the advertising content. With the completion of multimodal content phase compression, the expression frequency and semantic weights of some cultural symbols have dynamically changed. First, cultural category statistics are performed on the updated semantic boundary list to identify the frequency and duration of each cultural category's appearance in the multimodal content. Second, the average sentiment and semantic intensity of each cultural category in different time intervals are calculated to obtain its sentiment stability and semantic contribution. For example, when the sentiment curve of Eastern cultural semantics in multimodal content fluctuates less and lasts longer, its weight should be higher than that of Western cultural semantics, which are short-lived but emotionally intense. Third, the cultural semantics involving brand symbols are weighted and adjusted in conjunction with the constraint tags in the semantic guardrail set, ensuring that the core brand culture occupies a dominant position in the overall cultural semantics. Finally, the weights of all cultural categories are arranged in chronological order to form a cultural balance vector. The cultural balance vector describes the dynamic proportional relationship of different cultural expressions in the advertising content, providing a precise basis for cultural regulation in generating new semantic boundary data. Through this process, the cultural fingerprint database transforms from a static weighting system into a cultural expression structure that can be dynamically adjusted with semantic evolution, thereby ensuring that advertising content maintains cultural consistency and brand semantic stability in multimodal generation.
[0038] After recalculating the cultural fingerprint weights, the updated semantic boundary list is combined with the cultural balance vector to generate new semantic boundary data. This new semantic boundary data serves as the semantic control baseline in the advertising creative generation process, guiding the semantic distribution and brand expression of subsequent multimodal content generation. First, the time anchors, sentiment nodes, and cultural category identifiers in the semantic boundary list are mapped one-to-one with the weight coefficients in the cultural balance vector, generating semantic control entries for each time interval. Each semantic control entry includes a semantic level identifier, cultural attribute label, brand weight coefficient, sentiment curve parameter, and time position marker. Second, all semantic control entries are integrated chronologically to form a semantic boundary control sequence. This sequence is called sequentially during content generation to define the semantic scope of multimodal generation. When the generation process enters a specific time interval, the semantic boundary control sequence automatically loads the corresponding entries, ensuring that content generation only unfolds within that semantic interval, thus avoiding semantic drift. Third, within the brand symbol expression area, the new semantic boundary data continuously maintains the stability of the brand's visual characteristics and language style, ensuring a consistent expression of the brand symbol across different modalities. Finally, through a dynamic feedback mechanism of semantic boundary data, the subsequent content generation model maintains semantic convergence during continuous generation. Semantic convergence is manifested in the content remaining consistent with the initial brand value and cultural orientation after multiple rounds of generation, without semantic deviation or brand distortion. Thus, the new semantic boundary data becomes the core foundation for controlling the multimodal content generation process, ensuring that advertising creative output remains coordinated and stable at the semantic, cultural, and brand levels.
[0039] This invention establishes cross-modal temporal anchoring bands and semantic boundary lists before content generation, enabling precise mapping between textual and image information in both temporal and semantic dimensions. This maintains clarity of semantic hierarchy and consistency of expression during content generation. Through the synergistic effect of semantic guardrail sets and content gates, textual intent and visual elements are synchronized during the input stage, avoiding semantic overlap and cultural mismatch. This ensures that the generated content maintains coherence and adaptability across emotional rhythm, semantic imagery, and cultural context.
[0040] This invention utilizes breathing-style phase traction control and dynamic updates of cultural fingerprint weights to create a continuous convergence mechanism for multimodal content during rhythmic and emotional changes, maintaining consistency in brand symbolic expression at both the visual and auditory levels. The generated results achieve synchronous stability in semantics, emotion, and cultural expression, strengthening the brand identity of the advertising content while ensuring the accuracy and coherence of cultural semantic transmission, thereby enhancing the overall coordination and communication effectiveness of advertising creative generation.
[0041] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. An advertising creative matching method based on multimodal content generation, characterized in that, Includes the following steps: Establish a cross-modal temporal anchoring band for ad creative matching, perform metaphorical hierarchical decomposition on the input text information, and annotate the symbol axes on the image information to obtain an initial semantic boundary list; A cultural fingerprint database is constructed based on the initial semantic boundary list, and regional taboo information and brand symbol information are mapped as constraint labels to obtain a set of semantic guardrails. Based on the semantic guardrail set, a content gate is established at the content generation entry point to receive text and image information, synchronize text intent and visual elements, and generate an alignment instruction sequence. Breathing phase traction control is executed according to the alignment instruction sequence, and the rhythm and emotional curve of audio and image content are adjusted according to the time interval to obtain phase-compressed multimodal content results. The semantic boundary list is updated based on the phase-compressed multimodal content results, the cultural fingerprint weights are recalculated, new semantic boundary data is generated, and the content generation model is controlled to maintain semantic convergence and stable brand symbolic expression in subsequent generation processes.
2. The advertising creative matching method based on multimodal content generation according to claim 1, characterized in that, The steps to obtain the initial semantic boundary list are as follows: Collect complete text information used in the advertisement, identify semantic units in the text information and divide them into narrative semantic units, emotional semantic units, symbolic semantic units and logical semantic units, mark the time position according to the order of appearance and semantic intensity of the semantic units, and form a time anchoring band composed of multiple time anchor points; Based on the time anchoring band, the text information is decomposed into metaphorical levels, and the text semantics are divided into basic, intermediate and deep layers according to the degree of abstraction of expression. Each metaphorical level is then associated with a time anchor point. Based on the time anchor band, semantic extraction and symbol axis labeling of image information are performed to identify visual symbols in the image and draw the trajectory of symbol change, so that the symbol axis nodes correspond to the time anchor points of the text metaphor level. By integrating the temporal anchoring data of textual metaphor hierarchy and image symbol axis, semantic boundary markers are established and semantic interaction relationships are integrated to generate an initial semantic boundary list.
3. The advertising creative matching method based on multimodal content generation according to claim 2, characterized in that, The steps for generating a semantic guardrail set are as follows: After obtaining the initial list of semantic boundaries, cultural features are identified for the semantic boundary nodes, and semantic elements related to regional cultural symbols, traditional color habits, religious symbolic patterns, festival phrases and brand spirit are extracted to form a list of semantic elements. Based on the list of cultural semantic elements, a cultural fingerprint database structure is established, and a cultural semantic fingerprint composed of cultural category, symbolic semantics, visual features, emotional attributes and temporal features is constructed, which corresponds to the boundary nodes in the initial semantic boundary list. Regional taboo information and brand symbol information are mapped into semantic constraint tags and associated with cultural semantic fingerprints to form constraint data; Integrate cultural semantic fingerprints and semantic constraint tags to generate a set of semantic guardrails based on chronological order.
4. The advertising creative matching method based on multimodal content generation according to claim 3, characterized in that, When generating the semantic guardrail set, semantic constraint boundary values are assigned to semantic boundary nodes based on the constraint category and constraint strength in the semantic constraint tags. By limiting the expression range of textual metaphors and image symbols, a hierarchical semantic guardrail structure is constructed, so that the advertising content can maintain stable expression within the cultural semantic range, and strengthen the consistency of brand color, tone and visual composition within the brand symbol area.
5. The advertising creative matching method based on multimodal content generation according to claim 3, characterized in that, The steps for generating the alignment instruction sequence are as follows: Based on the boundary constraint information in the semantic guardrail set, a semantic input filtering layer is set at the content generation entry point to perform semantic screening on the input text and image information, blocking metaphorical words and culturally sensitive elements that exceed the allowed range of the semantic guardrail set; It receives semantically filtered text and image information, identifies textual intent from text information, extracts visual elements from image information, and performs corresponding matching based on cultural and brand constraints in the semantic guardrail set. Based on the time anchoring information in the semantic guardrail set, the text intent and visual elements are time-aligned and semantically registered to form semantic matching units. Transform semantic matching units into alignment instruction sequences.
6. The advertising creative matching method based on multimodal content generation according to claim 5, characterized in that, The generated alignment instruction sequence is arranged in chronological and cultural order, with connection information set between adjacent instructions. It records the semantic change direction and intensity trend of the time period before and after, so that the content generation process maintains semantic continuity and forms a stable alignment relationship between language expression, visual symbols and cultural semantics.
7. The advertising creative matching method based on multimodal content generation according to claim 5, characterized in that, The steps for obtaining multimodal content results are as follows: A breathing phase traction framework is established based on the time anchors in the alignment instruction sequence. Each time anchor is divided into a rhythm control zone and an emotion control zone, so that the audio content and the image content form a breathing rhythm structure in the time dimension. The audio content is rhythmically adjusted based on the semantic intensity and emotional tendency in the alignment instruction sequence. The audio rhythm, volume and timbre are adjusted in different time intervals to make the audio content correspond to the semantic rhythm. Based on the same time interval structure, the rhythm and emotional curve of the image content are adjusted to keep the rhythm of the picture, color brightness and changes in visual center of gravity synchronized with the rhythm of the audio. Phase compression is performed based on a breathing-style phase traction framework to align the rhythmic peaks and valleys of the audio and image content in time, resulting in phase-compressed multimodal content.
8. The advertising creative matching method based on multimodal content generation according to claim 7, characterized in that, The semantic boundary list is updated based on the phase-compressed multimodal content results. The cultural fingerprint weights are recalculated to generate new semantic boundary data. The steps to control the semantic distribution and brand expression of the content generation model are as follows: Based on the phase-compressed multimodal content results, semantic association information is extracted, and textual semantic cues, visual symbolic expressions, audio emotional features and cultural tendencies are identified to generate multimodal semantic association information; The original semantic boundary list is updated based on multimodal semantic association information, the semantic boundary nodes are repositioned and multimodal weight coefficients are added to keep the semantic space synchronized in terms of time and emotion. Based on the updated semantic boundary list, the cultural fingerprint weights are recalculated, and a cultural balance vector is formed by combining cultural category statistics and semantic strength calculation. The updated semantic boundary list is combined with the cultural balance vector to generate new semantic boundary data.
Citation Information
Cited By
Tuberculosis-based chest radiograph lesion detection method and system
CN122199530B
A brand constraint-based advertisement material generation method and system
CN122550233A