Content generation method and system based on multi-modal large model
By establishing a demand association network and a multimodal element coupling pool, and using a multimodal large model to generate a content framework that matches user needs, the problem of inconsistent content generation quality in existing technologies is solved, and high-quality, personalized content generation and dissemination are achieved.
Patent Information
- Application Number
- CN202511331884.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-09-18
AI Technical Summary
Existing multimodal content generation technologies struggle to deeply understand user needs and lack the inherent connections and logical consistency between different modalities, resulting in inconsistent content quality that fails to meet users' personalized needs.
Establish a demand association network, construct a multimodal element coupling pool, generate an initial content framework that matches user needs through a multimodal large model, and perform element collaborative optimization to generate an optimized content framework that conforms to the constraints of content element nodes.
It improves the quality and consistency of content generation, ensuring that the generated content accurately meets user needs, enhances the targeting and effectiveness of content dissemination, and provides a high-quality and personalized content experience.
Smart Images

Figure CN120832885B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital content creation and dissemination technology, and more specifically, to a content generation method and system based on a multimodal large model. Background Technology
[0002] In the field of digital content creation and dissemination, with the rapid development of information technology, users' demands for content are becoming increasingly diverse and personalized. Traditional content generation methods are often limited to a single modality, such as focusing solely on text creation or image design, making it difficult to meet users' needs for the integration of multiple modalities in different scenarios.
[0003] While existing multimodal content generation technologies have achieved a certain degree of integration of text, images, and audio, most lack a deep understanding and precise grasp of user needs. In the content generation process, they often simply pile up elements from different modalities without considering the inherent connections and logical consistency between these elements. This results in inconsistent content quality that fails to truly meet users' actual needs. Furthermore, traditional methods lack effective demand correlation analysis and element co-optimization mechanisms when handling complex content generation tasks. This makes it difficult to generate high-quality, personalized target content based on different scenario formats and content element constraints, significantly diminishing the effectiveness of content dissemination. Summary of the Invention
[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, a content generation method based on a multimodal large model is provided, the method comprising:
[0005] Establish a demand association network for content generation. The demand association network contains multiple nodes, including user demand nodes, content element nodes, and scene form nodes. Each node is connected by an association edge and the association strength is marked.
[0006] A multimodal element coupling pool is constructed based on the demand association network. The multimodal element coupling pool contains multiple element units, including text description units, visual display units, and semantic association units. Each element unit forms a mapping relationship with the nodes of the demand association network.
[0007] The multimodal large model is invoked to perform contextualized mapping processing on the multimodal element coupling pool, generating an initial content framework set that matches the demand-related network;
[0008] The initial content framework set is subjected to element collaborative optimization processing to obtain an optimized content framework that conforms to the content element node constraints;
[0009] A target content set is generated based on the optimized content framework, and the target content set is pushed to the corresponding content distribution node to complete the content dissemination operation.
[0010] In another aspect, the present invention also provides a content generation system based on a multimodal large model, including a processor and a machine-readable storage medium connected to the processor. The machine-readable storage medium is used to store programs, instructions or code, and the processor is used to execute the programs, instructions or code in the machine-readable storage medium to implement the above-described method.
[0011] Based on the above, by establishing a demand association network for content generation, user demand nodes, content element nodes, and scene form nodes are connected by association edges and the association strength is labeled. A multimodal element coupling pool is constructed based on this demand association network, enabling text description units, visual display units, and semantic association units to form mapping relationships with the nodes of the demand association network. This achieves the organic integration and orderly organization of multimodal elements. The multimodal large model is invoked to perform contextualized mapping processing on the multimodal element coupling pool, generating an initial content framework set that matches the demand association network. This fully utilizes the powerful capabilities of the multimodal large model to deeply integrate multimodal elements with user needs and scene forms, generating an initial content framework that conforms to the context. The initial content framework set undergoes element collaborative optimization processing to obtain an optimized content framework that conforms to the constraints of content element nodes, further improving the quality and consistency of the content. This ensures that the generated content accurately meets the requirements of user needs and content elements. Based on the optimized content framework, a target content set is generated and pushed to the corresponding content distribution node to complete the content dissemination operation. This effectively improves the efficiency and quality of content generation, enhances the targeting and effectiveness of content dissemination, and provides users with a higher quality and more personalized content experience. Attached Figure Description
[0012] Figure 1 This is a schematic diagram of the execution flow of the content generation method based on a multimodal large model provided in an embodiment of the present invention.
[0013] Figure 2 This is a schematic diagram of exemplary hardware and software components of the content generation system based on a multimodal large model provided in an embodiment of the present invention. Detailed Implementation
[0014] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating a content generation method based on a multimodal large model provided in an embodiment of the present invention. The following is a detailed description of this content generation method based on a multimodal large model.
[0015] Step S110: Establish a demand association network for content generation. The demand association network includes multiple nodes, including user demand nodes, content element nodes, and scene form nodes. Each node is connected by an association edge and the association strength is marked.
[0016] In this embodiment, the generation of promotional content for apparel products in the e-commerce sector is used as a unified application scenario for explanation. In this scenario, it is necessary to generate promotional content that meets the purchasing needs of different consumers, covers specific attributes of clothing, and is adaptable to various display scenarios. For example, generating multimodal promotional content for a summer dress that is adapted to different scenarios such as product detail pages, short video promotions, and live streaming displays.
[0017] Step S111: Collect user historical interaction records and content detail documents, extract user behavior keywords and feedback phrases from the user historical interaction records, and extract core content attribute words and feature descriptions from the content detail documents.
[0018] When collecting user interaction history records, consumer behavior data related to summer dresses on e-commerce platforms is collected. User interaction history records include the duration of time consumers browse summer dress pages on the platform, the number of times they click on different dress styles, the content of reviews posted in the product review section, and descriptions of purchased dresses. During the collection of this data, sensitive privacy data such as consumers' personal purchase records is involved. Data anonymization technology is used to anonymize consumers' personal identification information such as names and mobile phone numbers. Encryption algorithms are used to transform personal identifiers into meaningless character sequences. Simultaneously, data access control technology is employed to restrict access to sensitive data to only authorized personnel, ensuring that consumers' privacy information is not leaked.
[0019] Extract user behavior keywords and feedback phrases from user interaction history. For example, if a consumer enters "slimming summer dress" in the search bar, extract "slimming" and "summer dress" as user behavior keywords; if a consumer comments "I hope the dress fabric is breathable and suitable for summer," extract "breathable fabric" and "suitable for summer" as feedback phrases.
[0020] The detailed documentation includes product design specifications, fabric composition reports, size charts, and washing instructions for the summer dress. Core attribute terms and characteristic descriptions are extracted from these documents. Core attribute terms include "fabric material," "pattern design," "color and style," and "size." Characteristic descriptions include phrases such as "made of lightweight and breathable fabric" and "A-line design with a fitted waist."
[0021] Step S112: Cluster the user behavior keywords and feedback phrases to generate multiple user demand nodes. Each user demand node includes a demand type label and a demand description.
[0022] The extracted user behavior keywords and feedback phrases are clustered. First, these words and phrases are preprocessed to remove meaningless function words (such as "ah" and "ne") and to merge synonyms (e.g., "slimming" and "slender" are merged into "slimming and slender"). Then, a semantic similarity-based clustering algorithm is used to cluster words according to their semantic associations.
[0023] For example, keywords and phrases containing "slimming," "flattering," and "fitting" can be grouped into one category to generate a user demand node. This node's demand type tag is "slimming fit requirement," and the demand description is "Consumers want summer dresses with a slimming and fitted design to flatter their figure." Keywords and phrases containing "breathable fabric," "comfortable fabric," and "skin-friendly fabric" can be grouped into another category to generate a user demand node with the demand type tag "comfortable fabric requirement." The demand description is "Consumers require summer dresses to be made of breathable, comfortable, and skin-friendly fabric, suitable for summer wear." Additionally, user demand nodes such as "fashionable style requirement" and "reasonable price requirement" might be generated, corresponding to consumers' needs for fashionable and novel dress styles and prices within an acceptable range, respectively.
[0024] Step S113: The core attribute words and feature descriptions of the content are hierarchically divided to construct a content element node system, which includes main content category nodes, sub-content attribute nodes and content parameter nodes.
[0025] The core attribute terms and feature descriptions of summer dresses are hierarchically divided. The main content category node is set as "summer dresses".
[0026] Under the main content category node, multiple sub-content attribute nodes are further divided. For example, the "Fabric Attribute" sub-content attribute node covers content related to the dress fabric; the "Pattern Attribute" sub-content attribute node involves the dress pattern design; the "Style Attribute" sub-content attribute node contains the dress style details; and the "Size Attribute" sub-content attribute node involves various size parameters of the dress, etc.
[0027] Each sub-content attribute node is further divided into content parameter nodes. The content parameter nodes under the "Fabric Attributes" sub-content attribute node include "Fabric Composition", "Fabric Thickness", and "Fabric Breathability"; the content parameter nodes under the "Pattern Attributes" sub-content attribute node include "Overall Pattern", "Waist Design", and "Skirt Design"; the content parameter nodes under the "Style Attributes" sub-content attribute node include "Collar Design", "Sleeve Design", and "Pattern Decoration"; and the content parameter nodes under the "Size Attributes" sub-content attribute node include "Garment Length", "Bust", and "Waist".
[0028] Step S114: Analyze the content samples in the content library that match the user demand node and the content element node, and extract the presentation mode, structure type and modal combination mode of the content samples as the constituent elements of the scene form node.
[0029] The content samples related to summer dresses in the content library include existing product detail pages, promotional short videos, and live stream replay clips. The matching of these content samples with the aforementioned user demand nodes and content element nodes is analyzed.
[0030] For content samples that match the elements of "slimming fit" and "fit attributes," such as product detail pages and short videos showcasing dress fits, the presentation often involves multi-angle real-life photos combined with detailed descriptions, using images from different angles to demonstrate the dress fit effect; the structure type is usually "overall fit display - partial detail design - wearing effect demonstration"; the modal combination mode is a combination of text and static images.
[0031] For content samples that match the elements of "fabric comfort requirements" and "fabric attributes," such as live stream clips and details pages introducing dress fabrics, the presentation method is mostly to show the fabric texture up close and give a verbal explanation; the structure type is "fabric appearance display - fabric texture description - fabric characteristic description"; the modal combination mode is a combination of video, audio and text, where the video shows the fabric details, the audio provides an explanation, and the text presents information such as the fabric composition.
[0032] Based on these analyses, presentation methods, structural types, and modal combination patterns are extracted as constituent elements of scene format nodes, generating scene format nodes such as "Product Details Page Scene," "Short Video Promotion Scene," and "Live Streaming Display Scene." The constituent elements of the "Product Details Page Scene" are: presentation method combining multiple images with detailed text descriptions; structural type "overall product display - detailed feature introduction - specification parameter description - user review display"; and modal combination pattern "combination of text, static images, and dynamic images." The constituent elements of the "Short Video Promotion Scene" are: presentation method combining dynamic video display with background music; structural type "eye-catching opening - product highlight display - purchase guidance"; and modal combination pattern "combination of video, audio, and text subtitles." The constituent elements of the "Live Streaming Display Scene" are: presentation method combining real-time demonstration and interactive Q&A; structural type "product wearing effect display - detailed explanation - promotional activity introduction - Q&A interaction"; and modal combination pattern "combination of video, audio, and real-time interactive text."
[0033] Step S115: Calculate the co-occurrence frequency between the user demand node and the content element node using an association rule mining algorithm, and set a first association strength parameter based on the co-occurrence frequency as the annotation information for the association edge between the two.
[0034] The FP-Growth algorithm from association rule mining is used to calculate the co-occurrence frequency between user demand nodes and content element nodes. First, the behavioral data related to user demand nodes and content element nodes in the user's historical interaction records are transformed into transaction sets. Each transaction represents a combination of user demand nodes and content element nodes involved by a consumer during a shopping process.
[0035] For example, a consumer's shopping experience might involve both a "slimming fit" requirement and content elements such as "overall fit" and "waistline design," thus constituting a transaction. By statistically analyzing the number of times the "slimming fit" user requirement node and the "waistline design" content parameter node co-occur in multiple transactions, their co-occurrence frequency can be obtained.
[0036] The first association strength parameter is set based on the co-occurrence frequency. The higher the co-occurrence frequency, the larger the first association strength parameter, ranging from 0 to 1. For example, the user demand node "slimming fit requirement" and the content parameter node "waist-cinching design" have a high co-occurrence frequency, so the first association strength parameter is set to a high value; the co-occurrence frequency of "slimming fit requirement" and the content parameter node "loose cuffs" is low, so the first association strength parameter is set to a low value.
[0037] Step S116: Calculate the fit degree between the content element node and the scene form node, and set a second association strength parameter based on the fit degree as the annotation information of the association edge between the two.
[0038] When calculating the fit between content element nodes and scene format nodes, factors such as the clarity of content display and the effectiveness of information delivery are considered. For the "fabric composition" content parameter node and the "product details page scene" scene format node, "fabric composition" requires detailed text descriptions and close-up images, which the product details page can well meet, resulting in a high fit. However, the fit between "fabric composition" and the "short video promotion scene" is relatively low because short videos are too short to display fabric composition in detail.
[0039] In the specific calculation, an adaptation evaluation index system is constructed, including the matching degree between content information density and scene capacity, and the fit between content display needs and scene presentation capabilities. Each index is scored, and then the scores of each index are weighted according to a set weight to obtain the overall adaptation degree. A second association strength parameter is set based on the adaptation degree; the higher the adaptation degree, the larger the second association strength parameter, also ranging from 0 to 1. For example, the adaptation degree calculation result for "fabric composition" and "product details page scene" is relatively high, and the corresponding second association strength parameter is also relatively high; the adaptation degree calculation result for "overall pattern display" and "short video promotion scene" is relatively high, and the corresponding second association strength parameter is also relatively high.
[0040] Step S117: Calculate the satisfaction degree between the user demand node and the scene form node, and set a third association strength parameter based on the satisfaction degree as the annotation information of the association edge between the two.
[0041] Calculating the satisfaction level between user demand nodes and scenario format nodes primarily considers whether the scenario format node can effectively meet the needs described by the user demand node. For example, the user demand node "fashionable style requirement" and the scenario format node "short video promotion scenario" have a higher satisfaction level because short videos can highlight the fashionability of the dress through dynamic displays and fashionable shooting techniques. However, the satisfaction level between "fashionable style requirement" and "product details page scenario" is relatively low because the details page mainly uses static displays.
[0042] When calculating satisfaction, conversion rates and user satisfaction levels for the user's need node in the corresponding scenario are analyzed from historical data. If a high percentage of consumers in the "short video promotion scenario" purchase a dress because they are attracted by its style, and their satisfaction with the style is high, then the satisfaction level for "style fashion need" and "short video promotion scenario" is high. A third association strength parameter is set based on the satisfaction level; the higher the satisfaction level, the larger the third association strength parameter, ranging from 0 to 1. For example, the third association strength parameter corresponding to the satisfaction level of "style fashion need" and "short video promotion scenario" is set to a relatively high value; the satisfaction level of "material comfort need" and "product details page scenario" is relatively high because the details page can display detailed information such as the texture and composition of the material, and its corresponding third association strength parameter is also set to a relatively high value, while the satisfaction level of "material comfort need" and "short video promotion scenario" is relatively low, and the third association strength parameter is set to a relatively low value.
[0043] Step S118: The user demand node, content element node, scene form node, and the associated edges labeled with the first association strength parameter, the second association strength parameter, and the third association strength parameter are networked and integrated to generate a demand association network with a three-layer node structure.
[0044] The generated user demand nodes, content element nodes, and scene form nodes are connected according to their relationships to form an interconnected network structure. In this network, each node in the user demand node layer is connected to the corresponding node in the content element node layer through an association edge, with the association edge labeled with a first association strength parameter; nodes in the content element node layer are connected to nodes in the scene form node layer through an association edge, with the association edge labeled with a second association strength parameter; nodes in the user demand node layer are also directly connected to nodes in the scene form node layer through an association edge, with the association edge labeled with a third association strength parameter.
[0045] For example, the "Style and Fashion Needs" node is connected to content parameter nodes such as "Dress Style Design" and "Pattern Elements" via an edge labeled with the first correlation strength parameter; the "Dress Fabric Material" content parameter node is connected to the "Product Details Page Scenario" node via an edge labeled with the second correlation strength parameter; and the "Value for Money Needs" node is connected to the "Promotional Activity Page Scenario" node via an edge labeled with the third correlation strength parameter. Through this networked integration, a complete demand association network is formed, which can accurately reflect the relationship and strength between user needs, content elements, and scenario forms.
[0046] Step S120: Construct a multimodal element coupling pool based on the demand association network. The multimodal element coupling pool contains multiple element units, including text description units, visual display units, and semantic association units. Each element unit forms a mapping relationship with the nodes of the demand association network.
[0047] Step S121: Based on the user demand nodes in the demand association network, retrieve descriptive words, sentence templates and text structure fragments that match the demand type tags in the text corpus, and then classify them according to the demand description to generate text description units. Each text description unit contains a classification identifier and corresponding text content.
[0048] The text corpus stores a large amount of textual information related to dresses, including product titles, detailed descriptions, user reviews, and fashion news. Based on the demand type tags of user demand nodes in the demand association network, matching descriptive words, sentence templates, and text structure fragments are retrieved from the text corpus.
[0049] For the user demand node of "fashionable style", the retrieved descriptive terms include "trendy design", "fashionable tailoring", "unique pattern", and "unique neckline"; the sentence templates include "This dress adopts the currently popular [pattern] design, combined with [tailoring] technology, showing a fashionable atmosphere" and "The unique [neckline] design makes this dress stand out among many styles"; the chapter structure segments include "first introducing the overall style of the dress, then describing its design highlights in detail, and finally explaining the suitable wearing scenarios".
[0050] The requirements are categorized into three levels based on the level of detail and depth of the description: Basic, Intermediate, and Professional. Basic-level text descriptions primarily provide a simple style introduction, categorized as T1, such as T1-“This dress is stylish and suitable for everyday wear.” Intermediate-level text descriptions delve into more design details, categorized as T2, such as T2-“This dress features a fashionable A-line silhouette with a polka dot pattern and a round neckline, exuding elegance and style.” Professional-level text descriptions analyze fashion trends, categorized as T3, such as T3-“This dress follows the season's trends; its irregular hem design is inspired by the latest fashion week elements, and the contrasting trim showcases a unique fashion attitude.”
[0051] Step S122: Based on the content element nodes in the demand association network, after filtering image fragments, video clips, and display models in the visual resource library that correspond to the core attribute words and feature descriptions of the content, feature extraction is performed to generate visual display units containing content feature tags. Each visual display unit is associated with a corresponding content element node.
[0052] The visual resource library contains a large amount of visual materials related to dresses, such as product images from different angles, videos showing how they look when worn, close-ups of fabric details, and 3D display models. Based on the core attribute words and characteristic descriptions of the content element nodes, corresponding visual resources are selected.
[0053] For the "Dress Fabric Material" content parameter node, we filter out close-up image clips of the fabric (such as the texture of cotton, the lightness of chiffon), video clips showing the stretching of the fabric (demonstrating the elasticity of the fabric), and 3D fabric display models (allowing for 360-degree viewing of fabric details). We then extract features from these visual resources: for image clips, we extract features such as color (fabric color), texture (fabric surface texture), and gloss (fabric reflectivity); for video clips, we extract dynamic features (such as the fabric's movement) and keyframe image features; and for the display model, we extract structural features (whether it allows for interactive viewing of details) and material simulation features (such as simulating the tactile feedback visual effects of the fabric).
[0054] Generate visual display units containing content feature tags and associate them with corresponding content element nodes. For example, visual display unit V1 - image clip, with content feature tags of "cotton, light color, clear texture", is associated with the content parameter node "dress fabric material"; visual display unit V2 - video clip, with content feature tags of "chiffon fabric, light and flowing, good breathability", is associated with the content parameter node "dress fabric material"; visual display unit V3 - display model, with content feature tags of "knitted fabric, good elasticity, 360-degree view", is associated with the content parameter node "dress fabric material".
[0055] Step S123: Analyze the association edges and association strength parameters between nodes in the demand association network, extract the semantic correspondence rules between the user demand node and the content element node, the presentation logic rules between the content element node and the scene form node, and the matching rules between the user demand node and the scene form node, and then convert the semantic correspondence rules, the presentation logic rules and the matching rules into semantic association units, each of which includes a rule type description and a strength threshold range.
[0056] The association edges and association strength parameters between nodes in the demand association network are analyzed to extract various rules. For example, the semantic correspondence rules between user demand nodes and content element nodes are analyzed. The association strength between the content parameter nodes "material comfort demand" and "dress fabric material" is high. The extracted semantic correspondence rule is "when a user has a material comfort demand, prioritize associating content elements related to dress fabric material." The strength threshold range is determined based on the first association strength parameter; that is, this rule applies when the first association strength parameter is within this strength threshold range.
[0057] The presentation logic rules between content element nodes and scene form nodes are as follows: For example, the content parameter node "dress size parameter" has a high correlation with the "product details page scene" node. The extracted presentation logic rule is "in the product details page scene, the dress size parameter content element is presented in a table combined with a diagram first". The intensity threshold range is determined based on the second correlation intensity parameter.
[0058] Matching rules between user demand nodes and scenario format nodes. For example, the correlation strength between the "cost-effectiveness demand" and the "promotion activity page scenario" node is high. The extracted matching rule is "cost-effectiveness demand and promotion activity page scenario have a high degree of matching, and it is suitable to highlight price advantages and preferential information in this scenario". The strength threshold range is determined according to the third correlation strength parameter.
[0059] These rules are transformed into semantic association units. For example, semantic association unit R1 is described as "User Needs - Content Element Semantic Correspondence Rule: When a user has a need for material comfort, content elements related to the dress fabric material are prioritized for association," and the strength threshold range is the applicable range of the corresponding first association strength parameter. Semantic association unit R2 is described as "Content Element - Scene Presentation Logic Rule: In the product details page scenario, the dress size parameter content element is prioritized to be presented in a table combined with a diagram," and the strength threshold range is the applicable range of the corresponding second association strength parameter. Semantic association unit R3 is described as "User Needs - Scene Form Matching Rule: The matching degree between cost-effectiveness needs and promotional activity page scenarios is high, making it suitable to highlight price advantages and discount information in this scenario," and the strength threshold range is the applicable range of the corresponding third association strength parameter.
[0060] Step S124: Establish a mapping relationship table between the text description unit and the user demand node, content element node, and scene form node, and determine the node and mapping weight corresponding to each text description unit.
[0061] A mapping relationship is established between each text description unit and user requirement nodes, content element nodes, and scene form nodes. By analyzing the content of the text description unit, the degree of its correlation with each node is determined, and then the mapping weight is determined. The mapping weight reflects the closeness of the correlation between the text description unit and the corresponding node, and ranges from 0 to 1.
[0062] For example, text description unit T2 - "This dress features a stylish A-line silhouette with a polka dot pattern and a round neckline, exuding elegance and fashion" - is closely related to the user demand node "Style and Fashion Needs," and its mapping weight is set to a high value. It is also closely related to the content parameter nodes "Dress Style Design" and "Pattern Elements," and its mapping weight is set to a high value. Furthermore, it has some connection to the scenario format nodes "Short Video Promotion Scenarios" and "Product Details Page Scenarios," and its mapping weight is set to corresponding values based on their suitability. These mapping relationships are organized into a mapping relationship table, clearly recording the node and mapping weight corresponding to each text description unit.
[0063] Step S125: Establish a mapping relationship table between the visual display unit and the user demand node, content element node, and scene form node, and determine the node and mapping weight corresponding to each visual display unit.
[0064] Similarly, a mapping relationship is established between each visual display unit and user demand nodes, content element nodes, and scene form nodes. Based on the content feature tags of the visual display unit, the degree of association between it and each node is analyzed to determine the mapping weight.
[0065] For example, the visual display unit V1 - image fragment (content feature tags: "cotton, light color, clear texture") is closely related to the "material comfort requirement" user demand node, so its mapping weight is set to a high value; it is also closely related to the "dress fabric material" content parameter node, so its mapping weight is set to a high value; it is also relatively closely related to the "product details page scenario" node, so its mapping weight is set to a high value; while its relationship with the "short video promotion scenario" node is relatively weak, so its mapping weight is set to a low value. These mapping relationships are compiled into a mapping relationship table for later use.
[0066] Step S126: Establish a mapping table between the semantic association unit and the association edge in the demand association network, and determine the association edge and association strength parameter range corresponding to each semantic association unit.
[0067] Analyze the rules represented by each semantic association unit, determine the corresponding association edges in the demand association network, and record the association strength parameter range of the association edge.
[0068] For example, semantic association unit R1 corresponds to the association edge between the user demand node "material comfort requirement" and the content parameter node "dress fabric material," and its association strength parameter range is the applicable range of the first association strength parameter of this association edge; semantic association unit R2 corresponds to the association edge between the content parameter node "dress size parameter" and the "product details page scenario" node, and its association strength parameter range is the applicable range of the second association strength parameter of this association edge; semantic association unit R3 corresponds to the association edge between the user demand node "cost-effectiveness requirement" and the "promotional activity page scenario" node, and its association strength parameter range is the applicable range of the third association strength parameter of this association edge. These mapping relationships are organized into a mapping relationship table, clearly defining the association edge and association strength parameter range corresponding to each semantic association unit.
[0069] Step S127: Integrate the text description unit, visual display unit, semantic association unit and their mapping relationship table to construct a multimodal element coupling pool.
[0070] The generated text description units, visual display units, semantic association units, and their respective mapping relationship tables are integrated to form a unified multimodal element coupling pool. Within this coupling pool, various element units are interconnected, enabling rapid retrieval and invocation of corresponding element units based on the nodes and relationships within the required association network.
[0071] Step S130: Invoke the multimodal large model to perform contextualized mapping processing on the multimodal element coupling pool to generate an initial content framework set that matches the demand-related network.
[0072] Step S131: Input the demand association network into the context parsing layer of the multimodal large model, and use the node feature extraction algorithm to identify the core demand keywords of the user demand node, the key attribute features of the content element node, and the preference type of the scene form node to generate a context feature vector.
[0073] In this embodiment, the context parsing layer of the multimodal large model is used to perform deep analysis of the input demand association network to extract key features. First, feature extraction is performed on user demand nodes, identifying core demand keywords from the demand type labels and demand descriptions of each node. For example, "fashionable style demand" node extracts "style" and "fashionable," while "comfortable material demand" node extracts "material" and "comfortable." Next, content element nodes are processed, extracting key attribute features from main content category, sub-content attribute, and content parameter nodes, including core attributes such as "style," "fabric," and "size." Then, scenario format nodes are analyzed to determine types such as "short video promotion scenario" preferring dynamic displays and "product details page scenario" preferring a combination of text and images. After standardizing these features into computable feature values, target feature vectors are generated by assigning weights according to their importance, and finally mapped to a high-dimensional semantic space to obtain context feature vectors.
[0074] Step S1311: Feature extraction is performed on the user demand nodes in the demand association network. The demand type label, demand description and associated content element node information of each user demand node are extracted. The core demand keywords are mined through topic modeling algorithm. The core demand keywords are used to represent the demand points that users are most concerned about.
[0075] A comprehensive scan of user demand nodes in the demand association network is performed to extract basic information for each node. For example, for the "fashionable style demand" node, its demand type tag "fashionable style," demand description "consumers hope that dresses are novel in design and in line with current trends," and related content element node information such as "dress style design" and "pattern elements" are extracted. The LDA topic model algorithm is used to mine the topics of this information. By calculating the co-occurrence probability and topic distribution of words in the text, core demand keywords such as "style," "fashion," "trend," and "design" are identified to accurately locate the user's core demand for style. For the "material comfort demand" node, after topic model analysis, core demand keywords such as "material," "comfort," "breathable," and "skin-friendly" are extracted.
[0076] Step S1312: Extract features from the content element nodes in the demand association network, extract the category labels of the main content category nodes, the attribute descriptions of the sub-content attribute nodes, and the specific parameter values of the content parameter nodes, and use a feature selection algorithm to filter out the key attribute features that affect content generation. The key attribute features include the function, material, and specifications of the content.
[0077] Feature extraction was performed on the content element node system. The category tag "dress" was extracted from the main content category node "Summer Dress"; attribute descriptions such as "fabric material description" and "pattern design features" were extracted from sub-content attribute nodes such as "fabric attribute" and "pattern attribute"; and specific parameter values such as "80% cotton, 20% polyester fiber," "A-line fitted waist," and "length 85cm" were extracted from content parameter nodes. The ReliefF feature selection algorithm was used to calculate the weight of each feature, and the top 30% of features by weight were selected as key attribute features. For example, in the fabric attribute, "fabric composition," "breathability rating," and "skin-friendliness" were selected as key attribute features due to their significant impact on content generation; in the pattern attribute, "overall pattern type," "waist design," and "skirt shape" became key attribute features. These features directly determine the focus of content generation.
[0078] Step S1313: Extract features from the scene format nodes in the demand association network, extract feature parameters of the scene format node presentation mode, structure type and modal combination mode, and determine the scene format preference type through cluster analysis. The preference type includes detail page format, short video format and clip format.
[0079] The core feature parameters of nodes in each scenario were extracted. For the "product details page scenario," the presentation parameters were "multiple images + text description," the structure type parameter was "hierarchical and progressive," and the modal combination parameter was "text + static image + dynamic image." For the "short video promotion scenario," the presentation parameters were "dynamic video + background music," the structure type parameter was "highlight-first," and the modal combination parameter was "video + audio + text subtitles." K-means clustering was used to perform cluster analysis on these feature parameters, calculating the Euclidean distance between the feature parameters of nodes in different scenarios, and grouping nodes with closer distances into one category. Ultimately, three preference types were clustered: details page format (emphasizing information completeness), short video format (emphasizing visual impact), and live stream segment format (highlighting real-time interactivity). Each preference type corresponds to a set of typical feature parameter combinations.
[0080] Step S1314: Standardize the core requirement keywords, key attribute features, and scenario form preference types to generate computable feature values.
[0081] Unstructured text features are standardized by mapping the core requirement keyword "style" to the value 1 and "fashion" to the value 2; key attribute features such as "slim-fit A-line silhouette" are encoded as the value 101 and "cotton-linen blend fabric" as the value 203. One-hot encoding is used for scene format preference types: details page format is represented as [1, 0, 0], short video format as [0, 1, 0], and live stream segment format as [0, 0, 1]. For numerical parameters such as "breathability level," Min-Max standardization is used to transform them to the [0, 1] range, giving different types of features a unified computable form and laying the foundation for subsequent weighted processing.
[0082] Step S1315: Set weight parameters according to the importance of the core requirement keywords, key attribute features and scenario form preference types in content generation, perform weighted processing on the feature values, and generate target feature vector.
[0083] The weight parameters of each feature were determined using the Analytic Hierarchy Process (AHP). Core demand keywords, directly reflecting user needs, were weighted at 0.4; key attribute features, as fundamental elements of content generation, were weighted at 0.35; and scenario format preference types, influencing content presentation, were weighted at 0.25. The standardized feature values were then weighted. For example, in a demand association network, the core demand keyword feature vector was [0.8, 0.6, 0.9], the key attribute feature vector was [0.7, 0.85, 0.75], and the scenario format preference type feature vector was [1, 0, 0]. After weighted calculation, the target feature vector was obtained as: [0.8×0.4+0.7×0.35+1×0.25, 0.6×0.4+0.85×0.35+0×0.25, 0.9×0.4+0.75×0.35+0×0.25], forming a target vector containing weighted features across all dimensions.
[0084] Step S1316: Map the target feature vector to a high-dimensional semantic space to generate a context feature vector.
[0085] Using a pre-trained BERT model as a mapping tool, the target feature vector is input into the model's semantic encoding layer. Feature transformation and semantic enhancement are performed through a multi-layer Transformer structure. The model maps the original feature vector from a low-dimensional space to a 768-dimensional high-dimensional semantic space, preserving the semantic relationships between features during the mapping process. For example, the intrinsic connection between "breathable fabric" and "comfort requirements" is reflected in the high-dimensional space through vector distance. The final generated contextual feature vector not only contains the numerical information of the original features but also rich semantic association information, capable of representing the overall context of the demand association network.
[0086] Step S132: Input the text description unit, visual display unit and semantic association unit in the multimodal element coupling pool into the element encoding layer of the multimodal large model, and perform text encoding, visual encoding and semantic rule encoding processing respectively to generate an element feature vector set composed of text feature vector, visual feature vector and semantic association feature vector.
[0087] The feature encoding layer of the multimodal large model encodes various feature units in the multimodal feature coupling pool. For text description units, a pre-trained language model is used for text encoding, transforming the text content into a fixed-dimensional text feature vector that captures the semantic information of the text. For example, after encoding text description unit T2, the generated text feature vector can reflect semantic information such as "A-line shape," "polka dot pattern," and "round neck design."
[0088] For the visual display unit, visual encoding is performed using a convolutional neural network to extract deep features of the visual resources and generate visual feature vectors. For example, encoding image segments of visual display unit V1 generates visual feature vectors that reflect visual features such as the texture and color of the cotton material.
[0089] For semantic association units, a rule-based encoding method is used to transform the rule type description and intensity threshold range into a semantic association feature vector. This semantic association feature vector can represent the core meaning and applicable scope of the rule. For example, encoding the semantic association unit R1 generates a semantic association feature vector that reflects the semantic correspondence and intensity threshold range between "material comfort requirements" and "dress fabric material".
[0090] The generated text feature vectors, visual feature vectors, and semantic association feature vectors are integrated together to form a set of element feature vectors.
[0091] Step S133: Calculate the matching degree between each feature vector and the context feature vector using the mapping matching layer of the multimodal large model based on the mapping relationship table of the multimodal feature coupling pool, and filter out candidate feature vectors whose matching degree meets the preset conditions.
[0092] The main function of the mapping and matching layer is to match feature vectors of elements with feature vectors of context. Based on the mapping relationship table in the multimodal feature coupling pool, the association between each feature vector and the nodes in the demand association network is clarified. Then, the matching degree is determined by calculating the similarity between the feature vectors of elements and the feature vectors of context; the higher the similarity, the higher the matching degree.
[0093] For example, the cosine similarity between text feature vectors and context feature vectors can be calculated. The higher the cosine similarity value, the higher the matching degree between the two. Similarly, a suitable similarity calculation method can be used to obtain the matching degree for visual feature vectors and context feature vectors.
[0094] In this embodiment, a preset condition is set, such as a matching degree greater than or equal to a certain threshold, and feature vectors that meet the preset condition are selected as candidate feature vectors. The feature units corresponding to these candidate feature vectors have a high matching degree with the context of the demand association network, and can serve as the basis for subsequent content framework construction.
[0095] Step S134: Based on the association strength parameters of the association edges in the demand association network, the candidate element feature vectors are weighted and sorted to determine the priority of the text description unit, visual display unit and semantic association unit in content generation, and the priority ranking result is obtained.
[0096] Referring to the association strength parameter of the association edges in the demand association network, the feature vectors of candidate elements are weighted. The higher the association strength parameter, the greater the weight of the corresponding feature vector in the ranking. After weighted calculation, the feature vectors of candidate elements are ranked to determine the priority of their corresponding text description units, visual display units, and semantic association units in content generation.
[0097] For example, the feature vectors of candidate elements corresponding to edges with higher association strength parameters have higher priority; conversely, those with lower association strength parameters have lower priority. This weighted sorting yields a priority ranking result, clarifying the order and importance of various element units when generating the content framework.
[0098] Step S135: Based on the priority ranking result, the framework construction layer of the multimodal large model combines the text feature vector, visual feature vector, and semantic association feature vector according to the semantic logic of content presentation to generate multiple initial content frames. Each initial content frame contains an element combination sequence and a framework structure description.
[0099] The framework construction layer combines different types of feature vectors according to the semantic logic of content presentation, based on the priority ranking results. First, it determines the main logic of content presentation, such as following the logical order of "attracting attention - introducing core selling points - detailed explanation - promoting purchase". Then, it selects text feature vectors, visual feature vectors, and semantic association feature vectors according to their priority from high to low, and combines them into a content framework with a certain structure.
[0100] Each initial content frame contains an element combination sequence and a frame structure description. The element combination sequence records the order of the selected text description units, visual display units, and semantic association units; the frame structure description describes the overall structure of the content frame, such as how many parts it is divided into, the main content of each part, and its presentation method. For example, the element combination sequence of an initial content frame might be "Visual Display Unit V2 - Text Description Unit T2 - Semantic Association Unit R2 - Visual Display Unit V1 - Text Description Unit T1", and the frame structure description might be "First, a video is shown to demonstrate the wearing effect of the dress, then its fabric material and comfort are introduced, followed by the semantic association unit explaining the matching relationship between the wearing effect and the fabric material, then detailed design images of the dress are shown, and finally, the unique features of the detailed design are explained."
[0101] For example, another initial content framework's element combination sequence might be "text description unit T3 - visual display unit V4 - semantic association unit R1 - text description unit T5 - visual display unit V3". The framework structure is described as follows: "First, highlight the dress's fashionable design style with text, then show overall appearance images from different angles, link the design style with appearance features through semantic association units, then introduce matching suggestions with text, and finally show the overall effect image after matching."
[0102] During the generation of these initial content frameworks, the framework construction layer ensures semantic coherence between different element units, avoiding logical gaps. For example, when combining visual display units with textual description units, it guarantees that the information presented by the visual content corresponds to the content described in the text, preventing situations where the visual display shows style A of a dress while the text describes style B. Simultaneously, semantically related units are interspersed in appropriate positions, serving to connect and explain, making the logic of the entire content framework clearer.
[0103] Step S136: Perform integrity verification on each initial content framework to check whether it covers the core user demand nodes, key content element nodes, and important scene form nodes in the demand association network. If there are any missing elements, select the corresponding element feature vector from the multimodal element coupling pool to supplement them.
[0104] Each generated initial content framework needs to undergo integrity verification. First, identify the core user demand nodes in the demand association network, such as "fashionable style demand," "comfortable fabric demand," and "reasonable price demand"; key content element nodes include "dress style design," "fabric material," and "price range"; and important scenario format nodes include "short video promotion scenario" and "product details page scenario."
[0105] During validation, each initial content framework is checked to ensure it covers these core nodes. For example, for a given initial content framework, it is checked whether it contains elements related to "style and fashion requirements." If the framework only describes fabric materials and does not cover style design, meaning it lacks the core user requirement node for "style and fashion requirements" and the corresponding key content element node for "dress style design," then feature vectors of text description units, visual display units, and semantic association units related to style design need to be selected from the multimodal element coupling pool to supplement it.
[0106] During the supplementation process, it is essential to ensure that the newly added feature vectors are semantically consistent with the feature vectors in the original framework, without disrupting the original logical structure. For example, if the original framework focuses on introducing fabric comfort, when supplementing with content related to style design, the two should be connected through appropriate semantic association units to explain the advantages of combining comfortable fabrics with fashionable styles.
[0107] Step S137: The initial content frameworks that have passed the integrity check are integrated into an initial content framework set, and each initial content framework corresponds to a matching score with the requirement-related network in the initial content framework set.
[0108] All initial content frameworks that pass the integrity check are integrated together to form an initial content framework set. To facilitate subsequent collaborative optimization of elements, each initial content framework needs to be assigned a matching score to the requirement's relational network.
[0109] The matching score is mainly calculated based on the degree of association between the element units included in the initial content framework and each node in the demand association network, as well as the association strength parameter of each association edge. For example, if an initial content framework covers a large number of core user demand nodes, key content element nodes, and important scenario form nodes, and the association strength parameter between these nodes is high, then the matching score of the framework is relatively high; conversely, if it covers fewer nodes or the association strength parameter is low, then the matching score is relatively low.
[0110] The matching score of each initial content framework is recorded in the initial content framework set. For example, the matching score of initial content framework F1 is 0.85, and the matching score of initial content framework F2 is 0.72. These matching scores will intuitively reflect the degree of fit between each framework and the network of requirements.
[0111] Step S140: Perform element collaborative optimization processing on the initial content framework set to obtain an optimized content framework that meets the content element node constraints.
[0112] Step S141: Analyze the content element node constraints in the demand association network to determine the presentation requirements of the core content attributes, the matching standards of content specification parameters, and the rules for creating content style characteristics.
[0113] Analyzing the constraints of content element nodes in the demand-related network is a prerequisite for collaborative optimization of elements. These content element node constraints are formulated based on the actual needs of apparel product promotion and are used to ensure that the generated content framework conforms to the core attributes and promotional requirements of the product.
[0114] The presentation requirements for the core attributes of the content mainly involve the display methods of the dress's core attributes such as style, fabric, and color. For example, for the style attribute, it is required to clearly display details such as the dress's collar type, sleeve type, and skirt design; for the fabric attribute, it is required to explain the fabric's composition, texture, and care instructions.
[0115] The matching standard for content specifications mainly refers to the need for parameters such as dress size and style to match the characteristics of the target consumer group. For example, dresses targeting young women should cover common sizes such as small, medium, and large, and the style design should conform to the body characteristics and aesthetic preferences of this group.
[0116] The rules for creating content style characteristics are to ensure that the overall style of the content framework is consistent with the product positioning and promotional scenarios. For example, if the product positioning of a dress is high-end fashion, then the style of the content framework should lean towards elegance and sophistication, the language should be concise and elegant, and the visual presentation should emphasize texture and detail.
[0117] Step S142: Extract the text description unit, visual display unit and semantic association unit contained in each initial content frame from the initial content frame set, and analyze the matching degree between each element unit and the content element node constraint conditions. The matching degree is calculated by the fit between the content of the element unit and the constraint conditions.
[0118] The text description units, visual display units, and semantic association units contained in each initial content framework are extracted one by one from the initial content framework set. Then, the content of the above element units is compared with the constraints of the content element nodes to analyze their matching degree.
[0119] The degree of matching is determined by the fit between the content of the element unit and the constraints. For example, for a text description unit, if its content details the fabric composition and care methods of the dress, which highly matches the presentation requirements of the fabric attribute in the core content attributes, then the degree of matching of the text description unit is high; if its content only briefly mentions the fabric without involving specific composition and care methods, then the degree of matching is low.
[0120] For visual display units, if they clearly show the style details of the dress, such as the collar and sleeves, and meet the presentation requirements of the style attributes, the matching degree is relatively high; if the visual content is blurry and the style details cannot be clearly distinguished, the matching degree is relatively low.
[0121] For semantically related units, if they can accurately associate different attributes of the dress and conform to the rules for creating content style features, the matching degree is high; if their association logic is chaotic or does not match the overall style, the matching degree is low.
[0122] Step S143: Replace the element units with insufficient matching degree by selecting element units that better meet the content element node constraints from the multimodal element coupling pool and replacing them while maintaining the semantic association between the text description unit and the visual display unit during the replacement process.
[0123] If the matching degree of a certain element unit is lower than a preset threshold, it needs to be replaced. For example, if a certain text description unit does not describe the fabric composition of a dress in enough detail, only mentioning "comfortable fabric" without specifying the specific components, it does not match the presentation requirements of the core attributes of the content. In this case, a text description unit that can describe the fabric composition in detail (such as the proportion of cotton and polyester fibers) needs to be selected from the multimodal element coupling pool for replacement.
[0124] When replacing visual display units, if the color of the dress displayed in a certain visual display unit deviates from the actual product and does not meet the matching standard of the content specification parameters, then a visual display unit that can accurately present the product color should be selected for replacement.
[0125] During the replacement process, the semantic relevance between the text description unit and the visual display unit must be maintained. For example, if the replaced text description unit emphasizes the dress's "chiffon fabric, light and breathable," then the corresponding visual display unit should show images or videos that reflect the light texture of the chiffon fabric, such as the dynamic effect of the dress flowing in the wind, ensuring that the information conveyed by the text description and the visual display are consistent.
[0126] Step S144: Adjust the combination order and association strength of the text description unit, visual display unit and semantic association unit so that the content logic of the text description unit is consistent with the presentation logic of the visual display unit, and the semantic association unit can accurately reflect the logical relationship between the text and the visual.
[0127] After replacing the element units, the combination order and correlation strength of each element unit need to be adjusted. When adjusting the combination order, it is important to ensure that the content logic of the text description unit is consistent with the presentation logic of the visual display unit. For example, if the text description unit describes the fabric composition in the order of "fabric composition - style design - wearing effect", the visual display unit should also be presented in the order of showing fabric details - showing the overall style - showing the wearing effect.
[0128] When adjusting the association strength, enhance the association strength between element units that highlight the core selling points of the dress. For example, if "silk fabric" is one of the core selling points of the dress, then the association strength between the text description unit describing the silk fabric and the visual display unit showcasing the texture of the silk fabric should be appropriately increased, and the corresponding semantic association units should also emphasize the relationship between the two.
[0129] The semantic association unit should be adjusted to accurately reflect the logical relationship between the text and the visuals. For example, when the text description unit mentions "the waist-cinching design of the dress can accentuate the figure," and the visual display unit shows a detailed picture of the waist-cinching design, the semantic association unit should clearly state "the waist-cinching design shown in the picture echoes the figure-enhancing effect described in the text, and this design achieves the figure-enhancing effect through XX method."
[0130] Step S145: Perform element synergy evaluation on the adjusted content framework, and calculate the coherence score between the text description units, the style consistency score between the visual display units, and the semantic matching score between the text description units and the visual display units.
[0131] Step S1451: Extract the text description unit sequence in the adjusted content framework, analyze the lexical cohesion information, grammatical association information and semantic coherence information between adjacent text description units, calculate the coherence score through a pre-trained language model, and use the coherence score as the coherence score between text description units.
[0132] Extract the text description unit sequence from the adjusted content framework, such as "T1-T2-T3-T4". Analyze the lexical connection between T1 and T2. If T1 mentions "chiffon fabric" and T2 uses "this fabric" to refer to it, the lexical connection is smooth. If T1 discusses fabric and T2 suddenly shifts to size without a proper transition, then there is a problem with the lexical connection.
[0133] Analyze grammatical coherence information to check whether the sentence structures of adjacent text description units are consistent and whether there are any grammatical conflicts. For example, if T1 uses a declarative sentence to describe the fabric and T2 uses an interrogative sentence to introduce the style, the grammatical coherence is good if the transition is natural; if the sentence structure is chaotic and lacks logical connection, the grammatical coherence is poor.
[0134] Analyze semantic coherence information to determine whether the semantics expressed by adjacent textual descriptive units are related and whether they can form a coherent semantic flow. For example, if T1 describes the fabric as lightweight, and T2 then describes the fabric as suitable for summer wear, the semantics are coherent; if T1 says the fabric is heavy, but T2 says it is suitable for summer, the semantics are incoherent.
[0135] Input this information into the pre-trained language model, and the model will calculate a coherence score based on the language rules and semantic associations it has learned during training. This coherence score is the coherence score between text description units.
[0136] Step S1452: Extract the sequence of visual display units in the adjusted content framework, analyze the consistency of color tone, composition style and element layout of the visual display units, calculate the style similarity through a visual feature comparison algorithm, and use the style similarity as the style consistency score between visual display units.
[0137] Extract the sequence of visual display units in the adjusted content framework, such as "V1-V2-V3-V4". Analyze the color tone. If all visual display units are mainly in soft, light colors, the color tone is consistent. If V1 is a cool color and V2 suddenly becomes a warm color without a proper transition, the color tone is inconsistent.
[0138] Analyzing the composition style, if all visual display units use a centered composition to highlight the main body of the dress, then the composition style is consistent; if V1 uses a panoramic composition and V2 suddenly uses a close-up with a large difference in composition, then the composition style is inconsistent.
[0139] Analyzing the element layout, if the placement of the dresses and the matching of background elements in all visual display units are consistent, such as all using a simple solid color background, the element layout is consistent; if the background of V1 is complex and the background of V2 is simple and unrelated, the element layout is inconsistent.
[0140] By using a visual feature comparison algorithm, these visual features are compared and the style similarity is calculated. This style similarity is the style consistency score between visual display units.
[0141] Step S1453: Establish the correspondence between the text description unit and the visual display unit, with each text description unit corresponding to one or more visual display units, and extract the core semantics of the text description unit and the visual semantics of the visual display unit.
[0142] Establish a correspondence between text description units and visual display units. For example, text description unit T1 (introducing the light texture of chiffon fabric) corresponds to visual display units V1 (close-up image of chiffon fabric) and V2 (video clip of a dress flowing); text description unit T2 (introducing the waist-cinching design) corresponds to visual display unit V3 (detailed image of the waist-cinching area).
[0143] Extracting the core semantics of the text description units, the core semantics of T1 is "lightweight chiffon fabric"; the core semantics of T2 is "waist-cinching design to flatter the figure".
[0144] The visual semantics of the visual display units are extracted. The visual semantics of V1 is "the delicate texture and translucency of chiffon fabric"; the visual semantics of V2 is "the lightness of the dress during movement"; and the visual semantics of V3 is "the design of the waistline and the optimization of body proportions".
[0145] Step S1454: Calculate the semantic distance between the core semantics and the visual semantics. The semantic distance is calculated using Euclidean distance in the semantic vector space.
[0146] The core semantics of the text description unit and the visual semantics of the visual display unit are transformed into semantic vectors and mapped to the same semantic vector space. For example, the core semantics of "lightweight chiffon fabric" is transformed into vector A, the visual semantics of "delicate texture and translucency of chiffon fabric" is transformed into vector B, and the visual semantics of "the lightness of the dress during movement" is transformed into vector C.
[0147] Calculate the Euclidean distance between vector A and vector B, and the Euclidean distance between vector A and vector C. These distances are the semantic distances between the core semantics and the visual semantics.
[0148] Step S1455: The semantic distance is converted into a semantic matching score, and the semantic matching score is negatively correlated with the semantic distance.
[0149] The smaller the semantic distance, the higher the degree of matching between the core semantics and the visual semantics, and the higher the corresponding semantic matching score; conversely, the larger the semantic distance, the lower the semantic matching score. For example, if the semantic distance between vector A and vector B is small, the corresponding semantic matching score is high; if the distance between vector A and another visual semantic vector is large, its semantic matching score is low.
[0150] Step S1456: The coherence score between text description units, the style consistency score between visual display units, and the semantic matching score between text description units and visual display units are weighted according to preset weight parameters to obtain a comprehensive score for element synergy evaluation.
[0151] Preset the weight parameters for coherence score, style consistency score and semantic matching score. For example, the weight of coherence score is 0.3, the weight of style consistency score is 0.3, and the weight of semantic matching score is 0.4.
[0152] Assuming that the coherence score of a certain adjusted content framework is 0.8, the style consistency score is 0.7, and the semantic matching score is 0.9, then the comprehensive score = 0.8×0.3 + 0.7×0.3 + 0.9×0.4. The comprehensive score obtained by calculation is the comprehensive score of the element synergy assessment.
[0153] Step S146: If the element synergy assessment result meets the preset threshold, the adjusted content framework is used as a candidate optimization framework; if not, the replacement and adjustment operations are repeated until the element synergy assessment result meets the standard.
[0154] Set a preset threshold for the overall score of the element synergy assessment. If the overall score of the adjusted content framework reaches or exceeds the threshold, it indicates that the element synergy of the framework is good, and it is used as a candidate optimization framework. If the overall score is lower than the threshold, the element units need to be replaced and adjusted again, and the element synergy assessment needs to be carried out again until the overall score meets the standard.
[0155] For example, if the preset threshold is 0.75, and the overall score of a certain adjusted content framework is 0.82, it meets the threshold requirement and becomes a candidate optimization framework; if the overall score of another framework is 0.7, it needs to be replaced and adjusted again, such as replacing some visual display units to improve style consistency, or adjusting the order of text description units to enhance coherence, and re-evaluating until the score reaches 0.75 or above.
[0156] Step S147: Select the framework with the highest matching degree with the content element node constraint conditions from the candidate optimization frameworks as the optimized content framework. The optimized content framework includes a text description sequence, a visual display sequence, and a set of semantic association rules that have been collaboratively optimized.
[0157] The matching degree between all candidate optimization frameworks and content element node constraints is re-evaluated. The evaluation dimensions include the completeness of core attribute presentation, the accuracy of specification parameter matching, and the consistency of style features.
[0158] The candidate optimization framework that performs best across these dimensions, i.e., has the highest matching degree, is selected as the final optimized content framework. The text description sequence, visual display sequence, and semantic association rule set in this optimized content framework have all been collaboratively optimized and can well meet the constraints of the content element nodes.
[0159] Step S150: Generate a target content set based on the optimized content framework, and push the target content set to the corresponding content distribution node to complete the content dissemination operation.
[0160] Step S151: parse the text description sequence in the optimized content framework, combine each text description unit into complete text content according to semantic logic, and match the language style of the text content with the user demand nodes in the demand association network.
[0161] Step S1511: Extract each text description unit from the text description sequence in the optimized content framework, and parse the hierarchical identifier, text content, and corresponding user requirement node information of each text description unit.
[0162] Extract each text description unit from the optimized content framework's text description sequence, such as "T1-T2-T3-T4". Parse the hierarchical identifier of each unit, where T1 is identified as "Basic Level", T2 as "Advanced Level", T3 as "Basic Level", and T4 as "Advanced Level".
[0163] Analyzing the text content, T1: "This dress is made of high-quality chiffon fabric"; T2: "Chiffon fabric is lightweight and breathable, making it suitable for summer wear"; T3: "The dress features a waist-cinching design"; T4: "The waist-cinching design effectively flatters the figure and highlights an elegant temperament."
[0164] The corresponding user requirement node information is analyzed. T1 and T2 correspond to "fabric comfort requirements"; T3 and T4 correspond to "style requirements".
[0165] Step S1512: Determine the combination order of the text description units based on the semantic association between them.
[0166] Analyzing the semantic relationships between the text description units, T1 introduces the fabric material, and T2 further explains the advantages of this fabric in the dress design, such as "breathable cotton-linen fabric can better highlight the lightness of the wearer in a fitted A-line dress." Therefore, T1 and T2 are progressively related, and T1 should be placed before T2. T3 describes the wearing scenarios of the dress, and its content needs to be based on the basic information of the fabric and the design, such as "a breathable cotton-linen fitted A-line dress is suitable for commuting with a suit jacket, and for casual wear with canvas shoes." Therefore, T3 should follow T1 and T2. T4 involves washing and care methods, which are directly related to the fabric material, such as "cotton-linen fabric should be washed gently in cold water and avoid direct sunlight." Therefore, T4 should follow closely after T1, which introduces the fabric, forming a logical chain of "fabric characteristics - washing methods - design advantages - wearing scenarios."
[0167] Based on the semantic association analysis described above, the combination order of the text description units is determined to be T1-T4-T2-T3. This order not only conforms to the consumer's cognitive logic from basic attributes to usage details, but also ensures that the text content remains coherent in conveying information, avoiding information jumps or logical breaks.
[0168] Step S1513: Connect the text description units in the combination process by adding preset connecting words, transition sentences or paragraphs.
[0169] Between T1 and T4, add the transitional sentence "Having understood the fabric characteristics of this dress, the correct washing method can better maintain its texture," naturally transitioning the fabric introduction to washing and care instructions. Between T4 and T2, use the conjunction "In addition," paired with the transitional sentence "Besides the comfort provided by the fabric, its design is also worth noting," achieving a smooth transition from washing information to the advantages of the design. Between T2 and T3, use the conjunction "Based on the above design and fabric" to introduce the transitional sentence "This dress can showcase its unique charm in different scenarios," closely linking the design information with the description of the wearing scenarios.
[0170] These seamless transitions avoid the awkwardness of combining text units while strengthening the logical connections between different parts, making the entire text read smoothly and naturally, and facilitating consumer understanding and information reception.
[0171] Step S1514: Analyze the demand type labels and demand descriptions of user demand nodes in the demand association network, and determine the corresponding language style features, including formality, emotional tendency and conciseness.
[0172] In the demand association network, the demand description for the "fabric comfort demand" node is "Consumers want dresses made of skin-friendly, breathable fabric suitable for long-term wear," with the corresponding demand type label being "practical demand." This type of demand emphasizes the accuracy and practicality of information; therefore, the language style should be formal and objective, avoiding excessive embellishment. The demand description for the "slimming fit demand" node is "Consumers seek dresses that flatter their figure and highlight their best features," with the demand type label being "aesthetic demand." The language style can appropriately incorporate positive emotional inclinations, such as using words like "elegantly accentuates" and "cleverly enhances" to increase appeal. The demand description for the "scene adaptability demand" node is "Consumers need dresses that can meet the needs of various occasions and improve dressing convenience," with the demand type label being "multifunctional demand." The language style should be concise and clear, using refined language to explain dressing suggestions for different scenarios, allowing consumers to quickly obtain key information.
[0173] Based on the characteristics of each requirement node, the overall language style is determined to be formal with a moderately positive tone, while ensuring that the information is concise and easy to understand. This style meets the accuracy requirements of practical needs, the appeal of aesthetic needs, and the efficiency of information for multifunctional needs.
[0174] Step S1515: Adjust the vocabulary selection, sentence structure and tone of the text content according to the language style features, and output the complete text content, which contains multiple paragraphs, each paragraph corresponding to a core semantic theme.
[0175] To maintain a formal tone, avoid colloquial expressions in vocabulary selection. For example, instead of "very comfortable to wear," use "comfortable wearing experience"; instead of "has a good slimming effect," use "has a significant effect on body shape." Primarily use declarative sentences, reducing the use of exclamatory and interrogative sentences. For instance, change "This dress is so versatile!" to "This dress has strong versatility."
[0176] Considering emotional appeal, when describing the advantages of the design, use expressions with positive emotions such as "the elegant waist-cinching design cleverly outlines the waistline, and the A-line skirt falls naturally, showing a lively temperament"; when introducing the wearing scenarios, use statements that evoke positive associations in consumers, such as "whether it is a capable look for commuting to the workplace or a casual style for weekend leisure, it can be easily handled".
[0177] To maintain conciseness, each paragraph should be limited to three to five sentences, highlighting the core meaning. For example, the paragraph about the fabric: "This dress is made of breathable cotton-linen fabric, which is lightweight, skin-friendly, and suitable for extended wear. The fabric's loose fiber structure provides excellent moisture-wicking properties, keeping you dry and comfortable even in hot weather." The paragraph about the silhouette: "The dress features a fitted A-line design. The cinched waist flatters the waist and abdomen, while the A-line skirt flares out naturally from the waist, effectively concealing the hips and thighs, making it suitable for various body types."
[0178] After adjustments, the output text content is divided into four paragraphs, corresponding to four core semantic themes: fabric characteristics, washing and care, pattern advantages, and wearing scenarios. The language style of each paragraph is consistent with the defined features, and the overall content accurately conveys product information while meeting consumer needs and expectations.
[0179] Step S152: Analyze the visual display sequence in the optimized content framework, integrate each visual display unit according to the content composition rules, and generate visual content that corresponds to the text content. The presentation style of the visual content conforms to the content element node requirements in the requirement association network.
[0180] Analyze and optimize the visual display sequence within the content framework, clarifying the content and function of each visual display unit. For example, visual display unit V1 is a close-up image of the dress fabric details, V2 is an overall image of the dress worn (front view), V3 is an overall image of the dress worn (side view), and V4 is an illustration of how to wear the dress in different scenarios (workplace, casual).
[0181] Content composition rules include highlighting the main subject, complementary perspectives, and scene matching. When integrating the images, V1 is placed at the beginning of the visual content, using a close-up shot to emphasize the fabric's texture and feel, echoing the description of the fabric's characteristics in the text. V2 and V3 are arranged vertically, with the front view showcasing the overall style of the dress and the side view highlighting the fitted waist, complementing the description of the dress's silhouette in the text. V4 is arranged in the order of workplace scenes first, followed by casual scenes, corresponding to the order of descriptions of the outfit's style in the text.
[0182] The "fabric attributes" content element node in the demand association network requires the visual content to clearly display the fabric details. Therefore, the composition of V1 needs to focus on the fabric texture and avoid interference from irrelevant elements. The "pattern design" content element node requires highlighting the pattern lines. The compositions of V2 and V3 need to use simple backgrounds so that the silhouette and lines of the dress become the visual focus. The "scene adaptation" content element node requires the visual content to reflect the atmosphere of different scenes. In V4, the background of the workplace scene is set to an office environment, and the background of the leisure scene is set to a coffee shop or park, so that the visual presentation style is consistent with the requirements of each content element node.
[0183] The integrated visual content features natural transitions between units, maintaining overall harmony through a unified color scheme (such as soft, light colors) and a consistent shooting style (such as natural light photography). At the same time, each unit effectively echoes the corresponding part of the text content, enhancing the intuitiveness of information delivery.
[0184] Step S153: Based on the set of semantic association rules in the optimized content framework, the text content and visual content are associated and combined to form a multimodal content unit. Each multimodal content unit includes corresponding user demand node tags, content element node tags, and scene form node tags.
[0185] The set of semantic association rules in the optimized content framework includes rules such as "fabric characteristic text corresponds to fabric detail visuals," "pattern advantage text corresponds to pattern display visuals," and "outfit scenario text corresponds to scene outfit visuals." Based on these rules, the fabric paragraph of the text content is combined with V1 to form the first multimodal content unit; the pattern paragraph is combined with V2 and V3 to form the second multimodal content unit; and the outfit scenario paragraph is combined with V4 to form the third multimodal content unit.
[0186] Each multimodal content unit is labeled with corresponding tags. The first unit is labeled with user demand tags for "fabric comfort needs", content element tags for "fabric attributes", and scenario format tags for "product details page scenarios"; the second unit is labeled with user demand tags for "slimming fit needs", content element tags for "fit design", and scenario format tags for "product details page scenarios"; the third unit is labeled with user demand tags for "scenario adaptation needs", content element tags for "scenario adaptation", and scenario format tags for "short video promotion scenarios".
[0187] The addition of these tags enables multimodal content units to be precisely matched with nodes in the demand-related network, ensuring that different content units can be pushed to the appropriate scenarios and channels.
[0188] Step S154: Arrange and combine multiple multimodal content units according to the structural type of the scene form node to generate a target content set containing different modal combinations.
[0189] The structural types of scenario-based nodes include "product detail page structure" and "short video promotion structure". The "product detail page structure" requires content units to be arranged in the order of "basic attributes - core advantages - usage scenarios". Therefore, the first multimodal content unit (fabric), the second multimodal content unit (fit), and the third multimodal content unit (outfit scenarios) are arranged in sequence to form a content combination suitable for the product detail page.
[0190] The "short video promotion structure" requires an attention-grabbing visual opening, key information interspersed throughout, and action-oriented content at the end. Accordingly, V2 (a front view of the overall outfit) serves as the opening, followed by a simplified version of the second multimodal content unit (text highlighting the silhouette's advantages and V3), then the third multimodal content unit (outfit scenario), and finally, the guiding text "Click to view details" is added, forming a content combination suitable for short video promotion.
[0191] The target content set contains content in these two different modal combinations. Each combination maintains logical coherence between content units and conforms to the structural requirements of the corresponding scenario-based nodes, thus meeting the needs of different promotion scenarios.
[0192] Step S155: Retrieve content distribution node information that matches the multimodal content unit tags in the target content set. The content distribution node information includes node identifier, distribution range, and receiving preference.
[0193] Content distribution node information is retrieved through tag matching. Multimodal content units tagged "product details page scenario" are matched with distribution nodes identified as "PC details page node" and "mobile details page node," respectively. Their distribution scope is the website PC terminal and the mobile APP terminal, and the receiving preference is detailed content combining images and text.
[0194] Multimodal content units labeled "short video promotion scenario" are matched with distribution nodes identified as "short video platform node A" and "short video platform node B". The distribution scope is the corresponding short video platform user group, and the receiving preference is for short content with strong visual impact.
[0195] Meanwhile, the distribution node information also includes user profile data for each node. For example, users of the "mobile details page node" are mostly young women, while users of the "short video platform node A" are mainly young people from second- and third-tier cities.
[0196] Step S156: Based on the receiving preferences of the content distribution node, the multimodal content units in the target content set are filtered so that the multimodal content units pushed to the content distribution node meet the node's receiving requirements.
[0197] Both the "PC-based product detail page node" and the "mobile-based product detail page node" prefer detailed text and image content. Therefore, a combination of content suitable for the product detail page is selected from the target content set, while retaining complete text and visual units to ensure comprehensive information. For the "mobile-based product detail page node," considering the smaller screen size of mobile phones, the size of the visual content is adjusted to fit the mobile screen ratio, and the text font is set to an easily readable size.
[0198] Both "Short Video Platform Node A" and "Short Video Platform Node B" prefer short, visually appealing content. They filter content combinations suitable for short video promotion, simplify text content, retain core information, and keep video length within the platform's requirements, emphasizing the attractiveness of the visual content. For example, they delete text describing fabrics in detail, retaining only the key phrase "breathable cotton and linen, comfortable and versatile."
[0199] The filtered content units all conform to the receiving preferences of each distribution node and can be better presented in different distribution channels.
[0200] Step S157: Push the filtered target content set to the corresponding content distribution node. After receiving the target content set, the content distribution node performs content display and dissemination operations.
[0201] The content combination suitable for the product detail page is pushed to the "PC detail page node" and the "mobile detail page node" respectively. After receiving it, these two nodes display the content on the dress product page on the PC website and the mobile APP, including a complete text description and orderly arranged visual content, so that consumers can browse and understand the product details.
[0202] Content combinations suitable for short video promotion are pushed to "Short Video Platform Node A" and "Short Video Platform Node B". After receiving the content, the nodes process it into short videos that conform to the platform's format and recommend and spread it on the platform. The dynamic visuals and concise information attract user attention and guide users to click on the link to enter the product details page.
[0203] When performing display and dissemination operations, each content distribution node will also provide real-time feedback on data such as content exposure and clicks. This data can be used to optimize and adjust the content generation method in the future, thereby improving the effectiveness of content promotion.
[0204] Figure 2 The illustration shows exemplary hardware and software components of a multimodal large model-based content generation system 100 that can implement the ideas of this application, according to some embodiments of this application. For example, a processor 120 can be used in the multimodal large model-based content generation system 100 and to perform the functions in this application.
[0205] For example, a multimodal large-scale content generation system 100 may include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and various forms of storage media 140, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the multimodal large-scale content generation system 100 may also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The methods of this application can be implemented according to these program instructions. The multimodal large-scale content generation system 100 also includes an I / O interface 150 between the computer and other input / output devices.
[0206] Furthermore, embodiments of the present invention also provide a readable storage medium, wherein computer-executable instructions are preset in the readable storage medium, and when the processor executes the computer-executable instructions, the above-described content generation method based on a multimodal large model is implemented.
[0207] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
Claims
1. A content generation method based on a multi-modal large model, characterized in that, The method comprises: establishing a demand association network of content generation, the demand association network comprising a plurality of nodes, the plurality of nodes including user demand nodes, content element nodes and scene form nodes, each node being connected by an association edge and being marked with an association strength; the establishment of the demand association network of content generation comprises: collecting user historical interaction records and content detail documents, extracting user behavior keywords and feedback phrases from the user historical interaction records, and extracting content core attribute vocabularies and feature descriptions from the content detail documents; performing clustering processing on the user behavior keywords and feedback phrases to generate a plurality of user demand nodes, each user demand node comprising a demand type label and a demand description; performing hierarchical division on the content core attribute vocabularies and feature descriptions to construct a content element node system, the content element node system comprising main content category nodes, sub-content attribute nodes and content parameter nodes; analyzing content samples in a content library that match the user demand nodes and the content element nodes, and extracting presentation modes, structure types and modality combination modes of the content samples as constituent elements of scene form nodes; calculating the co-occurrence frequency between the user demand nodes and the content element nodes by an association rule mining algorithm, setting a first association strength parameter according to the co-occurrence frequency as the marked information of the association edge between them; calculating the adaptation degree between the content element nodes and the scene form nodes, setting a second association strength parameter according to the adaptation degree as the marked information of the association edge between them; calculating the satisfaction degree between the user demand nodes and the scene form nodes, setting a third association strength parameter according to the satisfaction degree as the marked information of the association edge between them; networking the user demand nodes, content element nodes, scene form nodes and association edges marked with the first association strength parameter, the second association strength parameter and the third association strength parameter to generate a demand association network with a three-layer node structure; constructing a multi-modal element coupling pool based on the demand association network, the multi-modal element coupling pool comprising a plurality of element units, the plurality of element units including text description units, visual display units and semantic association units, each element unit forming a mapping relationship with the nodes of the demand association network; the construction of the multi-modal element coupling pool based on the demand association network comprises: based on the user demand nodes in the demand association network, retrieving description vocabularies, sentence templates and chapter structure fragments in a text corpus that match the demand type labels, and then classifying according to the demand description to generate text description units, each text description unit comprising a classification identifier and corresponding text content; based on the content element nodes in the demand association network, filtering image segments, video clips and display models in a visual resource library that correspond to the content core attribute vocabularies and feature descriptions, and then performing feature extraction to generate visual display units comprising content feature labels, each visual display unit being associated with a corresponding content element node; The semantic correspondence rule, the presentation logic rule, and the matching rule are respectively converted into a semantic association unit after analyzing the association edges and the association strength parameters between the nodes in the demand association network, extracting the semantic correspondence rule between the user demand node and the content element node, the presentation logic rule between the content element node and the scene form node, and the matching rule between the user demand node and the scene form node, and the semantic association unit includes rule type description and strength threshold range; A mapping relationship table is established between the text description unit and the user demand node, the content element node, and the scene form node to determine the corresponding node and mapping weight of each text description unit; A mapping relationship table is established between the visual display unit and the user demand node, the content element node, and the scene form node to determine the corresponding node and mapping weight of each visual display unit; A mapping relationship table is established between the semantic association unit and the association edges in the demand association network to determine the corresponding association edge and association strength parameter range of each semantic association unit; The text description unit, the visual display unit, the semantic association unit, and their mapping relationship tables are integrated to construct a multi-modal element coupling pool; A multi-modal large model is called to perform situational mapping processing on the multi-modal element coupling pool to generate an initial content framework set matched with the demand association network; Element collaborative optimization processing is performed on the initial content framework set to obtain an optimized content framework that meets the constraints of the content element node; According to the optimized content framework, a target content set is generated, and the target content set is pushed to the corresponding content distribution node to complete the content propagation operation. 2.The content generation method based on a multi-modal large model according to claim 1, wherein, The calling of the multi-modal large model to perform situational mapping processing on the multi-modal element coupling pool to generate an initial content framework set matched with the demand association network includes: The demand association network is input into the situational analysis layer of the multi-modal large model to identify the core demand keywords of the user demand node, the key attribute features of the content element node, and the preference types of the scene form node through a node feature extraction algorithm to generate a situational feature vector; The text description unit, the visual display unit, and the semantic association unit in the multi-modal element coupling pool are input into the element coding layer of the multi-modal large model for text coding, visual coding, and semantic rule coding processing to generate an element feature vector set composed of a text feature vector, a visual feature vector, and a semantic association feature vector; The mapping matching layer of the multi-modal large model calculates the matching degree of each element feature vector and the situational feature vector based on the mapping relationship table of the multi-modal element coupling pool to filter out candidate element feature vectors that meet the preset conditions; According to the association strength parameters of the association edges in the demand association network, the candidate element feature vectors are weighted and sorted to determine the priority of the text description unit, the visual display unit, and the semantic association unit in content generation, and a priority sorting result is obtained. The framework construction layer using the multi-modal large model constructs a layer based on the priority sorting result, combines the text feature vector, visual feature vector and semantic association feature vector according to the semantic logic of content presentation, and generates a plurality of initial content frameworks, each initial content framework containing an element combination sequence and a framework structure description; Each initial content framework is subjected to integrity checking to check whether the core user demand node, the key content element node and the important scene form node in the demand association network are covered, and if there is a missing, the corresponding element feature vector is selected from the multi-modal element coupling pool for supplementing; The initial content frameworks passing the integrity checking are integrated into an initial content framework set, and each initial content framework corresponds to a matching degree score of the demand association network in the initial content framework set. 3.The content generation method based on a multi-modal large model according to claim 2, wherein, The context analysis layer of inputting the demand association network into the multi-modal large model identifies the core demand keywords of the user demand node, the key attribute features of the content element node and the preference type of the scene form node through a node feature extraction algorithm, and generates a context feature vector, including: Feature extraction is performed on the user demand nodes in the demand association network, and the demand type label, demand description and associated content element node information of each user demand node are extracted, and the core demand keywords are mined through a topic model algorithm, the core demand keywords are used to represent the demand points that the user pays most attention to; Feature extraction is performed on the content element nodes in the demand association network, and the class label of the main content category node, the attribute description of the sub-content attribute node and the specific parameter value of the content parameter node are extracted, and the key attribute features affecting content generation are selected through a feature selection algorithm, the key attribute features include the function, material and specification of the content; Feature extraction is performed on the scene form nodes in the demand association network, and the feature parameters of the presentation mode, structure type and modal combination mode of the scene form node are extracted, and the scene form preference type is determined through cluster analysis, the preference type includes the detail page form, short video form and fragment form; The core demand keywords, key attribute features and scene form preference type are standardized to generate computable feature values; According to the importance of the core demand keywords, key attribute features and scene form preference type in content generation, a weight parameter is set, the feature values are weighted to generate a target feature vector; The target feature vector is mapped to a high-dimensional semantic space to generate a context feature vector. 4.The content generation method based on a multi-modal large model according to claim 1, wherein, The initial content framework set is subjected to element collaborative optimization processing to obtain an optimized content framework conforming to the content element node constraint, including: Analyzing the content element node constraint condition in the demand association network, determining the presentation requirement of the content core attribute, the matching standard of the content specification parameter and the creation rule of the content style feature; extracting a text description unit, a visual presentation unit and a semantic association unit contained in each initial content framework from the initial content framework set, analyzing a matching degree of each element unit with the content element node constraint condition, the matching degree being calculated by a content fit degree of the element unit and the constraint condition; performing replacement processing on the element unit with insufficient matching degree, selecting an element unit more conforming to the content element node constraint condition from the multi-modal element coupling pool to replace, and maintaining semantic association between the text description unit and the visual presentation unit during the replacement process; adjusting a combination order and an association strength of the text description unit, the visual presentation unit and the semantic association unit, so that a content logic of the text description unit and a presentation logic of the visual presentation unit remain consistent, and the semantic association unit can accurately reflect a logical relationship between text and vision; performing element cooperativity evaluation on the adjusted content framework, calculating a coherence score between the text description units, a style consistency score between the visual presentation units and a semantic matching degree score between the text description units and the visual presentation units; if the element cooperativity evaluation result meets a preset threshold, taking the adjusted content framework as a candidate optimized framework; if not, repeatedly performing replacement and adjustment operations until the element cooperativity evaluation result meets the standard; selecting a framework with the highest matching degree with the content element node constraint condition from the candidate optimized framework as an optimized content framework, the optimized content framework containing a cooperatively optimized text description sequence, a visual presentation sequence and a semantic association rule set. 5.The content generation method based on a multi-modal large model according to claim 4, wherein, The element cooperativity evaluation on the adjusted content framework, the calculation of the coherence score between the text description units, the style consistency score between the visual presentation units and the semantic matching degree score between the text description units and the visual presentation units, includes: extracting a text description unit sequence in the adjusted content framework, analyzing lexical cohesion information, grammatical association information and semantic coherence information between adjacent text description units, calculating a coherence score through a pre-trained language model, and taking the coherence score as the coherence score between the text description units; extracting a visual presentation unit sequence in the adjusted content framework, analyzing consistency of color keynote, composition style and element layout of the visual presentation unit, calculating a style similarity through a visual feature comparison algorithm, and taking the style similarity as the style consistency score between the visual presentation units; establishing a corresponding relationship between the text description units and the visual presentation units, each text description unit corresponding to one or more visual presentation units, extracting core semantics of the text description units and visual semantics of the visual presentation units; calculating a semantic distance between the core semantics and the visual semantics, the semantic distance being calculated by an Euclidean distance in a semantic vector space; translating the semantic distance into a semantic matching degree score, the semantic matching degree score being in a negative correlation relationship with the semantic distance; The coherence score between the text description units, the style consistency score between the visual presentation units, and the semantic matching degree score between the text description units and the visual presentation units are weighted and calculated according to a preset weight parameter to obtain a comprehensive score of the element collaboration evaluation. 6.The content generation method based on a multi-modal large model according to claim 1, wherein, The generating a target content set according to the optimized content framework and pushing the target content set to a corresponding content distribution node to complete a content propagation operation, comprises: parsing a text description sequence in the optimized content framework, combining each text description unit into complete text content according to semantic logic, a language style of the text content matching a user demand node in the demand association network; parsing a visual presentation sequence in the optimized content framework, integrating each visual presentation unit according to a content composition rule to generate visual content corresponding to the text content, a presentation style of the visual content meeting a content element node requirement in the demand association network; according to a semantic association rule set in the optimized content framework, associating and combining the text content and the visual content to form a multi-modal content unit, each multi-modal content unit containing a corresponding user demand node label, a content element node label and a scene form node label; arranging and combining a plurality of the multi-modal content units according to a structure type of the scene form node to generate a target content set containing different modal combinations; retrieving content distribution node information matching the multi-modal content unit label in the target content set, the content distribution node information containing a node identifier, a distribution range and a receiving preference; according to the receiving preference of the content distribution node, screening the multi-modal content unit in the target content set, so that the multi-modal content unit pushed to the content distribution node meets the receiving requirement of the node; pushing the screened target content set to the corresponding content distribution node, the content distribution node receiving the target content set and performing content display and propagation operations.
7. The content generation method based on a multi-modal large model according to claim 6, characterized in that, The parsing a text description sequence in the optimized content framework, combining each text description unit into complete text content according to semantic logic, comprises: extracting each text description unit of the text description sequence in the optimized content framework, parsing the hierarchical identifier, text content and corresponding user demand node information of each text description unit; determining the combination order of the text description unit according to the semantic association relationship between the text description units; performing a connection process on the text description unit in the combination process, adding a preset conjunction, a transition sentence or a paragraph; analyzing the demand type label and demand description of the user demand node in the demand association network to determine the corresponding language style features, the language style features including formality, emotional tendency and conciseness; adjusting the vocabulary selection, sentence structure and expression tone of the text content according to the language style features to output complete text content, the text content containing a plurality of paragraphs, each paragraph corresponding to a core semantic theme.
8. A multi-modal large model based content generation system, characterized by, The application relates to a device for implementing the content generation method based on a multimodal large model, comprising a processor and a memory, wherein the memory is connected with the processor, the memory is used for storing programs, instructions or codes, and the processor is used for executing the programs, instructions or codes in the memory to realize the content generation method based on a multimodal large model according to any one of claims 1-7.
Citation Information
Patent Citations
AIGC content generation method and system based on multi-modal fusion
CN120578796A
Generating Context-Aware Rendering of Media Contents for Assistant Systems
US20220210111A1
Cited By
Software system development method based on multi-modal AI large model
CN121657970A