Content generation method, electronic device, computer storage medium, and program product

CN120407913BActive Publication Date: 2026-08-11DINGTALK (CHINA) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]虽然现在的AIGC工具可以向用户提供上述内容生成服务,但其仅具有通用内容生成能力,也即,如果不同的用户输入相同的需求,则AIGC工具会为这些不同的用户生成相同或相似的内容,而不会根据用户的不同而为用户生成符合其个性的个性化内容

Benefits of technology

[0009]根据本申请实施例提供的内容生成方案,获取用户的多个模态的个性化数据及待生成的个性化模板的信息,对多个模态的个性化数据分别进行特征提取,获取各个模态对应的特征数据,形成更为全面、丰富的用户个性描述;基于各个模态的特征数据、个性化模板的信息,生成模板描述信息,从而描述出符合用户个性化需求的个性化模板;进一步地,基于模板描述信息,从多个用于生成不同模板组件的机器学习模型中选取对应的机器学习模型,使用选取的机器学习模型生成对应的模板组件,并基于生成的模板组件生成个性化模板。由此,通过分析用户多个模态的个性化数据,获取用户多个维度的个性化信息,通过不同类型的机器学习模型、基于这些个性化信息生成的模板组件都符合用户的个性化需求,因而使得基于这些模板组件生成的个性化模板整体符合用户的个性,提高了为用户生成的内容的个性化程度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407913B_ABST
    Figure CN120407913B_ABST
Patent Text Reader

Abstract

This application provides a content generation method, an electronic device, a computer storage medium, and a program product. The content generation method includes: acquiring personalized data of multiple modalities of a user and information about a personalized template to be generated; extracting features from the personalized data of the multiple modalities to obtain feature data corresponding to each modality; generating template description information based on the feature data of each modality and the personalized template information; selecting a corresponding machine learning model from multiple machine learning models used to generate different template components based on the template description information; generating corresponding template components using the selected machine learning model; and generating a personalized template based on the generated template components. This application improves the personalization of content generated for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a content generation method, electronic device, computer storage medium, and computer program product. Background Technology

[0002] With the development of AI (Artificial Intelligence) technology, AIGC (Artificial Intelligence Generated Content) is increasingly being applied to people's work and lives. AIGC is a technology that uses AI to automatically generate various types of content, including text, images, audio, and video. It is usually based on deep learning and natural language processing and can create content of a certain quality without direct human intervention.

[0003] While current AIGC tools can provide users with the aforementioned content generation services, they only have general content generation capabilities. That is, if different users input the same requirements, the AIGC tool will generate the same or similar content for these different users, rather than generating personalized content that matches the individual user's personality. Summary of the Invention

[0004] In view of this, embodiments of this application provide a content generation scheme to at least partially solve the above-mentioned problems.

[0005] According to a first aspect of the embodiments of this application, a content generation method is provided, comprising: acquiring personalized data of multiple modalities of a user and information of a personalized template to be generated; performing feature extraction on the personalized data of the multiple modalities respectively to obtain feature data corresponding to each modality; generating template description information based on the feature data of each modality and the information of the personalized template; selecting a corresponding machine learning model from multiple machine learning models used to generate different template components based on the template description information; generating a corresponding template component using the selected machine learning model, and generating the personalized template based on the generated template component.

[0006] According to a second aspect of the present application, an electronic device is provided, including: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, which causes the processor to perform an operation corresponding to the method described in the first aspect.

[0007] According to a third aspect of the embodiments of this application, a computer storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.

[0008] According to a fourth aspect of the embodiments of this application, a computer program product is provided, including computer instructions that instruct a computing device to perform an operation corresponding to the method described in the first aspect.

[0009] According to the content generation scheme provided in this application, personalized data of multiple modalities of the user and information of the personalized template to be generated are obtained. Feature extraction is performed on the personalized data of each modality to obtain feature data corresponding to each modality, forming a more comprehensive and richer description of the user's personality. Based on the feature data of each modality and the information of the personalized template, template description information is generated, thereby describing a personalized template that meets the user's personalized needs. Further, based on the template description information, a corresponding machine learning model is selected from multiple machine learning models used to generate different template components. The selected machine learning model is used to generate the corresponding template component, and a personalized template is generated based on the generated template component. Therefore, by analyzing the personalized data of multiple modalities of the user, multi-dimensional personalized information of the user is obtained. Different types of machine learning models and template components generated based on this personalized information all meet the user's personalized needs, thus making the personalized template generated based on these template components generally conform to the user's personality, improving the personalization of the content generated for the user.

[0010] Therefore, it can be seen that the solution of this application embodiment greatly improves the personalization of the content generated by the AIGC tool for users and enhances the user experience of the AIGC tool. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0012] Figure 1 A schematic diagram of an exemplary system for the content generation method applicable to the embodiments of this application;

[0013] Figure 2A This is a flowchart illustrating the steps of a content generation method according to an embodiment of this application;

[0014] Figure 2B for Figure 2A A schematic diagram of a scenario example in the illustrated embodiment; Figure 3This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0015] To enable those skilled in the art to better understand the technical solutions in the embodiments of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art should fall within the protection scope of the embodiments of this application.

[0016] The specific implementation of the embodiments of this application will be further described below with reference to the accompanying drawings.

[0017] Figure 1 An exemplary system for a content generation method applicable to embodiments of this application is shown. For example... Figure 1 As shown, the system 100 may include a cloud server 102, a communication network 104, and / or one or more user devices 106. Figure 1 The example in the text shows multiple user devices.

[0018] The cloud server 102 can be any suitable device for storing information, data, programs, and / or any other suitable type of content, including but not limited to distributed storage system devices, server clusters, computing cloud server clusters, etc. In some embodiments, the cloud server 102 can perform any suitable function. For example, in some embodiments, the cloud server 102 is equipped with multiple machine learning models for generating personalized content, including: acquiring personalized data of multiple modalities of the user and information of the personalized template to be generated; performing feature extraction on the personalized data of multiple modalities respectively to obtain feature data corresponding to each modality; generating template description information based on the feature data of each modality and the information of the personalized template; selecting a corresponding machine learning model from multiple machine learning models used to generate different template components based on the template description information; generating the corresponding template component using the selected machine learning model, and generating a personalized template based on the generated template component.

[0019] In some embodiments, the communication network 104 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 104 can include any one or more of the following: the Internet, an intranet, a wide area network (WAN), a local area network (LAN), a wireless network, a digital subscriber line (DSL) network, a frame relay network, an asynchronous transfer mode (ATM) network, a virtual private network (VPN), and / or any other suitable communication network. The user equipment 106 can be connected to the communication network 104 via one or more communication links (e.g., communication link 112), and the communication network 104 can be linked to the cloud server 102 via one or more communication links (e.g., communication link 114). The communication link can be any communication link suitable for transmitting data between the user equipment 106 and the cloud server 102, such as a network link, a dial-up link, a wireless link, a hardwired link, any other suitable communication link, or any suitable combination of such links.

[0020] User device 106 may include any one or more user devices suitable for interacting with a user. In some embodiments, user device 106 may perform information input, such as inputting information for a personalized template, and may display a personalized template. In some embodiments, user device 106 may include any suitable type of device. For example, in some embodiments, user device 106 may include a mobile device, tablet computer, laptop computer, desktop computer, and / or any other suitable type of user device.

[0021] Based on the above system, this application provides a content generation scheme, which will be described below through several embodiments.

[0022] Reference Figure 2A This document illustrates a flowchart of a content generation method according to an embodiment of this application. The content generation method of this embodiment includes the following steps:

[0023] Step S202: Obtain personalized data of multiple user modalities and information on the personalized template to be generated.

[0024] In this embodiment, to ensure that the subsequently generated template better meets the user's personalized needs, the system first obtains the user's personalized data across multiple modalities, as well as information about the personalized template the user wants to generate. This personalized data across multiple modalities can be personalized data in various different formats, such as text, video, image, and audio data. Through this diverse personalized data, the user's information and characteristics can be effectively represented from multiple dimensions, thus providing a basis for generating a personalized template that meets the user's individual needs.

[0025] In one alternative approach, the user's multimodal personalized data includes, but is not limited to, at least one of the following: the user's identifier image, text data of the user's industry sector, the user's online information data, the user's template style preference data, the user's personalized audio data, and the user's personalized video data.

[0026] The user's identifier image can be any image that clearly identifies and represents the user, used for user identification and differentiation in different scenarios. For example, a user's identifier image can be their profile picture or a poster on a social media platform. For businesses, organizations, or brands, the identifier image can be their logo, using design elements to convey identifiable information such as product, service content, and industry attributes. The user's identifier image can intuitively showcase their personality and style, improving user recognizability. Using the user's identifier image as one of their multiple modalities of personalized data also makes the generated templates more recognizable when needed in subsequent template generation.

[0027] Textual data related to a user's industry sector can include, but is not limited to, any data in text form that reflects the user's industry (including data published by the user themselves and / or data related to the user published by other parties). It is typically related to the user's expertise and technical field, and is therefore more specialized. Examples include industry promotional content, industry reports, and content related to industry characteristics. This textual data reflects the user's professional background, potential work content, and research interests, helping AIGC tools understand the user's specific needs within their industry sector and thus provide more accurate and personalized content.

[0028] User online information data can represent data related to a user's online behavior, such as advertising data, review reports, and trending behavior data. This data may also relate to the user's industry, but its level of specialization is less than that of textual data related to the user's industry. Through user online information data, we can reflect a user's online behavior, thus providing a more intuitive understanding of their individual needs and the degree to which they are recognized and followed.

[0029] User template style preference data can characterize user preferences for templates, including but not limited to user preferences for template style such as color scheme, layout, and background material. User template style preference data can reflect users' personalized needs when using templates or template generation tools, and different users have different personalized characteristics. It should be noted that in this application embodiment, "template" can be any template applicable to the solution of this application embodiment, such as web page templates, report templates, PPT templates, report templates, etc., and this application embodiment does not limit the specific implementation form of "template".

[0030] Personalized audio data can be audio data that identifies and represents a user, such as a segment of speech that effectively represents and identifies the user, promotional audio from a company, or news report audio. However, it is not limited to these; it can also include user preference data regarding audio content, such as user preferences for audio content, including but not limited to the style, type, and creator of the audio content; or user listening behavior preferences, such as the number of times they listen to one or more audio tracks and the duration of their listening.

[0031] Personalized video data can be video data that identifies the user, such as a video that effectively represents and identifies the user, a corporate promotional video, or a product demonstration video. However, it is not limited to these. It can also include user preference data regarding video content, such as user preferences for video content, including but not limited to the style, creator, and duration of the video; or user viewing behavior preference data, including but not limited to the number of times and duration of viewing a particular video or groups of videos.

[0032] The personalized data of the above users can all be used as a reference and basis when generating personalized templates in the future.

[0033] The information for the personalized template to be generated effectively describes what kind of template the user wants to generate. It includes at least the content information of the personalized template, that is, the content that the personalized template needs to present. For example, the content information of a personalized PPT template to be generated may include information indicating that images and text for the user's product introduction need to be generated. Because the content information needs to be carried by components, determining the content information of the personalized template also means determining its required components and the content carried by those components. Taking the PPT template mentioned above as an example, if the user wants to generate a PPT template containing product introduction images and text, it means that the template needs to generate image components and text components. The image component is used to carry the product introduction images, and the text component is used to carry the product introduction text, and so on. Therefore, it can also be considered that in subsequent processes, corresponding content can be generated using methods such as machine learning models, such as using machine learning models to generate images, using machine learning models to generate text, and so on.

[0034] While personalized templates can be generated based on content information, to better match user preferences and styles, one option is to include layout information within the personalized template. Generally, templates are formed by template components; that is, a personalized template can include at least one template component. Different template components carry different types of content, such as text components for text display, image components for image display, and interactive components for receiving interactive operations. Therefore, layout information can describe the layout structure of the personalized template to be generated, the layout relationships between various template components, etc. For example, the position, size, and nesting relationship of multiple template components within the personalized template. Although a personalized template can be generated using default layout information based solely on the above content information, having the user input the layout information ensures that the generated personalized template better meets the user's individual needs.

[0035] Furthermore, to obtain information about the personalized template to be generated, in one alternative approach, the AIGC tool can provide a human-computer interaction interface for inputting information about the personalized template. This interface can be used to obtain the user's requirements for the personalized template to be generated. Within this interface, those skilled in the art can set any settings that can be input and / or selected by the user according to their needs. By operating these settings, the user can generate the information of the personalized template they desire. These settings include, but are not limited to: template name settings, template content settings, settings of components included in the template, template layout settings, template style settings, and so on.

[0036] It should be noted that when the template components include interactive components, in one optional way, the human-computer interaction interface can be configured to set the interactive items, so as to realize the user's interactive needs for certain components in the personalized template to be generated.

[0037] For example, in one optional approach, the interactive item includes a first interactive item. The first interactive item can be implemented as an interactive item that receives at least one of the following input: text content, voice content, video content, or image content from the user. Through the first interactive item, the generated personalized template can have human-computer interaction functionality. As another example, the first interactive item can also be implemented as an interactive item with corresponding functions set for template components in the personalized template to achieve interaction between components. For example, a button that implements a jump function; when the user clicks the button, they can jump from the current component to another component, thereby achieving interaction between components. In specific settings, the relevant setting information of the first interactive item can be displayed on the human-computer interaction interface, including but not limited to: the identifier, name, type, location, interactive input (acceptable input methods, such as the aforementioned text input or button clicks), and interactive response (the response after receiving the interactive input, such as displaying the input at the corresponding position in the template, or jumping to other components, etc.). In practical applications, those skilled in the art can also make other settings according to actual needs, all of which are within the protection scope of the embodiments of this application.

[0038] In addition, the interactive items in the human-computer interaction interface may optionally include at least one of the following: a second interactive item for the user to select the type of personalized template to be generated; a third interactive item for the user to select the components contained in the personalized template to be generated; a fourth interactive item for the user to set the layout of the personalized template to be generated; and a fifth interactive item for setting the content presented by the components contained in the personalized template to be generated.

[0039] By manipulating the first, second, third, fourth, and fifth interactive items, personalized template information can be generated.

[0040] Step S204: Perform feature extraction on the personalized data of multiple modalities to obtain the feature data corresponding to each modality.

[0041] To extract key information from personalized data across multiple modalities and reduce the complexity of the original data, one feasible approach is to perform feature extraction on the personalized data for each modality separately, obtaining feature data corresponding to each modality, thereby more accurately reflecting the user's personalized needs and preferences. Consequently, the feature data corresponding to each modality can be applied to personalized templates, making the personalized templates generated based on the feature data corresponding to each modality more in line with the user's personality. For example, feature extraction on personalized data across multiple modalities can be implemented based on machine learning models. However, this is not a limitation; feature extraction on personalized data across multiple modalities can also be implemented based on other appropriate methods such as statistical methods and data dimensionality reduction techniques. The embodiments of this application do not limit the specific implementation method of feature extraction.

[0042] It should be noted that image feature extraction can be achieved from multiple dimensions. Since the user's identification image has an identification function, in one feasible approach, if the personalized data includes the user's identification image, feature extraction is performed on the personalized data of multiple modalities, including: for the user's identification image, feature extraction of at least one of the image semantic features, color features, style features, and text features of the identification image.

[0043] Image semantic features of a labeled image refer to the image information features contained in the labeled image that can be understood by humans, enabling computers to understand and interpret the content of the image. In one example, when extracting image semantic features of a labeled image, a machine learning model can be used to extract semantic features from the labeled image. For example, deep learning models, such as convolutional neural network models and semantic segmentation models, can be used. Optionally, the extracted results can be further processed using methods such as principal component analysis and linear discriminant analysis to improve the effectiveness of the extracted image semantic features of the labeled image.

[0044] The color scheme features of a signage image describe its color matching rules. Based on these features, the visual expression of the signage image can be perceived, leading to a better description of the image. For example, the color scheme features of a signage image might be "black and white" or "blue and white." In one instance, extracting the color scheme features of a signage image can be achieved by using a machine learning model. However, this is not the only method; histograms of each color component in the signage image can also be calculated using image processing software such as Photoshop to reflect the composition and distribution of each color.

[0045] The style features of a signage image refer to the attributes that reflect its atmosphere and characteristics, making it more visually recognizable. For example, the style features of a signage image could be "oil painting style," "retro style," or "futuristic style." In one instance, extracting style features from a signage image can be done using machine learning models, such as deep learning models like convolutional neural networks. However, this is not limited to this; style features can also be extracted using statistical methods such as histogram of oriented gradients (HOR) and scale-invariant feature transform.

[0046] Text features of a sign image refer to the characteristics of the text content contained within the sign image, such as the text content, font size, and layout. In one instance, extracting text features from a sign image can be achieved using optical character recognition technology, machine learning models, such as deep learning models like convolutional neural networks.

[0047] By extracting these features, a basis can be provided for generating a label image that can effectively identify users and has obvious deformation compared to the original label image.

[0048] In addition, optionally, before feature extraction from the user's identification image, image preprocessing can be performed on the identification image, such as enhancement processing or noise reduction processing, to improve the image quality of the identification image and thus improve the reliability of the extracted features.

[0049] If a user's personalized data includes text data related to their industry sector, one feasible approach to feature extraction for this sector is to use machine learning models, such as deep learning models like convolutional neural networks, to learn and extract feature representations from the text data. However, this is not the only approach; other methods for extracting text features are also applicable. For instance, the TF-IDF (termfrequency–inverse document frequency) method can be used to measure the importance of each word in the text data and use it as one of the text data's features.

[0050] Optionally, before extracting features from text data related to the user's industry sector, the text data can be preprocessed. For example, preprocessing the text data may include word segmentation or removal of stop words, thereby reducing the difficulty of subsequent feature extraction.

[0051] If personalized data includes users' online information data, one feasible approach to feature extraction from this data is to use machine learning models. The extracted features can then be used as the feature data of the user's online information. For example, deep learning models such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) can be employed for feature extraction. Another feasible approach is to use topic modeling to extract the topic structure from the user's online information data and use the extracted results as its feature data.

[0052] If the personalized data includes users' template style preference data, in one feasible approach, feature extraction based on users' template style preference data can be implemented using machine learning models, probabilistic statistical models, or frequency distribution models that statistically analyze the number of times each template style is used.

[0053] If personalized data includes a user's personalized audio data, one feasible approach is to extract features from the audio data using a machine learning model. Another feasible approach is to use speech recognition technology to convert the specific content of the personalized audio data into text format, and then extract key information from the text format as feature data for the user's personalized audio data. However, this is not the only feasible approach. In another feasible approach, feature extraction from a user's personalized audio data can be achieved by using audio processing software such as Adobe Audition or Audacity to analyze speech rate, tone, volume, and energy levels in the personalized audio data, and then extracting speaker voice style features and emotional features based on the analysis results, which can then be used as feature data for the user's personalized audio data.

[0054] If the personalized data includes a user's personalized video data, one feasible approach to feature extraction for a user's personalized video is to first split the personalized video into multiple video frames before performing feature extraction. In one example, a machine learning model for feature extraction on video frames can be used to learn from the video frames of the user's personalized video data, thereby extracting feature data from the video frames. However, this is not the only approach. In another instance, based on multiple video frames, methods such as drawing color histograms and calculating color moments can be used to extract color features contained in the video. Methods such as edge detection and contour tracking can also be used to capture the shape and contour information of objects in the video frames. Furthermore, motion information of objects in the personalized video can be detected by comparing pixel differences between adjacent video frames, serving as feature data for the user's personalized video data.

[0055] Therefore, by extracting feature data from the user's personalized data across multiple modalities, we can obtain the user's personalized characteristics for the template from multiple dimensions, providing a basis for generating personalized templates that better meet the user's needs.

[0056] Step S206: Generate template description information based on the feature data of each modality and the information of the personalized template.

[0057] The template description information can be implemented as a data object or a custom data structure, which includes at least the feature data of each of the above modalities and the information of the personalized template, so as to facilitate subsequent processing and use.

[0058] In one example, the template description information can be implemented in any appropriate form, such as an XML (Extensible Markup Language) data structure or a JSON object. This not only effectively carries feature data and personalized template information, but also facilitates the adjustment of the template description information to express and implement information content more flexibly, and is more interpretable.

[0059] Furthermore, in this embodiment, the personalized template is implemented through a template component, which can be generated based on a machine learning model. That is, multiple machine learning models can be set to generate different template components. In this case, in one feasible approach, the template description information can also carry the information of the machine learning model to be used. Specifically, generating the template description information based on the feature data of each modality and the information of the personalized template can be achieved by generating the template description information based on the feature data of each modality, the information of the personalized template, and the obtained information of the machine learning model to be used. The information of the machine learning model to be used can, exemplarily, be input or set by the user through the aforementioned human-computer interaction interface, or it can be implemented through an automated selection method.

[0060] Based on this, in one feasible approach, the information of the machine learning model to be used is obtained by: acquiring the user's historical machine learning model usage data, and obtaining the information of the machine learning model to be used based on the historical machine learning model usage data; or, receiving the information of the machine learning model to be used input by the user through a human-computer interaction interface.

[0061] When users generate personalized templates in the past, they leave behind corresponding model usage data. By analyzing this data, we can obtain the user's model usage preferences (including preferences when generating various model components). Based on this, when the user generates a personalized template again, the preferred model can be used. However, this is not limited to this; the historical machine learning model usage data can also be data on the models used by a large number of users when generating different personalized templates. By analyzing this data, when a user (whether a new or returning user) generates a personalized template, information on the machine learning model to be used can be obtained based on the analysis of data from a large number of users.

[0062] Therefore, by using the information of the machine learning model to be used, the relevant machine learning model can be called quickly and efficiently to generate components of a personalized template that match the user's personality, thereby generating a personalized template.

[0063] Step S208: Based on the template description information, select the corresponding machine learning model from multiple machine learning models used to generate different template components.

[0064] In order to improve the efficiency and accuracy of generating template components that meet the individual needs of users, and to enhance the degree of personalization, the solution of this application embodiment is to pre-set multiple machine learning models for generating different template components.

[0065] In one feasible approach, multiple machine learning models for generating different template components include multiple classes of machine learning models for generating different types of components in the template, each class of machine learning models including at least one machine learning model. For example, four classes of machine learning models are included, respectively for generating text components, image components, video components, and audio components; each of these four classes of machine learning models can further include multiple different machine learning models. Although these multiple different machine learning models can all generate the same type of template component, each model has different characteristics; for example, model A is suitable for generating content images, model B is suitable for generating logo images, etc.

[0066] Based on this, in one feasible approach, a corresponding machine learning model is selected from multiple machine learning models used to generate different template components based on template description information, including: determining the information of the components in the personalized template to be generated based on the template description information; determining the type of machine learning model to be used based on the component information; and selecting the target machine learning model to be used from each determined type of machine learning model.

[0067] By generating different types of components based on corresponding types of machine learning models, each type of machine learning model generates template components based on the features of a specific type of template component. This approach is more targeted and can process relevant data and generate the required template components more quickly and accurately.

[0068] Alternatively, after confirming the corresponding type of machine learning model to be used based on the template component to be generated, the user's usage preferences for machine learning models in the past can be obtained based on the user's personalized data. In this way, a machine learning model that better matches the user's preferences can be selected from the pre-set machine learning models, so as to generate a personalized template component that better matches the user's personality in the future.

[0069] For example, when the information of the personalized template includes the content information of the personalized template, the types of multiple machine learning models include, but are not limited to, at least one of the following: a first type of machine learning model for generating content images (excluding identifier images) that match the content information and the user's personalized data; a second type of machine learning model for generating text that matches the content information and the user's personalized data; a third type of machine learning model for generating identifier images that match the content information and the user's personalized data; a fourth type of machine learning model for generating videos that match the content information and the user's personalized data; and a fifth type of machine learning model for generating interactive items that match the content information and the user's personalized data.

[0070] In one example, an image can be generated using a first-type machine learning model based on the description of the content image to be generated in the template description information. This first-type machine learning model can be any text-to-image model, including but not limited to LLM (Large Language Model) and Stable Diffusion. However, it is not limited to these; the description of the content image to be generated in the template description information can also be image-related information, including but not limited to image links, image identifiers, or any information about the image itself. In this case, when generating the content image using the first-type machine learning model, the corresponding image can first be obtained as a reference image based on the template description information, and then the content image can be generated based on the reference image. For example, the first-type machine learning model can be implemented as a GAN (Generative Adversarial Network) model or a Stable Diffusion model, used to combine the user's personalized data to generate personalized content images for the user. Thus, the function of generating image type components in a personalized template using a first-type machine learning model is realized.

[0071] In another example, text can be generated using a second-type machine learning model based on the description of the text content to be generated in the template description information. For example, the second-type machine learning model can be implemented as an LSTM (Long Short-Term Memory) model, a model based on a Transformer structure, or even an LLM model, etc., to combine the user's personalized data to generate personalized text for the user. Thus, the second-type machine learning model can be used to generate text-type components in personalized templates.

[0072] In another example, a third type of machine learning model can generate personalized icon images for a user based on the description of the icon image to be generated in the template description information and the user's original icon image. This could be achieved through a GAN model or a Diffusion model. Therefore, the third type of machine learning model can be used to generate icon images in personalized templates.

[0073] In another example, a fourth-type machine learning model can be used to generate a video that matches the user's personality, based on the description of the video to be generated in the template description information; alternatively, a fourth-type machine learning model can be used to generate a video that matches the user's personality, based on the description of the video to be generated in the template description information and reference video frames. For example, the fourth-type machine learning model can be implemented as an LVLM model, etc., to combine the user's personalized data to generate a personalized video for the user. Therefore, the fourth-type machine learning model can be used to generate video type components in personalized templates.

[0074] In another example, if the template components of the personalized template to be generated include interactive items, then the aforementioned determination of the type of machine learning model to be used based on the component information, and the selection of the target machine learning model from each determined type of machine learning model, can be implemented as follows: if the components of the personalized template to be generated include interactive items for human-computer interaction, the type of machine learning model to be used is determined to be the fifth type of machine learning model, and the target machine learning model to be used is selected from the fifth type of machine learning model. In this case, the step of generating the corresponding template component using the selected machine learning model can be implemented as follows: obtaining the interaction information corresponding to the interactive item from the template description information; generating the interactive item based on the interaction information, and generating corresponding interaction response information for the interactive item; or, determining the template component for responding to the interaction for the interactive item, and associating the interactive item with the determined template component. The personalized template information can include descriptions of interactive items, such as the type of interactive item (e.g., input text box, interactive button, etc.), the interaction chain (e.g., how to process input text, or how to process an interactive button after it is clicked), etc. Based on this description, a fifth-type machine learning model is used to generate interactive items that match the user's personality. These interactive items can be associated with corresponding code to implement their interactive functions. For example, the fifth-type machine learning model can be implemented as an LLM (Local Level Model).

[0075] With these machine learning models set up, when generating template components, a suitable machine learning model can be selected from them. For example, the corresponding machine learning model can be selected based on the information of the machine learning model to be used carried in the template description; or, a machine learning model that better suits the user's usage habits can be selected based on the user's historical usage data of machine learning models.

[0076] It should be noted that each type of machine learning model includes at least one machine learning model to adapt to different types of segmentation generation needs, while ensuring the robustness of the system.

[0077] Step S210: Use the selected machine learning model to generate the corresponding template component, and generate a personalized template based on the generated template component.

[0078] Once a specific machine learning model is determined, it can be used to generate corresponding template components. For specific generation methods, please refer to the aforementioned descriptions for each type of machine learning model.

[0079] After generating template components using various types of machine learning models, personalized templates can be further generated based on these components. In one feasible approach, if the information of the personalized template includes layout information, then the personalized template can be generated based on that layout information.

[0080] In one feasible approach, personalized templates can also be generated using a machine learning model. That is, a machine learning model is used to generate personalized templates based on the generated template components and layout information. This machine learning model can take any suitable form, including but not limited to machine learning models specifically designed for template generation, or LLM models, etc.

[0081] It should be noted that among the multiple machine learning models in the embodiments of this application, an LLM (Limited Learning Model) can be set. Due to the powerful generation capabilities of LLM, it can be used as a machine learning model for generating template components, or as a machine learning model for generating personalized templates. However, in addition to this, as mentioned above, the embodiments of this application also simultaneously set multiple dedicated models for generating template components. Thus, the combined use of large (LLM) and small (models used only for generating a certain type of component) models is achieved, ensuring both the generation quality of personalized models and improving the robustness of model usage.

[0082] As can be seen, through the embodiments of this application, personalized data of multiple user modalities and information of personalized templates to be generated can be obtained. Feature extraction is performed on the personalized data of multiple modalities to obtain feature data corresponding to each modality, forming a more comprehensive and richer description of user personality. Based on the feature data of each modality and the information of the personalized template, template description information is generated, thereby describing a personalized template that meets the user's personalized needs. Further, based on the template description information, a corresponding machine learning model is selected from multiple machine learning models used to generate different template components. The selected machine learning model is used to generate the corresponding template component, and a personalized template is generated based on the generated template component. Thus, by analyzing the personalized data of multiple user modalities, personalized information of the user in multiple dimensions is obtained. Different types of machine learning models and template components generated based on this personalized information all meet the user's personalized needs, thus making the personalized template generated based on these template components conform to the user's personality as a whole, improving the personalization of the content generated for the user.

[0083] The following example uses a specific scenario. Figure 2B The above-described content generation method of the embodiments of this application will be described by way of example.

[0084] In this example, suppose an art tutoring institution needs to generate a webpage template for corporate promotion. First, personalized data from multiple modalities of the art tutoring institution can be obtained, including, in this example, the institution's logo, performance reports, and news articles. Then, information about the personalized template the institution wants to generate can be obtained. For example, a user can input this information through a human-computer interaction interface, instructing the generation of a personalized template including the institution's logo, promotional text, and promotional images.

[0085] Furthermore, based on the personalized data of the art tutoring institution across multiple modalities, feature data corresponding to each modality's personalized data is extracted. For example, machine learning models can be used to extract features from the personalized data of multiple modalities to obtain the corresponding feature data. For instance, based on the art tutoring institution's logo, semantic features of the image, color scheme features combining red, yellow, and blue, and style features characteristic of oil painting style can be extracted. Figure 2B The diagram illustrates "identifying image features." For example, based on a performance report, extract the feature data of the top N performers in a given year. Figure 2B The diagram illustrates "industry characteristics." For example, based on news reports about this art tutoring institution, we can extract characteristic data from reports on the achievements of its outstanding teachers and students. Figure 2B The diagram illustrates "network information characteristics". Furthermore, based on the aforementioned characteristic data and personalized template information, template description information can be generated.

[0086] Assuming the machine learning model includes the aforementioned five types, in this example, to generate a personalized template including an organization logo, promotional text, and promotional image, we can select one model from the third type of machine learning model ("Model 3-1"), the first model from the three models from the second type of machine learning model ("Model 2-1"), and the second model from the two models from the first type of machine learning model ("Model 1-2"). The selected models will then be used to generate the components of the personalized template indicated by the personalized template information, namely the organization logo, promotional text, and promotional image. In this example, after generating the template components, the personalized template will be generated according to the default layout information. A simple example of a generated personalized template is shown in the figure, including a first image component carrying the organization logo, a first text component carrying the organization's promotional text, and a second image component carrying the organization's promotional image.

[0087] As can be seen, this example provides users with a content generation solution that better meets their personalized needs. It generates templates that match users' personalized needs based on their personalized data, thereby improving the personalization of the content generated for users.

[0088] Reference Figure 3 This document illustrates a schematic diagram of an electronic device according to an embodiment of this application. The specific embodiments of this application do not limit the specific implementation of the electronic device.

[0089] like Figure 3 As shown, the electronic device may include: a processor 302, a communications interface 304, a memory 306, and a communications bus 308.

[0090] in:

[0091] The processor 302, communication interface 304, and memory 306 communicate with each other via communication bus 308.

[0092] Communication interface 304 is used to communicate with other electronic devices or servers.

[0093] The processor 302 is used to execute program 310, specifically to perform the relevant steps in the above-described content generation method embodiment.

[0094] Specifically, program 310 may include program code that includes computer operation instructions.

[0095] Processor 302 may be a CPU, a GPU (Graphics Processing Unit), an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The electronic device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.

[0096] Memory 306 is used to store program 310. Memory 306 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0097] Program 310 may include multiple computer instructions. Specifically, program 310 may use multiple computer instructions to cause processor 302 to perform the operation corresponding to the content generation method described in any of the foregoing multiple method embodiments.

[0098] The specific implementation of each step in procedure 310 can be found in the corresponding descriptions of the steps and units in the above method embodiments, and has corresponding beneficial effects, which will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.

[0099] This application also provides a computer storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in any of the foregoing method embodiments. The computer storage medium includes, but is not limited to, compact disc read-only memory (CD-ROM), random access memory (RAM), floppy disk, hard disk, or magneto-optical disk.

[0100] This application also provides a computer program product, including computer instructions that instruct a computing device to perform an operation corresponding to any of the content generation methods in the above-described multiple method embodiments.

[0101] Furthermore, it should be noted that the user-related information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data used for training the model, data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0102] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of this application can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this application.

[0103] The methods described in the embodiments of this application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code downloaded over a network that is originally stored in a remote recording medium or a non-transitory machine-readable medium and will be stored in a local recording medium. Thus, the methods described herein can be stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA)). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., Random Access Memory (RAM), Read-Only Memory (ROM), Flash Memory, etc.) capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0104] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.

[0105] The above embodiments are only used to illustrate the embodiments of this application, and are not intended to limit the embodiments of this application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of this application, and the patent protection scope of the embodiments of this application should be defined by the claims.

Claims

1. A content generation method, comprising: Obtain personalized data from multiple user modalities and information about the personalized template to be generated; Feature extraction is performed on the personalized data of the multiple modalities to obtain the feature data corresponding to each modality; Based on the feature data of each modality and the information of the personalized template, template description information is generated; Based on the template description information, a corresponding machine learning model is selected from multiple machine learning models used to generate different template components; The selected machine learning model is used to generate corresponding template components, and the personalized template is generated based on the generated template components.

2. The method according to claim 1, wherein, The plurality of machine learning models for generating different template components include multiple types of machine learning models for generating different types of components in the template, and each type of machine learning model includes at least one machine learning model. The step of selecting a corresponding machine learning model from multiple machine learning models used to generate different template components based on the template description information includes: Based on the template description information, determine the information of the components in the personalized template to be generated; Based on the information of the components, determine the type of machine learning model to be used; From each identified type of machine learning model, select the target machine learning model to be used.

3. The method according to claim 2, wherein, The information of the personalized template includes at least the content information of the personalized template; The types of the multiple machine learning models include at least one of the following: A first type of machine learning model for generating content images that match the content information and the user's personalized data; A second type of machine learning model used to generate text that matches the content information and the user's personalized data; A third type of machine learning model for generating identifiable images that match the content information and the user's personalized data, wherein the third type of machine learning model is trained based on identifiable image samples of the user; A fourth type of machine learning model used to generate videos that match the content information and the user's personalized data; A fifth type of machine learning model used to generate interactive items that match the content information and the user's personalized data.

4. The method according to claim 3, wherein, The step of determining the type of machine learning model to be used based on the information of the component; and selecting the target machine learning model to be used from each determined type of machine learning model, includes: if it is determined based on the information of the component that the personalized template to be generated includes interactive items for human-computer interaction, then the type of the machine learning model to be used is determined to be the fifth type of machine learning model, and the target machine learning model to be used is selected from the fifth type of machine learning model. The step of generating a corresponding template component using a selected machine learning model includes: obtaining interaction information corresponding to the interaction item from the template description information; generating an interaction item based on the interaction information and generating corresponding interaction response information for the interaction item; or, determining a template component for responding to the interaction item and associating the interaction item with the determined template component.

5. The method according to claim 3, wherein, The personalized template information also includes the layout information of the personalized template; The process of generating the personalized template based on the generated template component includes: The personalized template is generated using a machine learning model that generates the template, based on the generated template components and the layout information.

6. The method according to claim 1, wherein, The step of generating template description information based on the feature data of each modality and the information of the personalized template includes: generating template description information based on the feature data of each modality, the information of the personalized template, and the information of the obtained machine learning model to be used; The step of selecting a corresponding machine learning model from multiple machine learning models used to generate different template components based on the template description information includes: selecting a corresponding machine learning model from multiple machine learning models used to generate different template components based on the information of the machine learning model to be used in the template description information.

7. The method according to claim 6, wherein, The information of the machine learning model to be used is obtained in the following ways: Obtain the user's historical machine learning model usage data, and based on the historical machine learning model usage data, obtain the information of the machine learning model to be used; or, Receive information about the machine learning model to be used, input by the user through the human-computer interaction interface.

8. The method according to claim 1, wherein, The user's multimodal personalized data includes at least one of the following: User's identifier image, text data of the user's industry sector, user's network information data, user's template style preference data, user's personalized audio data, and user's personalized video data.

9. The method according to claim 8, wherein, If the personalized data includes the user's identifier image, then the feature extraction of the personalized data for each of the multiple modalities includes: For the user's identification image, at least one of the image semantic features, color scheme features, style features, and text features of the identification image is extracted.

10. An electronic device, comprising: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform the operation corresponding to the method as described in any one of claims 1-9.

11. A computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any one of claims 1-9.

12. A computer program product comprising computer instructions that instruct a computing device to perform an operation corresponding to any one of the methods described in claims 1-9.

Citation Information

Patent Citations

  • Analysis template generation method and device and financial report comment template generation method and device

    CN118171642A

  • Automatic generation of transformations of formatted templates using deep learning modeling

    US20220147702A1