Content generation method, electronic equipment, computer storage medium and program product
By obtaining personalized data of multiple modalities of users, performing feature extraction and machine learning model selection, and generating personalized templates that conform to users' personalities, solving the problem that existing AIGC tools cannot generate personalized content and improving user experience.
Patent Information
- Application Number
- CN202510325193.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-03-18
AI Technical Summary
Existing AIGC tools are unable to generate personalized content that matches user personality, resulting in poor user experience.
By obtaining personalized data of multiple modalities of the user, performing feature extraction, generating template description information, and selecting appropriate models from multiple machine learning models to generate personalized template components, and finally generating personalized templates that conform to the user's personality.
It improves the personalization of content generated by AIGC tools and improves the user experience.
Smart Images

Figure CN120407913A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a content generation method, an electronic device, a computer storage medium, and a computer program product. Background Art
[0002] With the development of AI (Artificial Intelligence) technology, AIGC (Artificial Intelligence Generated Content) has been increasingly applied to people's work and life. AIGC is a technology that uses AI technology to automatically generate various types of content, including text, images, audio, video, etc. It is usually based on deep learning and natural language processing and can create content with a certain quality without direct human intervention.
[0003] Although current AIGC tools can provide the above content generation services to users, they only have general content generation capabilities. That is, if different users input the same requirements, the AIGC tools will generate the same or similar content for these different users, rather than generating personalized content that conforms to the individuality of users according to the differences of users. Summary of the Invention
[0004] In view of this, embodiments of the present application provide a content generation solution to at least partially solve the above problems.
[0005] According to a first aspect of embodiments of the present application, a content generation method is provided, including: obtaining personalized data of multiple modalities of a user and information of a personalized template to be generated; respectively performing feature extraction on the personalized data of the multiple modalities to obtain feature data corresponding to each modality; generating template description information based on the feature data of each modality and the information of the personalized template; selecting a corresponding machine learning model from multiple machine learning models for generating different template components based on the template description information; using the selected machine learning model to generate a corresponding template component, and generating the personalized template based on the generated template component.
[0006] According to a second aspect of embodiments of the present application, an electronic device is provided, including: a processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the method in the first aspect.
[0007] According to the third aspect of the embodiments of the present application, a computer storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method described in the first aspect is implemented.
[0008] According to the fourth aspect of the embodiments of the present application, a computer program product is provided, including computer instructions, and the computer instructions instruct a computing device to execute operations corresponding to the method described in the first aspect.
[0009] According to the content generation solution provided by the embodiments of the present application, personalized data of multiple modalities of a user and information of a personalized template to be generated are obtained, feature extraction is respectively performed on the personalized data of multiple modalities, feature data corresponding to each modality is obtained, and a more comprehensive and rich user personality description is formed; based on the feature data of each modality and the information of the personalized template, template description information is generated, so as to describe a personalized template that meets the personalized needs of the user; further, based on the template description information, a corresponding machine learning model is selected from multiple machine learning models for generating different template components, the selected machine learning model is used to generate the corresponding template component, and a personalized template is generated based on the generated template component. Thus, by analyzing the personalized data of multiple modalities of the user, personalized information in multiple dimensions of the user is obtained, and the template components generated based on these personalized information by different types of machine learning models all meet the personalized needs of the user, so that the personalized template generated based on these template components as a whole conforms to the personality of the user, and the personalization degree of the content generated for the user is improved.
[0010] It can be seen that through the solution of the embodiments of the present application, the personalization degree of the content generated by the AIGC tool for the user is greatly improved, and the use experience of the user for the AIGC tool is enhanced. Description of the Drawings
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0012] Figure 1 Schematic diagram of an exemplary system for applying the content generation method of the embodiments of the present application;
[0013] Figure 2A Flowchart of the steps of a content generation method according to an embodiment of the present application;
[0014] Figure 2B For Figure 2A Schematic diagram of a scenario example in the shown embodiment; Figure 3Schematic structural diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0015] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the embodiments of the present application.
[0016] The following further illustrates the specific implementation of the embodiments of the present application with reference to the accompanying drawings of the embodiments of the present application.
[0017] Figure 1 An exemplary system applicable to the content generation method according to the embodiment of the present application is shown. As Figure 1 shown, the system 100 may include a cloud server 102, a communication network 104, and / or one or more user devices 106, Figure 1 exemplified as multiple user devices herein.
[0018] The cloud server 102 may be any suitable device for storing information, data, programs, and / or any other suitable type of content, including but not limited to distributed storage system devices, server clusters, computing cloud server clusters, etc. In some embodiments, the cloud server 102 may perform any suitable functions. For example, in some embodiments, multiple machine learning models are set in the cloud server 102 to generate personalized content, including: obtaining personalized data of multiple modalities of a user and information of a personalized template to be generated; respectively performing feature extraction on the personalized data of multiple modalities to obtain feature data corresponding to each modality; generating template description information based on the feature data of each modality and the information of the personalized template; selecting a corresponding machine learning model from multiple machine learning models for generating different template components based on the template description information; using the selected machine learning model to generate a corresponding template component, and generating a personalized template based on the generated template component.
[0019] In some embodiments, the communication network 104 can be any suitable combination of one or more wired and / or wireless networks. For example, the communication network 104 can include any one or more of the following: the Internet, an intranet, a Wide Area Network (WAN), a Local Area Network (LAN), a wireless network, a Digital Subscriber Line (DSL) network, a Frame Relay network, an Asynchronous Transfer Mode (ATM) network, a Virtual Private Network (VPN), and / or any other suitable communication network. The user device 106 can be connected to the communication network 104 via one or more communication links (e.g., communication link 112), and the communication network 104 can be linked to the cloud server 102 via one or more communication links (e.g., communication link 114). The communication link can be any communication link suitable for transmitting data between the user device 106 and the cloud server 102, such as a network link, a dial-up link, a wireless link, a hard-wired link, any other suitable communication link, or any suitable combination of such links.
[0020] The user device 106 can include any one or more user devices suitable for interacting with the user. In some embodiments, the user device 106 can perform information input, such as inputting information of a personalized template, etc., and can display the personalized template. In some embodiments, the user device 106 can include any suitable type of device. For example, in some embodiments, the user device 106 can include a mobile device, a tablet computer, a laptop computer, a desktop computer, and / or any other suitable type of user device.
[0021] Based on the above system, an embodiment of the present application provides a content generation solution, which will be described below through multiple embodiments.
[0022] Referring to Figure 2A , a step flowchart of a content generation method according to an embodiment of the present application is shown. The content generation method of this embodiment includes the following steps:
[0023] Step S202: Obtain personalized data of multiple modalities of the user and information of the personalized template to be generated.
[0024] In the embodiments of the present application, in order to make the subsequent generated templates more in line with the personalized needs of users, multiple-modal personalized data of the user and information about the personalized template to be generated, that is, the personalized template to be generated, will be obtained first. Among them, the multiple-modal personalized data can be personalized data in various different data forms. For example, data in text form, data in video form, data in image form, data in audio form, and so on. Through these different-modal personalized data, the information and characteristics of the user can be effectively characterized from multiple dimensions, thus providing a basis for generating personalized templates that meet the personalized needs of users.
[0025] In an alternative manner, the multiple-modal personalized data of the user includes but is not limited to at least one of the following: the user's identification image, text data of the industry field to which the user belongs, the user's network information data, the user's template style preference data, the user's personalized audio data, and the user's personalized video data.
[0026] Among them, the user's identification image can be any image that can clearly identify and represent the user, and is used to identify and distinguish the user in different scenarios. For example, the user's identification image can be the user's avatar on a social media platform, a person poster, etc. For another example, for an enterprise, institution, brand, etc., the identification image can be the Logo (emblem) of this enterprise, this institution or this brand, and the design elements in the Logo can convey the attributes of products, service contents, the industry to which they belong, etc., which are content with identification. The user's identification image can more intuitively display the user's personality, style, etc., and improve the user's recognizability. Taking the user's identification image as one of the multiple-modal personalized data of the user can also make the generated template more recognizable when needed in the subsequent template generation.
[0027] The text data of the industry field to which the user belongs can include but is not limited to any data that reflects the user's industry in text form (including data published by the user himself and / or data related to the user published by other parties), and is usually related to the user's professional and technical fields, with more professional characteristics. For example, industry promotion content, industry reports, content related to industry characteristics, etc. of the industry field to which the user belongs. The text data of the industry field to which the user belongs reflects the user's professional background and possible work content, research direction, etc., and helps the AIGC tool understand the user's specific needs for the industry field to which the user belongs, so as to provide more accurate personalized content for the user.
[0028] The user's network information data can represent data related to the user's network behavior. For example, advertising data, evaluation report data, hot behavior data, etc. for the user. These data may also involve the industry where the user is located, but the degree of professionalism is weaker than that of the text data of the industry to which the user belongs. Through the user's network information data, the user's behavior on the network can be reflected, thus more intuitively reflecting information such as the user's individual needs and the degree of being recognized and concerned.
[0029] The user's template style preference data can represent the user's preferences for templates, including but not limited to preference data for template styles such as the color tone, layout style, background material, etc. of the template. The user's template style preference data can reflect the personalized needs of the user when using templates or template generation tools, and has different personalized characteristics among different users. It should be noted that in the embodiments of the present application, a "template" can be any template applicable to the solutions of the embodiments of the present application, such as web page templates, report templates, PPT templates, report forms templates, etc. in any appropriate form, and the embodiments of the present application do not limit the specific implementation form of the "template".
[0030] The user's personalized audio data can be audio data that can identify and represent the user. For example, a piece of voice that can effectively represent and identify the user, a promotional audio of an enterprise unit, a news report audio, etc. However, it is not limited to this, and it can also be the user's preference data for audio. For example, the user's preference data for audio content, including but not limited to the style type, creator, etc. of the audio content; for another example, the user's behavioral preferences for listening to audio, such as the number of times of listening to a certain or certain audios, the listening duration, etc.
[0031] The user's personalized video data can be video data that can identify the user. For example, a piece of video that can effectively represent and identify the user, a promotional video of an enterprise unit, a product demonstration video, etc. However, it is not limited to this, and it can also be the user's preference data for video. For example, the user's preference data for video content, including but not limited to the style type, creator, video duration, etc. of the video content; for another example, the user's behavioral preference data for watching videos, including but not limited to the number of times of watching a certain video or certain videos, the watching duration, etc.
[0032] The above-mentioned personalized data of the user can all be used as a reference and basis when generating a personalized template subsequently.
[0033] The information of the personalized template to be generated is used to effectively describe what kind of template the user wants to generate, and it includes at least the content information of the personalized template to be generated, that is, the content that the personalized template needs to present. For example, the content information of a personalized PPT template to be generated may include information indicating images, texts, etc. for presenting the user's product introduction. Since the content information needs to be carried by components, after the content information of the personalized template is determined, it also means the determination of the required components and the content carried by these components. Still taking the above PPT template as an example, if the PPT template that the user wants to generate contains images and texts for product introduction, it means that this template needs to generate an image component and a text component. Among them, the image component is used to carry and display the product introduction image, and the text component is used to carry and display the product introduction text, and so on. Therefore, it can also be considered that in the subsequent process, corresponding content can be generated using, for example, machine learning models, such as generating images using machine learning models, generating texts using machine learning models, and so on.
[0034] Although the corresponding personalized template can already be generated based on the content information. However, in order to make the personalized template more in line with the user's preferences and styles, in an optional way, the information of the personalized template can also include the layout information of the personalized template. Generally speaking, a template is usually formed by template components, that is, a personalized template can include at least one template component, and different template components carry different types of content, such as a text component for carrying text to display text, an image component for carrying an image to display an image, an interaction component for receiving interaction operations to perform human-computer interaction, and so on. Therefore, the layout structure of the personalized template to be generated, the layout relationship between each template component, etc. can be described through the layout information. For example, the position, size, nesting relationship, etc. of multiple template components in the personalized template to be generated. Although, based only on the above content information, a personalized template can also be generated through the default layout information, but by the way of the user inputting the layout information of the personalized template, the subsequent generated personalized template can better meet the user's personalized needs.
[0035] In addition, in an optional way, in order to obtain the information of the personalized template to be generated, a human-computer interaction interface for inputting the information of the personalized template can be provided in the AIGC tool. This human-computer interaction interface can be used to obtain the user's requirements for the personalized template to be generated. In this human-computer interaction interface, those skilled in the art can set any settings for the user to input and / or select according to the requirements. By operating these settings, the user can generate the information of the personalized template he hopes to generate. These settings include but are not limited to: template name setting, template content setting, setting of the components included in the template, setting of the layout of the template, setting of the style of the template, and so on.
[0036] It should be noted that when the template component includes components such as interaction items, in an optional manner, the human-computer interaction interface can set the settings for the interaction items to meet the user's interaction requirements for some components in the to-be-generated personalized template.
[0037] For example, in an optional manner, the interaction item includes a first interaction item. The first interaction item can be implemented as an interaction item that receives at least one of text content, voice content, video content, or image content input by the user. Through the first interaction item, the generated personalized template can have a human-computer interaction function. For another example, the first interaction item can also be implemented as an interaction item that sets corresponding functions for the template components in the personalized template to achieve interaction between components. For example, a button with a jump function, when the user clicks the button, can jump from the current component to another component, thereby achieving interaction between components. When specifically setting, the relevant setting information of the first interaction item can be displayed on the human-computer interaction interface, including but not limited to: the identifier, name, type, position, interaction input (the input methods that can be received, such as the aforementioned text input or button click, etc.), and interaction response (the response after receiving the interaction input, such as displaying the input at the corresponding position in the template, or jumping to other components, etc.). In practical applications, those skilled in the art can also make other settings according to actual needs, which are all within the protection scope of the embodiments of this application.
[0038] In addition, optionally, the interaction items in the human-computer interaction interface can also include at least one of the following: a second interaction item for the user to select the type of the to-be-generated personalized template; a third interaction item for the user to select the components included in the to-be-generated personalized template; a fourth interaction item for the user to set the layout of the to-be-generated personalized template; and a fifth interaction item for setting the content presented by the components included in the to-be-generated personalized template.
[0039] By operating on the first interaction item, the second interaction item, the third interaction item, the fourth interaction item, and the fifth interaction item, information about the personalized template can be generated.
[0040] Step S204: Respectively perform feature extraction on the personalized data of multiple modalities to obtain the feature data corresponding to each modality.
[0041] In order to extract key information from personalized data in multiple modalities and reduce the complexity of the original data, in a feasible approach, the personalized data in multiple modalities can be separately subjected to feature extraction to obtain the feature data corresponding to each modality, thereby more accurately reflecting the personalized needs and preferences of the user. Thus, the feature data corresponding to each modality can be applied to the personalized template, making the personalized template generated based on the feature data corresponding to each modality more in line with the user's personality. Exemplarily, separately performing feature extraction on the personalized data in multiple modalities can be implemented based on a machine learning model. However, this is not limited thereto. Separately performing feature extraction on the personalized data in multiple modalities can also be implemented based on other appropriate means such as statistical methods and data dimensionality reduction techniques. The embodiments of the present application do not limit the specific implementation manner of feature extraction.
[0042] It should be noted that the extraction of image features can be achieved from multiple dimensions. Moreover, due to the identification function of the user's identification image, in a feasible approach, if the personalized data includes the user's identification image, separately performing feature extraction on the personalized data in multiple modalities includes: for the user's identification image, performing feature extraction on at least one of the image semantic features, color matching features, style features, and text features of the identification image.
[0043] The image semantic features of the identification image refer to the image information features contained in the identification image that can be understood by humans, enabling the computer to understand and interpret the content of the image. In one example, when extracting the image semantic features of the identification image, the semantic features of the identification image can be extracted based on a machine learning model. Exemplarily, a deep learning model can be used, such as a convolutional neural network model, a semantic segmentation model, etc. Optionally, for the obtained extraction results, methods such as principal component analysis and linear discriminant analysis can be further used to further process the extraction results to improve the effectiveness of the image semantic features of the extracted identification image.
[0044] The color matching features of the identification image describe the color matching rules of the identification image. Based on the color matching features of the identification image, the visual expression of the identification image can be perceived, and thus the identification image can be better described. For example, the color matching features of the identification image are "black and white color matching", "blue and white color matching", etc. In one example, extracting the color matching features of the identification image can be implemented as extracting the color matching features of the identification image based on a machine learning model. However, this is not limited thereto. It is also possible to calculate the histograms of the respective color components in the identification image based on image processing software such as Photoshop to reflect the composition and distribution of each color in the identification image.
[0045] The style features of the identification image refer to the attributes that can reflect the atmosphere characteristics of the identification image, and these attributes make the identification image have higher visual recognition. For example, the style features of the identification image can be "oil painting style", "retro style", "future style", etc. In one instance, extracting the style features of the identification image can be implemented by extracting based on a machine learning model in the identification image, such as a deep learning model like a convolutional neural network. However, it is not limited to this, and the style features of the identification image can also be extracted based on statistical methods such as histogram of oriented gradients and scale-invariant feature transform.
[0046] The text features of the identification image refer to the characteristics of the text content contained in the identification image. For example, the content of the text, font size, layout, etc. In one instance, extracting the text features of the identification image can be implemented by extracting based on optical character recognition technology and machine learning models, such as deep learning models like convolutional neural networks, etc. in the identification image.
[0047] By extracting these features, it can provide a basis for generating an identification image that can effectively identify the user and has an obvious deformation from the original identification image in the subsequent process.
[0048] In addition, optionally, before extracting the features of the user's identification image, the identification image can also be preprocessed, such as performing enhancement processing, denoising processing, etc. on the identification image to improve the image quality of the identification image, thereby improving the reliability of the extracted features.
[0049] If the user's personalized data includes text data in the industry field to which the user belongs, in a feasible way, when extracting features from the text data in the industry field to which the user belongs, a machine learning model, such as a deep learning model like a convolutional neural network, can be used to learn and extract the feature representation in the text data to obtain the corresponding text features. However, it is not limited to this, and other ways of extracting text features are also applicable. For example, the TF-IDF (term frequency–inverse document frequency) method can be used to measure the importance of each word in the text data and thus be used as one of the features of the text data.
[0050] Optionally, before extracting the features of the text data in the industry field to which the user belongs, the text data can also be preprocessed first. Exemplarily, preprocessing the text data can include: segmenting the text data or removing stop words, etc., so as to reduce the difficulty of subsequent feature extraction.
[0051] If the personalized data includes the user's network information data, when extracting features from the user's network information data, in one feasible way, a machine learning model can be used to extract features from the user's network information data, and the extraction result can be used as the feature data of the user's network information data. For example, deep learning models such as convolutional neural network models and recurrent neural network models can be used to extract features from the user's network information data. In another feasible way, the method of topic modeling can also be used to extract the topic structure in the user's network information data, and the extraction result can be used as the feature data of the user's network information data.
[0052] If the personalized data includes the user's template style preference data, in one feasible way, when extracting features from the user's template style preference data, it can be implemented based on a machine learning model, or a probability statistical model, or a frequency distribution model that counts the number of times each template style is used, etc.
[0053] If the personalized data includes the user's personalized audio data, in one feasible way, a machine learning model can be used to extract features from the audio data, so as to extract the feature data in the user's personalized audio data. In another feasible way, based on speech recognition technology, the specific content in the personalized audio data can be converted into text format, and further, the key information in the personalized audio data can be extracted based on the content in text format as the feature data of the user's personalized audio data. However, it is not limited to this. In one feasible way, when extracting features from the user's personalized audio data, it can also be implemented by analyzing the speech rate, intonation, volume, energy, etc. in the personalized audio data through audio processing software such as Adobe Audition and Audacity, and extracting the speaker's speech style features, emotional features, etc. in the personalized audio data based on the analysis results as the feature data of the user's personalized audio data.
[0054] If the personalized data includes the user's personalized video data, in one feasible way, when extracting features from the user's personalized video, the personalized video can be first split into multiple video frames, and then features can be extracted. In an example, a machine learning model used for feature extraction of video frames can be used to learn the video frames of the user's personalized video data, so as to extract the feature data in the video frames. However, it is not limited to this. In another example, based on multiple video frames, methods such as drawing color histograms and calculating color moments can be used to extract the color features contained in the video, and methods such as edge detection and contour tracking can also be used to capture the object shape, contour information, etc. in the video frames, and in addition, the motion information of the objects in the personalized video can be detected by comparing the pixel differences between adjacent video frames as the feature data of the user's personalized video data.
[0055] Thus, by extracting the feature data from the personalized data of multiple modalities of the user, personalized features of the user for the template can be obtained from multiple dimensions, providing a basis for generating a personalized template that better meets the user's needs.
[0056] Step S206: Generate template description information based on the feature data of each modality and the information of the personalized template.
[0057] Among them, the template description information can be implemented in the form of a data object or a custom data structure, which at least contains the feature data of each of the above modalities and the information of the personalized template, so as to facilitate subsequent processing and use.
[0058] In one example, the template description information can be implemented in any appropriate form such as an XML (Extensible Markup Language) data structure or a JSON object. Thus, it can not only effectively carry the feature data and the information of the personalized template, but also facilitate the adjustment of the template description information to more flexibly express and implement the information content, and is more interpretable.
[0059] Moreover, in the embodiments of the present application, the personalized template is implemented through template components, and the template components can be generated based on machine learning models. That is, multiple machine learning models for generating different template components can be set. In this case, in a feasible manner, the information of the machine learning model to be used can also be carried in the template description information. That is, generating the template description information based on the feature data of each modality and the information of the personalized template can be implemented as: generating the template description information based on the feature data of each modality, the information of the personalized template, and the obtained information of the machine learning model to be used. The information of the machine learning model to be used can be input or set by the user through the aforementioned human-computer interaction interface, or can be implemented through an automated selection method.
[0060] Based on this, in a feasible manner, the information of the machine learning model to be used is obtained in the following way: obtaining the user's historical machine learning model usage data, and obtaining the information of the machine learning model to be used according to the historical machine learning model usage data; or receiving the information of the machine learning model to be used input by the user through the human-computer interaction interface.
[0061] When the user generates a personalized template in the previous period, corresponding model usage data will be left. By analyzing this data, the user's model usage preferences (including model usage preferences when generating various model components, etc.) can be obtained. Based on this, when the user generates a personalized template again, the model preferred by the user can be used for generation. However, it is not limited to this. The historical machine model usage data can also be the data of the models used by a large number of users when generating different personalized templates. By analyzing this data, when a certain user (who may be a new user or an old user) generates a personalized template, information about the machine learning model to be used can be obtained based on the analysis of the data of the above-mentioned large number of users.
[0062] Thus, through the information about the machine learning model to be used, the relevant machine learning model can be quickly and efficiently called to generate the components of the personalized template that conforms to the user's personality, and then the personalized template can be generated.
[0063] Step S208: Based on the template description information, select the corresponding machine learning model from multiple machine learning models for generating different template components.
[0064] In order to improve the efficiency and accuracy of generating template components that meet the user's personalized needs and improve the degree of personalization, in the solution of the embodiment of the present application, multiple machine learning models for generating different template components are preset.
[0065] In a feasible manner, the multiple machine learning models for generating different template components include multiple types of machine learning models for generating different types of components in the template, and each type of machine learning model includes at least one machine learning model. Exemplarily, it includes four types of machine learning models, which are respectively used to generate text components, image components, video components, and audio components; for each of these four types of machine learning models, it can further include multiple different machine learning models. Although these multiple different machine learning models can all generate template components of the same type, their model characteristics are different. For example, model A is suitable for generating content images, and model B is suitable for generating Logo images, etc.
[0066] Based on this, in a feasible manner, selecting the corresponding machine learning model from multiple machine learning models for generating different template components based on the template description information includes: determining the information of the components in the personalized template to be generated based on the template description information; determining the type of the machine learning model to be used according to the information of the components; and selecting the target machine learning model to be used from each type of machine learning model determined.
[0067] By means of generation by different types of components based on corresponding types of machine learning models, each type of machine learning model is generated for the characteristics of a specific type of template component, with stronger pertinence, and can process relevant data more quickly and accurately to generate the required template components.
[0068] Further optionally, after confirming the corresponding type of machine learning model to be used based on the template component to be generated, the usage preferences of the user for the machine learning model during historical use can be further obtained based on the user's personalized data, so as to select a machine learning model that better conforms to the user's preferences from the pre-set machine learning models, in order to generate components of a more personalized template that better conforms to the user's personality in the subsequent process.
[0069] Exemplarily, when the information of the personalized template includes the content information of the personalized template, the types of multiple types of machine learning models include but are not limited to at least one of the following: the first type of machine learning model for generating a content image (excluding the logo image) that matches the content information and the user's personalized data; the second type of machine learning model for generating text that matches the content information and the user's personalized data; the third type of machine learning model for generating a logo image that matches the content information and the user's personalized data; the fourth type of machine learning model for generating a video that matches the content information and the user's personalized data; the fifth type of machine learning model for generating an interaction item that matches the content information and the user's personalized data.
[0070] In one example, an image can be generated by the first type of machine learning model based on the description of the content image to be generated in the template description information. The first type of machine learning model can be any text-to-image model, including but not limited to LLM (Large Language Model), Stable Diff, etc. However, this is not limited to this. The description of the content image to be generated in the template description information can also be information related to the image, including but not limited to the link of the image, the identifier of the image, or the image itself, or any information that can obtain the image. In this case, when generating the content image by the first type of machine learning model, the corresponding image can be obtained according to the template description information as a reference image first, and then the content image can be generated based on the reference image. In this case, exemplarily, the first type of machine learning model can be implemented in the form of, for example, a GAN (Generative Adversarial Network) model, a Stable Diffusion model, etc., to generate a content image that conforms to the user's personality for the user by combining the user's personalized data. Thus, the function of generating the component of the image type in the personalized template by the first type of machine learning model is realized.
[0071] In another example, text can be generated by a second type of machine learning model based on the description of the text content to be generated in the template description information. Exemplarily, the second type of machine learning model can be implemented as an LSTM (Long Short-Term Memory) model, a model based on the Transformer architecture, or even an LLM model, etc., for generating personalized text for the user in combination with the user's personalized data. Thus, the second type of machine learning model can be used to generate components of the text type in the personalized template.
[0072] In yet another example, a third type of machine learning model can generate a logo image that meets the personalized needs of the user based on the description of the logo image to be generated in the template description information and the user's original logo image. For example, it can be generated through a GAN model or a Diffusion model. Thus, the third type of machine learning model can be used to generate the logo image in the personalized template.
[0073] In still another example, a video that conforms to the user's personality can be generated by a fourth type of machine learning model based on the description of the video to be generated in the template description information; or, a video that conforms to the user's personality can be generated by a fourth type of machine learning model based on the description of the video to be generated in the template description information and the reference video frame. Exemplarily, the fourth type of machine learning model can be implemented as an LVLM model, etc., for generating a video that conforms to the user's personality in combination with the user's personalized data. Thus, the fourth type of machine learning model can be used to generate components of the video type in the personalized template.
[0074] In another example, if the template components of the personalized template to be generated include interaction items, then the type of machine learning model to be used is determined according to the information of the components; selecting the target machine learning model to be used from each determined type of machine learning model can be implemented as follows: when it is determined according to the component information that the components of the personalized template to be generated include interaction items for human-computer interaction, determine that the type of machine learning model to be used is the fifth type of machine learning model, and select the target machine learning model to be used from the fifth type of machine learning model. In this case, this step of using the selected machine learning model to generate the corresponding template component can be implemented as follows: obtaining the interaction information corresponding to the interaction item from the template description information; generating the interaction item based on the interaction information and generating the corresponding interaction response information for the interaction item; or, determining the template component for responding to the interaction for the interaction item and associating the interaction item with the determined template component. The information of the personalized template may include relevant descriptions for the interaction item, such as the type of the interaction item (such as input text box, interaction button, etc.), the interaction chain (such as how to process the received input text, or how to process it after the interaction button is clicked, etc.). Based on this description, the fifth type of machine learning model is used to generate an interaction item that conforms to the user's personality, and the interaction item may be associated with corresponding code to implement its interaction function. Exemplarily, the fifth type of machine learning model can be implemented as an LLM, etc.
[0075] In the case where these machine learning models are set, when a template component needs to be generated, a suitable machine learning model can be selected from them. For example, a corresponding machine learning model can be selected based on the information of the machine learning model to be used carried in the template description information; or, according to the usage data of the user's historical use of the machine learning model, a machine learning model that better conforms to the user's usage habit can be selected.
[0076] It should be noted that each type of machine learning model includes at least one machine learning model to adapt to different types of detailed generation requirements while ensuring the robustness of the system.
[0077] Step S210: Use the selected machine learning model to generate the corresponding template component and generate a personalized template based on the generated template component.
[0078] After determining the specific machine learning model, these machine learning models can be used to generate the corresponding template components, and the specific generation method can refer to the foregoing description of each type of machine learning model.
[0079] After generating the template components through various types of machine learning models, a personalized template can be further generated based on the template components. In a feasible way, if the information of the personalized template includes the layout information of the personalized template, the personalized template can be generated based on this layout information.
[0080] In a feasible manner, the personalized template can also be generated in the form of a machine learning model, that is, through the machine learning model for generating the template, the personalized template is generated according to the generated template components and layout information. The machine learning model can adopt any appropriate form, including but not limited to a machine learning model dedicated to generating templates, or an LLM model, etc.
[0081] It should be noted that in multiple machine learning models of the embodiments of the present application, an LLM can be set. Due to the powerful generation function of the LLM, it can be used as the machine learning model for generating template components and also as the machine learning model for generating personalized templates. However, in addition to this, as mentioned above, multiple dedicated models for generating template components are also set in the embodiments of the present application. Thus, the combination of large (LLM) and small (models only for generating certain types of components) models is realized, which not only ensures the generation quality of the personalized model but also improves the robustness of model use.
[0082] It can be seen that through the embodiments of the present application, personalized data of multiple modalities of the user and information of the personalized template to be generated can be obtained, feature extraction is respectively performed on the personalized data of multiple modalities to obtain the feature data corresponding to each modality, and a more comprehensive and rich user personality description is formed; based on the feature data of each modality and the information of the personalized template, template description information is generated, so as to describe the personalized template that meets the personalized needs of the user; further, based on the template description information, the corresponding machine learning model is selected from multiple machine learning models for generating different template components, the selected machine learning model is used to generate the corresponding template components, and the personalized template is generated based on the generated template components. Thus, by analyzing the personalized data of multiple modalities of the user, personalized information in multiple dimensions of the user is obtained, and the template components generated based on these personalized information by different types of machine learning models all meet the personalized needs of the user, so that the overall personalized template generated based on these template components conforms to the user's personality, and the personalization degree of the content generated for the user is improved.
[0083] Hereinafter, taking a specific scenario as an example, with reference to Figure 2B , an exemplary description of the above content generation method of the embodiments of the present application will be given.
[0084] In this example, it is assumed that an art tutoring institution needs to generate a web page template for corporate promotion. Then, first, personalized data in multiple modalities of the art tutoring institution can be obtained. In this example, it includes the logo of the art tutoring institution, the performance report of the art tutoring institution, and the news publicity reports of the art tutoring institution. Also, information on the personalized template that the art tutoring institution wants to generate is obtained. For example, the user can input this information through a human-computer interaction interface, indicating to generate a personalized template including the institution logo, publicity text, and publicity images.
[0085] Furthermore, based on the above-mentioned personalized data in multiple modalities of the art tutoring institution, characteristic data corresponding to the personalized data in multiple modalities are respectively extracted. Exemplarily, machine learning models can be used to extract features from the personalized data in multiple modalities to obtain the corresponding characteristic data. For example, based on the logo of the art tutoring institution, semantic features of the image, color matching features of the combination of red, yellow, and blue, and style characteristic data of the oil painting style, etc. are extracted. Figure 2B It is shown as "logo image features" in the middle. For another example, based on the performance report, characteristic data such as ranking among the top N in XX year are extracted. Figure 2B It is shown as "industry features" in the middle. For another example, based on the news reports of the art tutoring institution, characteristic data of the deeds reports of the excellent teachers and excellent students of the art tutoring institution are extracted. Figure 2B It is shown as "network information features" in the middle. Furthermore, based on the above-mentioned characteristic data and the information of the personalized template, template description information can be generated.
[0086] Suppose the set machine learning models include the aforementioned five types of machine learning models. Then, in this example, since a personalized template including the institution logo, publicity text, and publicity images needs to be generated, therefore, 1 model included in the third type of machine learning model ("Model 3-1") can be selected, the first model among the 3 models included in the second type of machine learning model ("Model 2-1") can be selected, and the second model among the 2 models included in the first type of machine learning model ("Model 1-2") can be selected. The selected models are used to generate the components in the template indicated by the information of the personalized template, namely the institution logo, the institution publicity text, and the institution publicity image. In this example, it is set that after generating the template components, the personalized template is generated according to the default layout information. A simple example of a generated personalized template is shown in the figure, including a first image component carrying the institution logo, a first text component carrying the institution publicity text, and a second image component carrying the institution publicity image.
[0087] It can be seen that through this example, a content generation solution that better meets the personalized needs of users is provided. Based on the personalized data of users, templates that meet the personalized needs of users are generated, improving the personalization degree of the content generated for users.
[0088] Referring to Figure 3 , a schematic structural diagram of an electronic device according to an embodiment of the present application is shown. The specific implementation of the electronic device is not limited in the specific embodiments of the present application.
[0089] As Figure 3 shown, the electronic device may include: a processor 302, a communication interface 304, a memory 306, and a communication bus 308.
[0090] Wherein:
[0091] The processor 302, the communication interface 304, and the memory 306 communicate with each other through the communication bus 308.
[0092] The communication interface 304 is used to communicate with other electronic devices or servers.
[0093] The processor 302 is used to execute the program 310, and specifically can execute the relevant steps in the above-mentioned content generation method embodiment.
[0094] Specifically, the program 310 may include program code, and the program code includes computer operation instructions.
[0095] The processor 302 may be a CPU, or a GPU (Graphic Processing Unit), or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the electronic device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0096] The memory 306 is used to store the program 310. The memory 306 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0097] The program 310 may include multiple computer instructions. Specifically, the program 310 may cause the processor 302 to execute the operations corresponding to the content generation method described in any one of the foregoing multiple method embodiments through the multiple computer instructions.
[0098] For the specific implementation of each step in the program 310, reference may be made to the corresponding descriptions in the corresponding steps and units in the foregoing method embodiments, and there are corresponding beneficial effects, which will not be elaborated herein. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be repeated herein.
[0099] The embodiments of the present application further provide a computer storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method described in any one of the foregoing multiple method embodiments. The computer storage medium includes, but is not limited to: Compact Disc Read-Only Memory (CD-ROM), Random Access Memory (RAM), floppy disk, hard disk, magneto-optical disk, etc.
[0100] The embodiments of the present application further provide a computer program product, including computer instructions, which instruct a computing device to execute the operations corresponding to any content generation method in the foregoing multiple method embodiments.
[0101] In addition, it should be noted that the information related to users (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data for training the model, data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the users or fully authorized by all parties. And the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to select authorization or rejection.
[0102] It should be noted that according to the needs of implementation, each component / step described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of the components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.
[0103] The method according to the embodiments of the present application can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code that is originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and will be stored in a local recording medium. Thus, the method described herein can be stored in such a software process on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA)). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a Random Access Memory (RAM), a Read-Only Memory (ROM), a flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.
[0104] Those of ordinary skill in the art can realize that the units and method steps of the examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such an implementation should not be considered to exceed the scope of the embodiments of the present application.
[0105] The above embodiments are only used to illustrate the embodiments of the present application, rather than to limit the embodiments of the present application. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application. The patent protection scope of the embodiments of the present application shall be defined by the claims.
Claims
1. A content generation method, comprising: Obtaining personalized data of multiple modalities of a user and information of a personalized template to be generated; Respectively performing feature extraction on the personalized data of the multiple modalities to obtain feature data corresponding to each modality; Generating template description information based on the feature data of each modality and the information of the personalized template; Based on the template description information, selecting a corresponding machine learning model from multiple machine learning models for generating different template components; Using the selected machine learning model to generate a corresponding template component and generating the personalized template based on the generated template component.
2. The method according to claim 1, wherein The multiple machine learning models for generating different template components include multiple types of machine learning models for generating different types of components in the template, and each type of machine learning model includes at least one machine learning model; The selecting a corresponding machine learning model from multiple machine learning models for generating different template components based on the template description information includes: Based on the template description information, determining information of components in the personalized template to be generated; According to the information of the components, determining the type of the machine learning model to be used; Selecting a target machine learning model to be used from each determined type of machine learning model.
3. The method according to claim 2, wherein, The information of the personalized template at least includes content information of the personalized template; The types of the multiple types of machine learning models include at least one of the following: The first type of machine learning model for generating a content image matching the content information and the personalized data of the user; The second type of machine learning model for generating text matching the content information and the personalized data of the user; The third type of machine learning model for generating an identification image matching the content information and the personalized data of the user, wherein the third type of machine learning model is trained based on an identification image sample of the user; The fourth type of machine learning model for generating a video matching the content information and the personalized data of the user; The fifth type of machine learning model for generating an interaction item matching the content information and the personalized data of the user.
4. The method according to claim 3, wherein, The determining the type of the machine learning model to be used according to the information of the components; selecting a target machine learning model to be used from each determined type of machine learning model includes: if it is determined according to the information of the components that the components in the personalized template to be generated include an interaction item for human-computer interaction, determining that the type of the machine learning model to be used is the fifth type of machine learning model, and selecting a target machine learning model to be used from the fifth type of machine learning model; Said generating a corresponding template component using the selected machine learning model includes: obtaining interaction information corresponding to the interaction item from the template description information; generating an interaction item based on the interaction information and generating corresponding interaction response information for the interaction item; or determining a template component for responding to the interaction for the interaction item and associating the interaction item with the determined template component.
5. The method according to claim 3, wherein The information of the personalized template further includes layout information of the personalized template; Said generating the personalized template based on the generated template component includes: generating the personalized template according to the generated template component and the layout information through a machine learning model for generating a template.
6. The method according to claim 1, wherein said generating template description information based on the feature data of each modality and the information of the personalized template includes: generating template description information based on the feature data of each modality, the information of the personalized template, and the information of the machine learning model to be used obtained; said selecting a corresponding machine learning model from multiple machine learning models for generating different template components based on the template description information includes: selecting a corresponding machine learning model from multiple machine learning models for generating different template components based on the information of the machine learning model to be used in the template description information.
7. The method according to claim 6, wherein The information of the machine learning model to be used is obtained by the following method: obtaining historical machine learning model usage data of the user and obtaining the information of the machine learning model to be used according to the historical machine learning model usage data; or receiving the information of the machine learning model to be used input by the user through a man-machine interface.
8. The method according to claim 1, wherein, The personalized data of multiple modalities of the user includes at least one of the following: the user's identification image, text data of the industry field to which the user belongs, the user's network information data, the user's template style preference data, the user's personalized audio data, the user's personalized video data.
9. The method according to claim 8, wherein, If the personalized data includes the user's identification image, then said respectively performing feature extraction on the personalized data of multiple modalities includes: performing feature extraction of at least one of the image semantic feature, color matching feature, style feature, and text feature of the identification image for the user's identification image.
10. An electronic device, comprising: A processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used for storing at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the method according to any one of claims 1-9.
11. A computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method according to any one of claims 1-9.
12. A computer program product, including computer instructions, and the computer instructions instruct a computing device to perform the operations corresponding to the method according to any one of claims 1-9.
Citation Information
Patent Citations
Analysis template generation method and device and financial report comment template generation method and device
CN118171642A
Automatic generation of transformations of formatted templates using deep learning modeling
US20220147702A1
Method for generating personalized product description based on multi-source crowd data
US20220245676A1
Systems and methods for generating personalized content items
US20220414754A1
Image processing method and apparatus, computer device, and storage medium
WO2025036359A1