Content generation method and device, storage medium and electronic equipment
By acquiring user comments from existing technologies and using large-scale language and image generation models to generate multimodal comment display content, the problem of single-presentation of user comments is solved, enabling interaction between users and the content of the book or readers, and enhancing the reading experience.
Patent Information
- Application Number
- CN202511449930.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2026-01-16
AI Technical Summary
In existing technologies, user comments can only be displayed in text boxes, which is a limited presentation method and prevents users from interacting with the content, thus affecting the reading experience.
By acquiring user comments, extracting prompt words, and inputting them into a large language model and image generation model for semantic analysis and image generation, multimodal comment display content is generated, including reply content and reply images.
It enhances the fun and interactivity of comments, enabling users to interact directly with the book's content or other readers, thus improving the user experience.
Smart Images

Figure CN121352001A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a content generation method, apparatus, storage medium and electronic device. Background Technology
[0002] Book reviews can facilitate communication between book editors and readers, influence and guide readers' reading activities, and promote and improve the quality of book editing, writing, and publishing.
[0003] Currently, the main approach is to directly identify the user's comment text or emoji when the user enters a comment for a certain chapter or section, and then display the user's comment in the corresponding text box.
[0004] However, this content generation method can only display the comments or emojis entered by different users in the text box. The presentation method is limited, and users cannot get corresponding replies by commenting on the content, which affects the user's reading experience. Summary of the Invention
[0005] In view of this, this application provides a content generation method, apparatus, storage medium, and electronic device, the main purpose of which is to improve the technical problem that the existing technology can only present the comment text or emoticons entered by different users in the text box, the presentation method is monotonous and users cannot get corresponding replies by commenting on the content, which affects the user reading experience.
[0006] Firstly, this application provides a content generation method, including: Get the user's comment content based on the target content; Based on the target content and the comment content, prompt words are extracted; The extracted prompt words are input into the first model and the second model respectively. Based on the prompt words, semantic analysis is performed in the first model to obtain the reply content of the comment. The screen is generated in the second model to obtain the reply screen corresponding to the reply content. Based on the comment content, the preset comment screen corresponding to the comment content, the reply content, and the reply screen, the target comment display content for the user is generated.
[0007] Secondly, this application provides a content generation apparatus, comprising: The acquisition module is configured to acquire the comment content entered by the user based on the target content; The extraction module is configured to extract prompt words based on the target content and the comment content; The input module is configured to input multiple extracted prompt words into a first large model and a second large model respectively, perform semantic analysis on the first large model based on the multiple prompt words to obtain the reply content of the comment content, and perform image generation on the second large model to obtain the reply image corresponding to the reply content; The generation module is configured to generate the user's target comment display content based on the comment content, the preset comment screen corresponding to the comment content, the reply content, and the reply screen.
[0008] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the content generation method described in the first aspect.
[0009] Fourthly, this application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the content generation method described in the first aspect.
[0010] Fifthly, embodiments of this application provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the content generation method described in the first aspect.
[0011] By employing the above technical solutions, this application provides a content generation method, apparatus, storage medium, and electronic device. Compared with existing technologies, this application obtains comment content input by a user based on target content; extracts prompt words based on the target content and comment content; inputs multiple extracted prompt words into a first model and a second model respectively; performs semantic analysis on the multiple prompt words in the first model to obtain reply content to the comment content; and generates a screen in the second model to obtain a reply screen corresponding to the reply content; and generates the user's target comment display content based on the comment content, the preset comment screen corresponding to the comment content, the reply content, and the reply screen. This allows the application to generate corresponding reply content and reply screen based on the comment content input by the user. Combining the comment content and the preset comment screen allows for multiple modal comment content to be displayed after the user inputs a comment, and the comment content and reply content can be displayed in combination with different screens, enhancing the interest of the comment. By generating comment display content based on the reply content and comment content, users can quickly receive corresponding replies after inputting comment content, improving interactivity and allowing users to directly interact with the content in the book or with comments between readers, thereby improving the user experience. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 A flowchart illustrating a content generation method provided in an embodiment of this application is shown; Figure 2 A flowchart illustrating a content generation method provided in an embodiment of this application is shown; Figure 3 A schematic diagram illustrating an example provided in an embodiment of this application is shown; Figure 4 A schematic diagram illustrating an example provided in an embodiment of this application is shown; Figure 5 A schematic diagram illustrating an example provided in an embodiment of this application is shown; Figure 6 This illustration shows a structural schematic diagram of a content generation apparatus provided in an embodiment of this application; Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0015] The embodiments of this application will now be described in more detail with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0016] To address the technical problem that existing technologies can only display user-inputted comments or emoticons in text boxes, resulting in a limited presentation method and the inability for users to receive responses by commenting on the content, thus negatively impacting the user reading experience, this embodiment provides a content generation method, such as... Figure 1 As shown, the method includes: Step 101: Obtain the comment content entered by the user based on the target content.
[0017] In this embodiment of the application, the target content can be content that the user is reading, content that the user is browsing, content that the user wants to comment on, etc.
[0018] In some examples, the comment content can be the content that the user wants to comment on. For example, if a user wants to comment "content a" on the first paragraph of article 1 while reading article 1, then the first paragraph of article 1 can be the target content in this application embodiment, and content a can be the comment content in this application embodiment.
[0019] Step 102: Extract prompt words based on the target content and comment content.
[0020] In this application embodiment, prompt extraction can be the process of automatically identifying and extracting keywords or phrases that have a key impact on the generated content from a piece of text, user input, history, or dialogue. These prompt keywords / phrases are typically used to guide AI models (such as Large Language Models (LLM), image generation models, etc.) to generate content that meets expectations.
[0021] In some examples, a prompt is a command or description input by the user into the AI system to guide the AI in generating a specific type of output. Correspondingly, prompt extraction involves identifying keywords or expressions from a large amount of data that are highly instructive, semantically clear, and reusable. For example, if the input command is "Help me write an explanatory article about climate change, with a more formal tone," the corresponding prompt could be "climate change, explanatory article, formal tone." It should be noted that in this embodiment, the target content and comment content are the inputs for prompt extraction.
[0022] Step 103: Input the extracted prompt words into the first and second large models respectively. Perform semantic analysis on the first large model based on the prompt words to obtain the reply content of the comment content, and generate the screen in the second large model to obtain the reply screen corresponding to the reply content.
[0023] In this embodiment, the first large model can be a large model obtained by training semantic analysis, and the second large model can be a large model obtained by training image generation.
[0024] In some examples, semantic analysis is a crucial task in natural language processing (NLP), aiming to understand the meaning of text rather than simply recognizing surface-level lexical or grammatical structures. For instance, the first major model can be a large language model (LLM). Large language models (LLMs) are deep learning models with a very large parameter set, trained on massive corpora, and possessing powerful language understanding and generation capabilities. They can perform various NLP tasks, such as text generation, question answering, translation, summarization, and reasoning.
[0025] For example, large-scale language models are language models built on neural networks, typically using the Transformer architecture, and are pre-trained on large amounts of text data to gain a deep understanding of language structure, semantics, and logic. In the embodiments of this application, semantic analysis can be performed using large-scale language models to obtain the response content to the comment content.
[0026] In this embodiment, the second major model can be a deep learning model trained on a large-scale image dataset to ultimately obtain image generation capabilities. For example, the second major model in this embodiment can be an AI large model. Specifically, the AI large model can be a neural network model based on a deep learning architecture, trained using a large-scale corpus or data, which has strong generalization ability and multi-task adaptability. In this embodiment, the AI large model can be used to generate images and obtain the response images corresponding to the response content.
[0027] Step 104: Generate the user's target comment display content based on the comment content, the preset comment screen corresponding to the comment content, the reply content, and the reply screen.
[0028] In this embodiment of the application, the comment content is the comment content entered by the user, and the preset comment screen corresponding to the comment content can be the comment content screen preset by the user. For example, the comment content screen can be presented in different modalities according to different user settings, such as video, audio, static image, dynamic image, etc.
[0029] In some examples, the response content can be a system-generated response based on the prompt words, and the corresponding response screen can also be a system-generated response screen based on the prompt words.
[0030] It should be noted that the target comment display content generated by this application includes two parts: comment display content and reply display content. The comment display content includes a combination of comment content and a preset comment screen, while the reply display content includes a combination of reply content and a reply screen.
[0031] Compared with existing technologies, this embodiment obtains the comment content input by the user based on the target content; extracts prompt words based on the target content and comment content; inputs the extracted prompt words into a first model and a second model respectively; performs semantic analysis on the prompt words in the first model to obtain the reply content of the comment content; and generates a screen in the second model to obtain the reply screen corresponding to the reply content; and generates the user's target comment display content based on the comment content, the preset comment screen corresponding to the comment content, the reply content, and the reply screen. This embodiment can generate corresponding reply content and reply screen based on the comment content input by the user. Combining the comment content and the preset comment screen allows for multiple modal comment content to be displayed after the user inputs a comment, and can combine different screens to display the comment content and reply content, enhancing the interest of the comment; by generating the comment display content based on the reply content and comment content together, the user can quickly get the corresponding reply after inputting the comment content, improving the interactivity, allowing the user to directly interact with the content in the book or with comments between readers, thereby improving the user experience.
[0032] As a refinement and extension of the above embodiments, this embodiment provides a content generation method, such as... Figure 2 As shown, the method includes: Step 201: Obtain the comment content entered by the user based on the target content.
[0033] For example, taking user 1 reading article 1 as an example, when user 1 reads an interesting part, they can choose the current chapter or a certain paragraph in article 1 to post a comment expressing their current mood and thoughts; specifically, after user 1 submits a comment in the reading app of article 1, the client sends the comment to the server, and the server saves the comment.
[0034] Step 202: Obtain user profile information containing user reading characteristics.
[0035] In this embodiment of the application, the user profile information containing user reading characteristics can be constructed on the server side by analyzing data such as user reading behavior, interest preferences, and reading habits to create a digital profile that can represent the user's reading behavior characteristics.
[0036] For example, based on the example in step 201, the user profile information can be generated based on the user 1's historical reading and commenting information and stored on the server. That is, in this embodiment of the application, obtaining user profile information containing user reading characteristics can be done by obtaining user profile information 1 containing user 1's reading characteristics from the server.
[0037] Step 203: Determine the context and content category corresponding to the target content.
[0038] In this embodiment of the application, the content category can specifically be the classification of the article in which the target content is located. Different article classifications reflect different styles. For example, they can include urban, fantasy, martial arts, historical, suspense and reasoning, etc. It should be noted that the classification information can be a key factor in generating prompt words, reply content and reply screen.
[0039] In some examples, the contextual content can be the context of the target content's location. Specifically, the comment content is based on the user's actual feelings and experiences within the context of the article. It should be noted that the contextual content is the same as the content category and can be the data source for the prompt word extraction model. This provides the prompt word extraction model with literary world scenarios and contextual information for training and inference. Prompt words generated based on the context are more reasonable and resonate with the user.
[0040] For example, such as Figure 3 As shown, after a user submits comment data in the APP, the server saves the comment content and submits it as input to the large model (including the prompt word extraction model, the first large model and the second large model in this application embodiment). After being processed by the large model, multimodal content is output (i.e., the reply screen or target comment display content in this application embodiment).
[0041] It should be noted that the output of the large model in this embodiment can be a reply screen, and the server generates the target comment display content based on the reply screen and reply content and returns it to the client; the output of the large model in this embodiment can also be the target comment display content, and the server returns the target comment display content output by the large model to the client. That is, the target comment display content can be generated by the large model or by the server, and no specific limitation is made here.
[0042] Step 204: Input user profile information, context content, content category and comment content into the prompt word extraction model. Extract prompt words from the reading feature dimension and reading content dimension in the prompt word extraction model to obtain multiple prompt words corresponding to multiple attributes and priority information corresponding to multiple prompt words.
[0043] In this embodiment, the Prompt Extraction Model is used to automatically identify and extract semantically guiding keywords or phrases from natural language text. These keywords are typically used as prompts to guide large language models (LLMs), image generation models (such as Stable Diffusion and Midjourney), content recommendation systems, and other systems in task understanding and content generation.
[0044] In some examples, the core objective of the cue word extraction model is to extract keywords or phrases that have a key impact on subsequent AI behavior from text such as user input, conversation records, article content, and search history, thereby improving the model's understanding ability and output quality.
[0045] For example, prompts describe the generated content. Different styles of content will have different prompts. The core attributes of prompts may include, but are not limited to, the following: Subject: The theme or content of the image; this is the most crucial part of all prompts. Style: The artistic style of the image. Action / Scene: Primarily used to describe the scene in which the subject is located and what it is doing. Artist: Emphasizes a specific style, such as Van Gogh's style. Filters: Add richer, more personalized, and controllable details to the image.
[0046] It should be noted that the order in which the prompt words are placed in the first and second major models determines their priority, with the weight values decreasing from the first to the last. For example, in the prompt words "[Cartoon Style][Humorous][Ming Dynasty][Zhu Yuanzhang finally became emperor]", the comment "Zhu Yuanzhang finally became emperor" is the subject. The server generates other key content for the prompt words based on the book category, the context of the current comment in the book's chapter, and the user's interests.
[0047] Step 205: Input the extracted prompt words into the first and second large models respectively. Perform semantic analysis on the first large model based on the prompt words to obtain the reply content of the comment. And generate the screen in the second large model to obtain the reply screen corresponding to the reply content.
[0048] Optionally, before executing "inputting the extracted prompt words into the first and second large models respectively", the method of this embodiment further includes: determining the weights corresponding to the multiple prompt words based on priority information; accordingly, when executing "inputting the extracted prompt words into the first and second large models respectively", the following methods can be used, but are not limited thereto, including: inputting the multiple prompt words into the first large model, the first large model is used to perform semantic analysis training on the multiple prompt words to obtain semantic analysis results; inputting the multiple prompt words and the weights corresponding to the multiple prompt words into the second large model, the second large model is used to perform vectorization processing on the multiple prompt words based on the weights corresponding to the multiple prompt words, and perform screen generation training based on the vectorization processing results to obtain the response screen.
[0049] Optionally, when performing the process of "obtaining reply content from comment content through semantic analysis in the first model based on multiple prompt words, and generating a display screen corresponding to the reply content in the second model", the following methods can be used, but are not limited to these: in the first model, performing semantic extraction processing on multiple prompt words to obtain semantic information corresponding to each prompt word, and generating reply content based on the semantic information; in the second model, decomposing and vectorizing multiple prompt words based on their respective weights to obtain multiple prompt word vectors, and generating a reply screen based on the multiple prompt word vectors.
[0050] For example, the first large model can be a large language model. For instance, the obtained prompt words, before weighting, are input into the large language model LLM. The large language model performs semantic extraction on the comment content to obtain information such as "who, where, what did, and content" (i.e., semantic information in the embodiments of this application). This information is used as input to generate prompts and generate the answer content of the comment data.
[0051] Optionally, when performing the "Generate a response screen based on multiple prompt word vectors" function, the following methods may be used, but are not limited to these: determining the response subject based on multiple prompt word vectors, and determining the target response format corresponding to the response content from multiple response formats, wherein the multiple response formats include at least one or more of static image formats, dynamic image formats, and video formats; and generating a response screen containing the response subject based on the target response format.
[0052] The content output of multimodal scene capabilities can take many forms; here, we'll illustrate it using a combination of images and text. The large model generates content in different styles based on the book type, such as martial arts, fantasy, humor, animation, etc. It also uses contextual information to assist in generating image content; for example, on a sunny day, the generated image would show sunlight and leaves swaying in the wind.
[0053] As an optional approach, such as Figure 4 As shown, different text content generates different multimodal comments (i.e., the reply screen in this embodiment) based on user profile tags, book categories, etc. For example, some people like domineering CEOs, while others prefer the humor in the original work. Factors affecting the AI large model include... Figure 4 As shown, there can be many other factors, such as age, weather, and a film or television series of the same name. The content output of the comments is also dynamic, such as videos and GIFs.
[0054] For example, the second large model can be an AI large model. In this model, a text understanding component decomposes and quantizes the input prompts to capture the intended meaning of the text. If the generated response is determined to be an image, the vectorized data generated by the text understanding component can be used as input to an image information generator. After processing by a neural network and model, an information vector is output. Based on the image encoding of this output information vector, the final image is drawn. For example, a 4×64×64 dimension information array can be used as input (from the image generator), and the output image can be 3×512×512.
[0055] Step 206: Generate the user's target comment display content based on the comment content, the preset comment screen corresponding to the comment content, the reply content, and the reply screen.
[0056] Optionally, when performing the "generating the user's target comment display content based on the comment content, the preset comment screen corresponding to the comment content, the reply content, and the reply screen", the following methods can be used, but are not limited to these, including: embedding the comment content into a preset position in the preset comment screen to obtain the first comment display content; embedding the reply content into a target position in the reply screen to obtain the second comment display content, wherein the target position is determined based on the reply subject; merging the first comment display content and the second comment display content to obtain the target comment display content, which is used to display in the comment position corresponding to the target content.
[0057] In this embodiment of the application, the first comment display content can be the content after embedding the comment content into a preset comment screen, and correspondingly, the second comment display content can be the content after embedding the reply content into a reply screen.
[0058] As an optional approach, such as Figure 5 As shown, the target comment display content can be broken down. Specifically, the text data of the comment reply and the AI-generated image (i.e., the reply screen in this embodiment) can be merged. Facial recognition technology is used to determine whether there is a face in the image and to calculate the coordinates of the mouth (i.e., the target position in this embodiment). If there is no person in the AI-generated image, the author's portrait is stylized to generate a portrait that matches the current reader's style and is embedded into the AI image. The text content is placed near the mouth, and the final synthesized image (i.e., the second comment display content in this embodiment) is synchronized.
[0059] It should be noted that the preset comment screen may include user avatars, user images, etc., and the first comment displayed can be generated by combining the preset comment screen with the comment content.
[0060] It should be noted that existing technical solutions suffer from poor user interactivity: readers cannot interact with exciting scenes in the book; they can only express their feelings and insights by sending simple text or emojis. There is no vivid and engaging medium to promote active interaction among readers. The user experience is also poor: existing technical solutions present information in a generic way, failing to highlight unique user behaviors. Users desire distinctive and thought-provoking actions. This hinders content dissemination and artistic creation, fails to increase reader engagement, and is susceptible to replacement. Therefore, the embodiments of this application can significantly enhance engagement, increase user stickiness, provide feedback channels between readers and authors, help artists enhance their creative inspiration and motivation, promote book dissemination and reader interaction, improve user experience and stickiness, and provide users with a completely new way to interact. Besides expressing their current reading feelings in real time, users can also directly interact with the content in the book or with comments from other readers, making it highly engaging and increasing user stickiness.
[0061] Compared with existing technologies, this embodiment obtains the comment content input by the user based on the target content; extracts prompt words based on the target content and comment content; inputs the extracted prompt words into a first model and a second model respectively; performs semantic analysis on the prompt words in the first model to obtain the reply content of the comment content; and generates a screen in the second model to obtain the reply screen corresponding to the reply content; and generates the user's target comment display content based on the comment content, the preset comment screen corresponding to the comment content, the reply content, and the reply screen. This embodiment can generate corresponding reply content and reply screen based on the comment content input by the user. Combining the comment content and the preset comment screen allows for multiple modal comment content to be displayed after the user inputs a comment, and can combine different screens to display the comment content and reply content, enhancing the interest of the comment; by generating the comment display content based on the reply content and comment content together, the user can quickly get the corresponding reply after inputting the comment content, improving the interactivity, allowing the user to directly interact with the content in the book or with comments between readers, thereby improving the user experience.
[0062] Furthermore, as Figure 1 and Figure 2 To provide a specific implementation of the method shown, this embodiment offers a content generation device, such as... Figure 6 As shown, the device includes: an acquisition module 31, an extraction module 32, an input module 33, and a generation module 34.
[0063] Module 31 is configured to acquire comment content input by the user based on the target content; Extraction module 32 is configured to extract prompt words based on the target content and the comment content; The input module 33 is configured to input multiple extracted prompt words into the first large model and the second large model respectively, perform semantic analysis on the first large model based on the multiple prompt words to obtain the reply content of the comment content, and perform screen generation on the second large model to obtain the reply screen corresponding to the reply content; The generation module 34 is configured to generate the user's target comment display content based on the comment content, the preset comment screen corresponding to the comment content, the reply content, and the reply screen.
[0064] In some examples of this embodiment, the extraction module 32 is further configured to obtain user profile information containing user reading characteristics; correspondingly, the extraction module 32 is specifically configured to determine the context content and content category corresponding to the target content; input the user profile information, the context content, the content category and the comment content into the prompt word extraction model, and extract from the reading feature dimension and the reading content dimension in the prompt word extraction model to obtain multiple prompt words corresponding to multiple attributes and priority information corresponding to the multiple prompt words respectively.
[0065] In some examples of this embodiment, the input module 33 is further configured to determine the weights corresponding to the plurality of prompt words based on the priority information; correspondingly, the input module 33 is specifically configured to input the plurality of prompt words into the first large model, which is used to perform semantic analysis training on the plurality of prompt words to obtain semantic analysis results; input the plurality of prompt words and the weights corresponding to the plurality of prompt words into the second large model, which is used to perform vectorization processing on the plurality of prompt words based on the weights corresponding to the plurality of prompt words, and perform screen generation training based on the vectorization processing results to obtain the response screen.
[0066] In some examples of this embodiment, the input module 33 is further configured to perform semantic extraction processing on the plurality of prompt words in the first large model to obtain semantic information corresponding to the plurality of prompt words respectively, and generate the response content based on the semantic information; in the second large model, the plurality of prompt words are decomposed and vectorized based on the weights corresponding to the plurality of prompt words to obtain a plurality of prompt word vectors, and the response screen is generated based on the plurality of prompt word vectors.
[0067] In some examples of this embodiment, the input module 33 is further configured to determine the response subject based on the plurality of prompt word vectors, and to determine the target response form corresponding to the response content from a plurality of response forms, wherein the plurality of response forms include at least one or more of static image form, dynamic image form, and video form; and to generate a response screen containing the response subject based on the target response form.
[0068] In some examples of this embodiment, the generation module 34 is further configured to embed the comment content into a preset position in the preset comment screen to obtain first comment display content; embed the reply content into a target position in the reply screen to obtain second comment display content, wherein the target position is determined based on the reply subject; and merge the first comment display content and the second comment display content to obtain the target comment display content, which is used to display at the comment position corresponding to the target content.
[0069] It should be noted that other corresponding descriptions of the functional units involved in the content generation apparatus provided in this embodiment can be found in [reference needed]. Figure 1 and Figure 2 The corresponding descriptions in [the document] will not be repeated here.
[0070] Based on the above, Figure 1 and Figure 2 Accordingly, this embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. Figure 1 and Figure 2 The method shown.
[0071] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0072] like Figure 7 The diagram shown is a hardware structure schematic of an electronic device according to the present invention, comprising: At least one processor 401; and, A memory 402 is communicatively connected to at least one of the processors 401; wherein, The memory 402 stores instructions that can be executed by at least one of the processors to enable at least one of the processors to perform the content generation method as described above.
[0073] Figure 7 Take a processor 401 as an example.
[0074] The electronic device may also include an input device 403 and a display device 404.
[0075] The processor 401, memory 402, input device 403, and display device 404 can be connected via a bus or other means. Figure 7 Taking the example of a connection between China and Israel via a bus.
[0076] Memory 402, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the content generation method in the embodiments of this application, for example, Figure 1 and Figure 2 The method flow is shown. The processor 401 executes various functional applications and data processing by running non-volatile software programs, instructions, and modules stored in the memory 402, thereby implementing the content generation method in the above embodiments.
[0077] Memory 402 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created according to the use of the content generation method, etc. Furthermore, memory 402 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 402 may optionally include memory remotely located relative to processor 401, and these remote memories may be connected to the apparatus performing the content generation method via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0078] Input device 403 can receive user clicks and generate signal inputs related to user settings and function control for content generation methods. Display device 404 may include display devices such as a display screen.
[0079] When one or more modules are stored in the memory 402, and are run by one or more processors 401, the content generation method in any of the above method embodiments is executed.
[0080] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0081] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0082] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0083] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms, or it can be implemented by hardware. By applying the solution of this embodiment, compared with the existing technology, this embodiment obtains the comment content input by the user based on the target content; extracts prompt words based on the target content and the comment content; inputs the extracted prompt words into the first large model and the second large model respectively; performs semantic analysis on the prompt words in the first large model to obtain the reply content of the comment content; and generates a screen in the second large model to obtain the reply screen corresponding to the reply content; generates the user's target comment display content based on the comment content, the preset comment screen corresponding to the comment content, the reply content, and the reply screen; this embodiment can generate corresponding reply content and reply screen based on the comment content input by the user. Combining the comment content and the preset comment screen can enable multiple modal comment content to be displayed after the user inputs a comment, and can combine different screens to display the comment content and reply content, thereby improving the fun of the comment; by generating the comment display content based on the reply content and the comment content, the user can quickly get the corresponding reply after inputting the comment content, improving the interactivity, allowing the user to directly interact with the content in the book or with comments between readers, thereby improving the user experience.
[0084] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0085] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A content generation method characterized by, The method comprises the following steps: obtaining the comment content input by the user based on the target content; extracting prompt words based on the target content and the comment content; inputting the extracted prompt words into a first large model and a second large model respectively, performing semantic analysis on the prompt words in the first large model based on the prompt words to obtain the reply content of the comment content, and performing picture generation in the second large model to obtain the reply picture corresponding to the reply content; generating the target comment display content of the user based on the comment content, the preset comment picture corresponding to the comment content, the reply content and the reply picture.
2. The method of claim 1, wherein, Before the step of extracting prompt words based on the target content and the comment content, the method further comprises the following steps: obtaining user portrait information containing user reading characteristics; The step of extracting prompt words based on the target content and the comment content comprises the following steps: determining the context content and the content category corresponding to the target content; inputting the user portrait information, the context content, the content category and the comment content into a prompt word extraction model, and extracting from the reading characteristic dimension and the reading content dimension in the prompt word extraction model to obtain a plurality of prompt words corresponding to a plurality of attributes and priority information corresponding to the plurality of prompt words.
3. The method of claim 2, wherein, Before the step of inputting the extracted prompt words into the first large model and the second large model, the method further comprises the following steps: determining the weights corresponding to the plurality of prompt words based on the priority information; The step of inputting the extracted prompt words into the first large model and the second large model comprises the following steps: inputting the plurality of prompt words into the first large model, and the first large model is used for semantic analysis training of the plurality of prompt words to obtain a semantic analysis result; inputting the plurality of prompt words and the weights corresponding to the plurality of prompt words into the second large model, and the second large model is used for vectorization processing of the plurality of prompt words based on the weights corresponding to the plurality of prompt words, and picture generation training based on the vectorization processing result to obtain the reply picture.
4. The method of claim 3, wherein, The step of performing semantic analysis on the plurality of prompt words in the first large model to obtain the reply content of the comment content, and performing display picture generation in the second large model to obtain the reply picture corresponding to the reply content comprises the following steps: in the first large model, performing semantic extraction processing on the plurality of prompt words to obtain semantic information corresponding to the plurality of prompt words, and generating the reply content based on the semantic information; in the second large model, performing disassembly and vectorization processing on the plurality of prompt words based on the weights corresponding to the plurality of prompt words to obtain a plurality of prompt word vectors, and generating the reply picture based on the plurality of prompt word vectors.
5. The method of claim 4, wherein, The step of generating the reply picture based on the plurality of prompt word vectors comprises the following steps: determining a reply subject based on the plurality of prompt word vectors, and determining a target reply form corresponding to the reply content from a plurality of reply forms, and the plurality of reply forms at least include one or more of a static image form, a dynamic image form and a video form; Generate a reply screen containing the reply body based on the target reply form.
6. The method of claim 5, wherein, The generating the target comment display content of the user based on the comment content, the preset comment screen corresponding to the comment content, the reply content and the reply screen comprises: Embedding the comment content in a preset position in the preset comment screen to obtain a first comment display content; Embedding the reply content in a target position in the reply screen to obtain a second comment display content, wherein the target position is determined based on the reply body; Merging the first comment display content and the second comment display content to obtain the target comment display content, and the target comment display content is used for display at a comment position corresponding to the target content.
7. A content generation apparatus characterized by comprising: Comprise: The acquisition module is configured to acquire comment content input by a user based on target content; The extraction module is configured to extract prompt words based on the target content and the comment content; The input module is configured to input the extracted prompt words into a first large model and a second large model respectively, perform semantic analysis on the prompt words in the first large model based on the prompt words to obtain reply content of the comment content, and perform screen generation in the second large model to obtain a reply screen corresponding to the reply content; The generation module is configured to generate target comment display content of the user based on the comment content, the preset comment screen corresponding to the comment content, the reply content and the reply screen.
8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1-6.
9. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, The processor executes the computer program to implement the method of any one of claims 1-6.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1-6.