User experience-based content generation platform server and platform providing method

The user experience-based content generation platform server uses multiple learning models and user feedback to generate content that accurately reflects personal experiences and preferences, addressing the limitations of existing AI-based methods by iteratively refining the creative output to align with user intentions.

JP2025533628APending Publication Date: 2025-10-07LG MANAGEMENT DEV INST CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2025518769
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-26
Publication Date
2025-10-07

AI Technical Summary

Technical Problem

Existing AI-based content generation methods struggle with generating creative content that accurately reflects a user's personal experiences and preferences due to unsophisticated language input or mismatched information, leading to outputs that do not align with the user's intended creative vision.

Method used

A user experience-based content generation platform server utilizing multiple learning models to process user input, including a first learning model for text reconstruction, a second learning model for image captioning, and a third learning model for morphological analysis, along with a user interface for feedback and word bag generation, to create content that aligns with the user's intentions.

Benefits of technology

The platform enables the creation of content that reflects the user's own concept and experiences, maximizing creativity by iteratively refining the content generation process through user feedback and learning, ensuring alignment with the user's intended creative output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025533628000001_ABST
    Figure 2025533628000001_ABST
Patent Text Reader

Abstract

The present invention aims to provide a process for generating images that maximize creativity in the direction desired by the user by repeating the process of forming an archive that reflects the user's own concept based on the user's experience and thoughts, using a learning model that generates images based on text.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a user experience-based content generation platform server and a platform providing method. [Background technology]

[0002] In today's environment where a wide variety of content is provided, creative content generated by users is not something that is created out of nothing, but rather something that is generated by incorporating existing content and reflecting the user's personal inclinations, such as their experiences, thoughts, and preferences.

[0003] Specifically, it is a method of generating creative content by linking existing content in various ways, and for this purpose, artificial intelligence-based content generation methods have been proposed.

[0004] However, learning for AI-based content generation can be difficult because unsophisticated language is input during learning, or information that differs from the user's needs is input, making it difficult to output the creative content intended by the user. Summary of the Invention [Problem to be solved by the invention]

[0005] The embodiments disclosed in the present disclosure aim to provide a user experience-based content generation platform server and platform provision method for forming a user's own concept based on the user's experience and ideas and providing content that maximizes creativity.

[0006] The problems to be solved by the present disclosure are not limited to those mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the following description. [Means for solving the problem]

[0007] A content generation platform server according to an embodiment of the present invention includes: a memory including a first learning model trained to generate reconstructed content based on text; and a processor communicating with the memory and controlling the first learning model to output at least one reconstructed content corresponding to text when text is input from a user terminal; wherein the processor receives input content and a first text matching the input content from the user terminal, generates a word bag based on the first text, determines caption data using a second text determined based on the first text included in the word bag and a predetermined sentence structure, inputs a sentence represented by the caption data into the first learning model, generates the at least one reconstructed content corresponding to the sentence, and outputs the at least one reconstructed content and the input content concatenated to the user terminal.

[0008] In this case, the memory includes a second learning model that outputs at least one text representing an input image, and when the input content is an input image, the processor, when the first text matching the input content is input, generates a recommendation text including sentences or words including objects, appearances, and backgrounds by captioning the input image using the second learning model, and provides the recommendation text to the user terminal, and at least one final text determined or modified by the user based on the recommendation text can be input as the first text.

[0009] Also, when the input content is an input image, the processor may provide the user terminal with a process for generating the caption data, the process including the predetermined sentence structure constituting the caption data and word categories for each item constituting the sentence structure, wherein the word categories for each item are composed of at least one or more blanks filled in by user settings, and linking words may be formed between the blanks of the word categories for each item.

[0010] In addition, the processor may provide the user terminal with recommended words that can be entered into each of the plurality of blank spaces, and may provide the recommended words taking into consideration their relevance to the input image and whether or not they match the word category.

[0011] Furthermore, when the reconstructed content is a reconstructed image, the processor may receive feedback from the user terminal regarding a reconstructed image selected by the user from the at least one reconstructed image, additionally generate the second text of the word bag based on the feedback, and reconstruct a sentence of the caption data based on the feedback.

[0012] The processor may also generate additional second text for at least one word category of the word bag based on the feedback.

[0013] In addition, the processor may provide a user interface to the user terminal so that when feedback is input from the user terminal, a feedback opinion on the reconstructed image and a pinpoint within the reconstructed image that matches the feedback opinion can be input by the user's operation.

[0014] The memory also includes a third learning model which is a morphological analyzer trained to preprocess text and separate it into morphemes, and after receiving the feedback, the processor inputs the reconstructed image to the third learning model to determine common morpheme tokens associated with the reconstructed image, reconstructs sentences of the caption data using the determined common morpheme tokens, and provides the reconstructed sentences to the user terminal.

[0015] In addition, the user terminal modifies the reconstructed sentence through the user's operation and transmits the final reconstructed sentence to the content generation platform server, and when the processor receives the final reconstructed sentence, it inputs it into the first learning model, re-outputs the reconstructed image, and provides it to the user terminal.

[0016] The processor may also generate an archive of each of the initial generation of the reconstructed image and the reconstructed image generated by the repeated execution of the feedback.

[0017] The processor may also search for real images having a similarity greater than or equal to a predetermined reference value based on the at least one reconstructed image, and provide the real images to the user terminal.

[0018] The processor may also extract caption data including sentences or words from captioning each of the at least one reconstructed image, and compare the extracted caption data with the actual image to recommend new caption data.

[0019] A method for providing a content generation platform according to an embodiment of the present invention is a method for generating user experience-based content by a processor of a user terminal in cooperation with a content generation platform server including at least one learning model, and may include the steps of: determining an input image and first text for the input image in response to a user's input; image captioning the input image to provide recommendation text including sentences or words for at least one category of objects, appearances, or backgrounds; additionally determining the first text for the input image in response to a user's input for the recommendation text; generating a word bag based on the determined first text; providing the user with a process for setting caption data based on the word bag; determining second text for the caption data in response to the user's input, and setting the caption data in accordance with the determined second text and a predetermined sentence structure; inputting a sentence represented by the set caption data, generating a reconstructed image using the at least one learning model, and outputting the generated reconstructed image.

[0020] In this case, the step of providing the user with a process for setting caption data based on the word bag may include a step of outputting a process for generating the caption data including the predetermined sentence structure constituting the caption data and word categories for each item constituting the sentence structure.

[0021] In addition, each of the itemized word categories in the process for generating the caption data may be composed of at least one or more blank spaces that are filled in by user settings, and linking words may be formed between the blank spaces in each of the itemized word categories.

[0022] The method may also include a step of receiving feedback from the user regarding a reconstructed image selected by the user from the at least one reconstructed image output, and a step of reconstructing the text of the caption data based on the feedback.

[0023] In addition, the step of inputting feedback on the reconstructed image may include a step of inputting a feedback opinion on the reconstructed image and a pinpoint within the reconstructed image that matches the feedback opinion in response to an operation of the user.

[0024] Furthermore, the step of reconstructing the sentence of the caption data based on the feedback may include inputting the reconstructed image into the third learning model to determine common morpheme tokens associated with the reconstructed image, reconstructing the sentence of the caption data using the determined common morpheme tokens, and providing the reconstructed sentence to the user terminal.

[0025] Additionally, reconstructing the sentences of the caption data based on the feedback may include generating and providing additional second text for at least one word category of a word bag based on the feedback. [Effects of the Invention]

[0026] According to the above-mentioned problem-solving means of the present disclosure, an archive that reflects the user's own concept based on the user's experience and thoughts can be created, and this can be used to provide images that maximize creativity.

[0027] According to the above-described problem-solving means of the present disclosure, the learning model is trained using text in a bag of words and a previously determined sentence structure, so that images that match the user's intentions can be generated.

[0028] The effects of the present disclosure are not limited to those mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the following description. [Brief explanation of the drawings]

[0029] [Figure 1] FIG. 2 is a block diagram of a content creation platform server of the present disclosure. [Figure 2] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 3] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 4] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 5] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 6] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 7] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 8] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 9] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 10] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 11] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 12] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 13] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 14] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 15] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 16]FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 17] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 18] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 19] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 20] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 21] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 22] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 23] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 24] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 25] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 26] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 27] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 28] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 29] FIG. 1 is an exemplary diagram illustrating a content generation method according to the present disclosure. [Figure 30] 1 is a flowchart illustrating a platform providing method according to the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0030] The same reference numerals refer to the same elements throughout this disclosure. This disclosure does not describe all elements of the embodiments, and content that is common in the technical field to which the disclosure belongs or redundant content in the embodiments will be omitted. The terms "unit, module, component, block" used in the specification can be realized as software or hardware, and depending on the embodiment, multiple "units, modules, components, blocks" may be realized as one component, or one "unit, module, component, block" may include multiple components.

[0031] Throughout this specification, when a part is said to be "coupled" to another part, this includes not only direct coupling but also indirect coupling, and indirect coupling includes coupling via a wireless communication network.

[0032] Furthermore, when a part is described as "comprising" certain elements, this does not mean that other elements are excluded, but that other elements may also be included, unless otherwise specified.

[0033] Throughout this specification, when an element is said to be "on" another element, this includes not only when the element is in contact with the other element, but also when there is another element between the two elements.

[0034] The terms "first," "second," etc. are used to distinguish one component from another, and do not limit the components to the terms mentioned above.

[0035] The singular expression includes the plural expression unless the context clearly indicates otherwise.

[0036] The identification numbers in each step are used for convenience of explanation, and do not describe the order of each step. The steps may be performed in an order different from the order specified unless the context clearly dictates a specific order.

[0037] Hereinafter, the working principle and embodiments of the present disclosure will be described with reference to the accompanying drawings.

[0038] In this specification, the term "content generation platform server according to the present disclosure" includes various devices capable of performing computations and providing results to users. For example, the content generation platform server according to the present disclosure may include all or any one of a computer, a server device, and a portable terminal.

[0039] Here, the computer may include, for example, a notebook PC, a desktop PC, a laptop, a tablet PC, a slate PC, or the like equipped with a web browser.

[0040] The server device is a server that communicates with external devices and processes information, and may include an application server, a computing server, a database server, a file server, a game server, a mail server, a proxy server, a web server, and the like.

[0041] The portable terminal is, for example, a wireless communication device that ensures portability and mobility, and may include any kind of handheld-based wireless communication device such as PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), WiBro (Wireless Broadband Internet) terminal, smartphone, etc., as well as wearable devices such as watches, rings, bracelets, necklaces, glasses, contact lenses, or head-mounted devices (HMDs).

[0042] The user experience-based content generation platform may include a content generation platform server and a terminal (not shown) that generates user experience-based content in response to user input.

[0043] In an embodiment, the content generating terminal runs a content generation program, communicates with the content generation platform server, and transmits data input by a user. The content generation platform server analyzes the transmitted data, generates content based on a learning model, and then transmits the content to the terminal, thereby generating user experience-based content.

[0044] In another embodiment, the content generating terminal performs user input and some data processing, transmits analysis using a learning model to a content generation platform server, and receives data output from the learning model and provides it to the user, thereby generating user experience-based content.

[0045] In another embodiment, the content generating terminal receives a content generation program including a learning model from a content generation platform server, and then executes the received content generation program to analyze data directly input by a user, thereby performing a user experience-based content generation method.

[0046] In the following description, a method will be described in which a content generating terminal receives user input, receives data analyzed in a learning model from a content generation platform server, and generates user experience-based content. However, in other embodiments, a process in which some of the processes of the content generation method described below as being performed by the content generation platform server are executed in the content generation terminal is also included in the embodiments of the present invention. Here, the terminal may include a computer, a portable terminal, etc. as described above.

[0047] FIG. 1 is a block diagram of a content creation platform server of the present disclosure.

[0048] The following description will be made with reference to FIGS. 2 to 29, which are illustrative diagrams for explaining the content generation method of the present disclosure.

[0049] As shown in FIG. 1 , the content generation platform server 100 includes a processor 110, a memory 140, and a communication processor 150. Here, the processor 110 executes a content generation program that generates content based on text to perform a content generation method. Such content generation program includes a content generation module 120 and a feedback processing module 130, which are executed by the processor 110. The content generation module 120 and the feedback processing module 130 can control the operation of each corresponding component. The components shown in FIG. 1 are not essential to realizing the content generation platform server 100 according to the present disclosure, and therefore the content generation platform server 100 described herein may have more or fewer components than those listed above.

[0050] The processor 110 is in communication with the memory 140, and the content generation program may include a first learning model that is trained to output at least one reconstructed content corresponding to text when text based on user input is input from a terminal.

[0051] On the other hand, users may find it difficult to create text that will generate the reconstructed content they desire using this first learning model, and when the reconstructed content is not generated in the direction they desire, it may be difficult to find a solution.

[0052] The present invention aims to provide a user experience-based content generation method by archiving a content generation process or a process in which a user directly inputs content and text to be matched thereto, and assisting the user in composing text based on this.

[0053] To this end, the processor 110 may receive input content from a user and the first text to be matched to the input content.

[0054] Here, the input content and the first text can be selected by the user through direct archive input, and the first text to be matched with the selected input content can be directly input (typed).

[0055] In addition, the first text may be a user-input text input into the first learning model to determine an initial concept of the reconstructed content to be generated, and the reconstructed content generated based on the first text may be the input content. That is, the user may input the first text into the first learning model for the initial concept and set the generated reconstructed content as the input content.

[0056] In the following, an example will be described in which the input content is an input image and the reconstructed content is a reconstructed image.

[0057] The first text may include personal reasons and opinions related to the input image, but is not limited to these, and may be any information that reflects the user's experience. For example, if the input image is a fox image, the processor 110 may register the "silly fox face" and "strong touching texture" related to the fox image input by the user as the first text. In this way, the first text may not only describe the subject of the input image, but may also include various opinions, such as the drawing style of the input image.

[0058] The processor 110 may receive input content and first text through a user's operation to generate an archive for setting an initial concept that reflects the user's experience. The archive may match and store various content (e.g., images) and associated first text based on the user's experience. The processor 110 may then use the archive to generate reconstructed content that matches the user's intention based on the archive including at least one input content and the matched first text.

[0059] 2, 6 to 9, the processor 110 provides items for registering an input image, and can register a fox image as shown in FIG. 9 according to an item selected by a user's operation (Experience Archive → Image Archive → Upload Input Image). At this time, the processor 110 can receive first text for the input image input by a user's operation. For example, the processor 110 can obtain a fox image and associated first text such as "Silly fox face" and "Strong touching texture" by a user's operation.

[0060] 10, the processor 110 may match and display the input image and the first text, where the first text may be one or more.

[0061] Meanwhile, the present disclosure can provide the following recommendation service in relation to the input of the first text through a learning model.

[0062] Figure 3 is a diagram illustrating in more detail the Documentation (Media to Text) process of Figure 2. As shown in Figure 3, when a first text matching the input content is input, the processor 110 can generate and output a recommendation text including sentences or words including objects, appearances, and backgrounds by captioning the input image using the second learning model. At this time, the processor 110 can obtain the recommendation text by searching a word bag after captioning the input image.

[0063] For example, the second learning model may include a first image captioning learning model that, when inputted with an input image, generates text describing each object and situation in the image.

[0064] The second learning model may also include a second image captioning learning model that, when an image is input, generates text describing the type, characteristics, or other word categories of attributes of the image.

[0065] 3, the processor 110 can provide at least one or more recommended texts to the user when the input image is input to the image captioning learning model. The processor 110 can also provide a user interface for re-editing the first text for the input image on a screen displaying the text input by the user for initial concept setting and at least one or more recommended texts (or recommended words).

[0066] More specifically, the processor 110 may filter the text input by the user for initial concept setting and the recommended text into a base text for generating the first text using a BOW filter. The processor 110 may then input the filtered base text into a sentence construction learning model to generate a suggested first text for the input image and provide it to the user.

[0067] The user can modify and edit the provided primary text to generate a final primary text for the input image.

[0068] That is, the processor 110 may receive the final text determined or modified based on the user's input test and the recommended text as the first text. That is, the recommended text may be selected as is or modified according to the user's operation and reflected in the first text.

[0069] 11 to 13, various types of input images that can serve as references for a concept that a user wishes to express and first text that matches the input images are repeatedly input to the processor 110. In addition, the processor 110 may link the input images, first text, other input images, and the first texts of the other input images according to the user's input, and store them in an archive. The state in which a plurality of input images and first texts are linked in this manner may be represented by a tree structure as shown in FIG. 13. When the processor 110 stores the archive including the input images and first texts in the memory 140, it may store them reflecting such linkages.

[0070] The processor 110 can then generate a bag of words based on the first text. The aforementioned iterative process results in a plurality of first texts, and the processor 110 can generate a word bag based on the first texts. At this time, the word bag is formed based on the first text input by the user, reflecting the characteristics of each individual user. Furthermore, the processor 110 can recommend first texts for generating a word bag by learning the input image.

[0071] The processor 110 may determine caption data using the second text included in the word bag and a predetermined sentence structure. In this case, the caption data may refer to a sentence expressing the content of the image. The second text may refer to text included in the word bag formed based on a plurality of first texts.

[0072] The processor 110 may output, via an output unit (not shown), a predetermined sentence structure including a plurality of blanks to which word categories for describing the input image are matched and linking words between the plurality of blanks. The word categories indicate that words to be input into the blanks are classified by attribute.

[0073] 14 to 20, when processor 110 executes the reconstructed image generation process, it generates caption data (e.g., sentences) for generating a desired reconstructed image, and generates and provides a reconstructed image using a first learning model based on the generated caption data. To assist in generating the caption data, processor 110 provides items for each word category for generating basic sentences for generating a reconstructed image and at least one or more first texts or / and words included in the first texts that are matched with each item, and determines second texts according to the first texts (or words belonging to the first texts) of each item selected by a user operation (AI Image Generation → Sentence Builder → Word Selection), thereby determining caption data.

[0074] At this time, the processor 110 may provide recommended words that can be entered into each of the blanks corresponding to each of the plurality of word category items based on the first text. To this end, the processor 110 may provide the first text and / or recommended words included in the first text, taking into consideration the relevance of the first text to the input image and whether or not the first text matches each word category.

[0075] For example, the processor 110 may provide a predetermined sentence structure such as "A 'TYPE' of 'BASE' that is 'DETAIL' in the 'STYLE'." In this case, "TYPE," "BASE," "DETAIL," and "STYLE" may be implemented as blank fields for inputting words. Here, a word may refer to a phrase or clause consisting of multiple words, and may be a word included in a first text input by the user for the input image or / and a first text recommended for the input image. Each blank field may display first text that matches a word category required to describe the image, such as type, base, detail, and style.

[0076] In addition, the processor 110 may additionally provide reference words representing the attributes of each blank item. In particular, the processor 110 may recommend reference words representing the attributes of each word category item that is not entered in the archive by the user but is typically used for each type, base, detail, and style item. When recommending reference words, the processor 110 may provide images that match the reference words so that the user can intuitively understand the meaning of the reference words.

[0077] The caption data may include the aforementioned A, of, that is, and in the connectives between multiple blanks, and may include articles, particles, etc., for completing a sentence when a word is input into the blank. In other words, the connectives of the present disclosure refer to words that complete a sentence when a word is input into a blank in a predetermined sentence structure.

[0078] Specifically, the processor 110 of the present disclosure can prevent a user from inputting a sentence that does not meet the standard in terms of format and content, or from inputting a sentence that does not accurately reflect the user's intention because the user is unsure of what sentence to input, thereby allowing the user to input a high-quality sentence that is useful for learning a learning model. That is, the processor 110 guides the user on what words to include in combining sentences.

[0079] As shown in Figures 16 to 20, the processor 110 can display a previously determined sentence structure, such as A TYPE of BASE that is DETAIL in the STYLE, through the output unit, and recommend and display at least one word that matches the word category of each blank space so that the user can input it.

[0080] For example, if the input image is a fox and a predetermined sentence structure such as "A "TYPE" of "BASE" that is "DETAIL" in the "STYLE" is output on the screen, processor 110 can provide a plurality of recommended words related to fox that can be entered in the blanks for "TYPE," "BASE," "DETAIL," and "STYLE" by searching a word bag. Processor 110 enters the words "oil-color painting," "a fox," "sitting in a field at sunrise," and "realism art" selected by the user in correspondence with "TYPE," "BASE," "DETAIL," and "STYLE" from the plurality of recommended words, so that they correspond to each blank. Processor 110 can determine caption data such as "A oil-color painting of a fox that is sitting in a field at sunrise in the style of realism art." based on the entered words.

[0081] That is, the processor 110 may provide words belonging to the first text corresponding to the TYPE item from a previously recorded archive or recommended words for the input image so that the user can select a word corresponding to the TYPE item. Also, the processor 110 may provide reference words typically used for TYPE and related reference images for understanding the expression meaning of the reference words so that the user can select a word corresponding to the TYPE item.

[0082] 19 and 20, the processor 110 may provide words belonging to the first text corresponding to the STYLE item from a previously recorded archive or recommended words for the input image so that the user can select a word corresponding to the STYLE item. The processor 110 may also provide reference words typically used for STYLE and related reference images for understanding the expression meaning of the reference words so that the user can select a word corresponding to the STYLE item. The time required to generate a reconstructed image varies depending on each word specifying STYLE. Therefore, when determining a word corresponding to STYLE, the processor 110 may provide a forecast image generation time for the corresponding word in advance depending on the number of images to be generated.

[0083] As shown in Figures 2 and 21, the categories of the word bag may include type, figuration, base, description, action, and style.

[0084] Here, type, shape, base, description, action and style can be derived from the codes for object, expression, atmosphere, style, background and others.

[0085] The type can be the final visible result (e.g., pattern, illustration, manipulation, photograph), the form can be simplified, shaped, abstract, etc.

[0086] Also, the base may be the core object and the description may be a description of the object.

[0087] The action may also be a behavior that the subject takes.

[0088] The first codes of type, shape, base, description, action and style mentioned above can be subdivided and applied to the second codes to obtain prototype codes of type, base, detail and style.

[0089] Such a prototype code may be a word category that guides input into a blank space of a predetermined sentence structure. The word categories and word bag categories are not limited to the above description and may be changed according to the needs of the operator.

[0090] Meanwhile, the processor 110 of the present disclosure can process the determined caption data as follows to improve the sentence completeness.

[0091] As an example, the processor 110 of the present disclosure includes a grammar check function, and can therefore perform a grammar check on the determined caption data. The processor 110 checks a sentence for potential grammatical errors using only the input words searched from the word bag and the determined connectives, and can correct the connectives or the form of the words input into the blanks so that the combination of words is correct. For example, the processor 110 can correct the part of speech of the words input into the blanks to match the sentence structure, or correct (e.g., add, delete, or change) the connectives.

[0092] As another example, when presenting recommended words searched from a word bag, the processor 110 of the present disclosure may modify the part of speech of the recommended word in consideration of the sentence structure between the blank into which the corresponding recommended word is input and the pre-set connectives, and present the recommended word. For example, the processor 110 may modify the part of speech of A according to the connectives adjacent to the blank into which the word A is input, or may add connectives such as articles and particles to A and present the recommended word.

[0093] The processor 110 provides a tool that allows a user to freely input sentences based on the word bag and generate caption data. For example, as shown in Figures 22 and 23, the processor 110 can select a word in the word bag by selecting the Direct Generator item. The grammar check function described above can also be applied in this process.

[0094] The processor 110 may input caption data into a learning model to generate at least one reconstructed content corresponding to the first text. The reconstructed content may be a reconstructed image. In this case, the processor 110 may generate the reconstructed content corresponding to the first text using a multimodal AI that simultaneously inputs and learns various modalities.

[0095] As shown in Figures 24 to 26, processor 110 can generate 16 reconstructed images based on the caption data "A painting of a fox sitting in a field at sunrise in the style of realism art." generated by the caption data generation process for generating reconstructed images.

[0096] The processor 110 may receive feedback on a reconstructed image selected by a user from at least one reconstructed image, as shown in Fig. 27. The processor 110 may then store the feedback on the reconstructed image in an archive as input image-first text, and use it to help generate a reconstructed image.

[0097] The above-mentioned feedback process can be performed after generating the at least one reconstructed content and before linking and outputting at least one reconstructed content and the input content, as described below, but is not limited to this, and can also be performed after linking and outputting the reconstructed content and the input content.

[0098] The processor 110 may additionally generate a second text of a word bag based on the archive containing the first text with the reconstructed image as the input image based on the feedback. That is, the processor 110 may receive feedback on the reconstructed image, input text and pinpoints as feedback for the reconstructed image, generate the reconstructed image as input image-first text, and store the reconstructed image in the archive. The processor 110 may then provide a process for generating caption data based on the additionally stored input image-first text, and repeat the process of generating a reconstructed image that gradually more closely matches the user's intentions.

[0099] To this end, the processor 110 may additionally generate second text according to categories of the word bag based on the feedback.

[0100] As shown in FIG. 28, when feedback is input, the processor 110 can input a feedback opinion (O) for the reconstructed image and a pinpoint (P) within the reconstructed image that matches the feedback opinion in accordance with the user's operation.

[0101] 4, when a feedback is input, the processor 110 inputs the reconstructed image to a learning model, identifies common morpheme tokens related to the reconstructed image, reconstructs the sentence of the caption data using the identified common morpheme tokens, and outputs the reconstructed sentence through an output unit. At this time, the processor 110 can modify the reconstructed sentence according to a user operation.

[0102] 2, the processor 110 can generate an archive (1st Image Archiving...Nth...Nth Image Archiving) of each reconstructed image generated by the initial generation of the reconstructed image and the repeated execution of the feedback. The multiple archives generated in this way can be stored in the memory 140 and used to generate images that better match the user's intentions.

[0103] As shown in FIG. 5, the processor 110 can search for and provide real images whose similarity is equal to or exceeds a predetermined reference value based on at least one reconstructed image.

[0104] Additionally, as shown in FIG. 5, the processor 110 can extract caption data including sentences or words from the captioning of each of at least one reconstructed image, and compare the extracted caption data with the actual image to recommend new caption data.

[0105] As shown in FIG. 29, the processor 110 can combine and output at least one reconstructed content (A-1) and the input content (A).

[0106] The processor 110 of the present disclosure may be configured with one or more cores and may include a processor for data analysis and deep learning, such as a central processing unit (CPU), a general-purpose graphics processing unit, or a tensor processing unit of a computing device. The processor 110 may read a computer program stored in the memory 140 to perform data processing for machine learning according to the present disclosure. According to the present disclosure, the processor 110 may perform calculations for neural network training. The processor 110 may perform calculations for neural network training, such as processing input data for training in deep learning, extracting features from the input data, calculating errors, and updating neural network weights using backpropagation. Although not shown, the processor 110 of the present disclosure may input noisy training data including clean-label data and label-noise data to a neural network model to select label noise, and mix the label-noise data with the clean-label data to train a classifier.

[0107] The neural network model may be a deep neural network. In this disclosure, the terms neural network, network function, and neural network may be used interchangeably. A deep neural network (DNN) may refer to a neural network that includes multiple hidden layers in addition to an input layer and an output layer. A deep neural network can be used to understand the latent structures of data. That is, it can understand the latent structures of photos, text, videos, audio, and music (e.g., what objects are in the photo, what is the content and emotion of the text, what is the content and emotion of the audio, etc.). Deep neural networks may include convolutional neural networks (CNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), Q-networks, U-networks, Siamese networks, etc.

[0108] A convolutional neural network is a type of deep neural network that includes a neural network with convolutional layers. A convolutional neural network is a type of multilayer perceptron designed to use minimal preprocessing. A CNN may consist of one or more convolutional layers and an artificial neural network layer connected to them. A CNN can additionally utilize weight and pooling layers. This structure allows CNNs to fully utilize two-dimensional input data. A convolutional neural network can be used to recognize objects from images. A convolutional neural network can process image data by representing it as a matrix with dimensions. For example, image data encoded in RGB (red-green-blue) can be represented as a two-dimensional matrix for each of the R, G, and B colors (e.g., for a two-dimensional image). In other words, the color value of each pixel in the image data can be an element of the matrix, and the size of the matrix can be the same as the size of the image. Therefore, image data can be represented as three two-dimensional matrices (a three-dimensional data array).

[0109] In a convolutional neural network, the convolution process (input / output of a convolutional layer) can be performed by moving the convolutional filter and multiplying the matrix elements at each position of the image. The convolutional filter may be configured as an n*n matrix. The convolutional filter may generally be configured as a fixed-type filter smaller than the total number of pixels in the image. That is, when an m*m image is input to a convolutional layer (e.g., a convolutional layer with an n*n convolutional filter), a matrix representing n*n pixels, including each pixel of the image, can be the Hadamard product (i.e., the product of each element of the matrix) of the convolutional filter. Components that match the convolutional filter can be extracted from the image by multiplying it with the convolutional filter. For example, a 3*3 convolutional filter for extracting vertical linear components from an image can be configured as [[0,1,0], [0,1,0], [0,1,0]]. When a 3*3 convolution filter for extracting upper and lower linear components from an image is applied to an input image, the upper and lower linear components that match the convolution filter can be extracted and output from the image. A convolution layer can apply a convolution filter to each matrix for each channel representing the image (i.e., R, G, and B colors in the case of an R, G, B coded image). A convolution layer can apply a convolution filter to an input image to extract features that match the convolution filter from the input image. The filter values ​​of the convolution filter (i.e., the values ​​of each component of the matrix) can be updated by backpropagation during the training process of the convolutional neural network.

[0110] A subsampling layer may be connected to the output of a convolutional layer to simplify the output of the convolutional layer and reduce memory usage and computational complexity. For example, if the output of the convolutional layer is input to a pooling layer having a 2*2 max pooling filter, the maximum value contained in each 2*2 patch for each pixel of the image may be output to compress the image. The pooling may be a method of outputting the minimum value from the patch or an average value of the patch, and any pooling method may be included in the present disclosure.

[0111] A convolutional neural network may include one or more convolutional layers and subsampling layers. A convolutional neural network can extract features from an image by repeatedly performing convolution and subsampling processes (e.g., max pooling, as mentioned above). Through repeated convolution and subsampling processes, the neural network can extract global features of the image.

[0112] The output of a convolutional layer or a subsampling layer can be input to a fully connected layer, which is a layer in which all neurons in one layer are connected to all neurons in the adjacent layer. A fully connected layer is a structure in a neural network in which all nodes in each layer are connected to all nodes in other layers.

[0113] At least one of the CPU, GPGPU, and TPU of the processor 110 may process the training of the network function. For example, the CPU and GPGPU may work together to train the network function and classify data using the network function. In addition, in one embodiment of the present disclosure, processors of multiple computing devices may be used together to train the network function and classify data using the network function. In addition, a computer program executed in a computing device according to one embodiment of the present disclosure may be a CPU-, GPGPU-, or TPU-executable program.

[0114] The memory 140 can store a computer program for providing a platform provision method, and the stored computer program can be read and executed by the processor 150. The memory 170 can store any type of information generated or determined by the processor 110 and any type of information received by the communication processor 150.

[0115] The memory 140 can store data supporting various functions of the content generation platform server 100 and programs for the operation of the processor 110, can store input and output data (e.g., input image, first text, reconstructed image, second text of word bag, etc.), can store a number of application programs (or applications) run in the content generation platform server 100, and data and instructions for the operation of the content generation platform server 100. At least some of these applications can be downloaded from an external server via wireless communication.

[0116] Thus, the memory 140 may include at least one type of storage medium selected from the group consisting of flash memory, hard disk, solid state disk, silicon disk drive, multimedia card micro, card-type memory (e.g., SD or XD memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, and optical disk. The memory may also be a database separate from the device but connected by wire or wirelessly.

[0117] The communications processor 150 may include one or more components that enable communication with external devices, such as at least one of a broadcast reception module, a wired communications module, a wireless communications module, a near-field communications module, and a location information module.

[0118] Although not shown, the content generation platform server 100 of the present disclosure may further include an output section and an input section.

[0119] The output unit may display a user interface (UI) for providing label noise selection results, learning results, etc. The output unit may output any type of information generated or determined by the processor 110 and any type of information received by the communication processor 150.

[0120] The output unit may include at least one of a liquid crystal display (LCD), a thin film transistor-liquid crystal display (TFT LCD), an organic light-emitting diode (OLED), a flexible display, and a 3D display. Some of these display modules may be transparent or light-transmitting so that the outside can be seen through them. This is called a transparent display module, and a representative example of the transparent display module is a TOLED (Transparent OLED).

[0121] The input unit can receive information input by a user. The input unit can include keys and / or buttons on a user interface for receiving information input by a user, or physical keys and / or buttons. User input via the input unit can execute a computer program for controlling a display according to an embodiment of the present disclosure.

[0122] FIG. 30 is a flowchart illustrating the platform providing method of the present disclosure.

[0123] The following describes a platform providing method in which a content generation platform server 100 receives text from a user's terminal and provides a learning model that has been trained to output at least one reconstructed image corresponding to the text.

[0124] The processor 110 of the content generation platform server 100 receives (310) input content and first text to be matched to the input content from the user terminal via the content generation module 120. The input content may be an input image.

[0125] Alternatively, the input content may be a reconstructed image generated by an existing user through a first learning model, and the first text may be text matched to the reconstructed image.

[0126] In step 310, the processor 110 generates a recommendation text including sentences or words containing objects, appearances, and backgrounds by captioning the input image using the second learning model of the content generation module 120, and sends and outputs the recommendation text to the user's terminal.

[0127] The processor 110 also receives as the first text a final text determined or modified based on text input by a user from an existing user via the content generation module 120 and / or recommended text output based on an input image from an image captioning model.

[0128] Next, the processor 110 generates 320 a bag of words based on the first text via the content generation module 120.

[0129] In detail, the processor 110 can classify words (here, words include words, phrases, and clauses that make up the first text) contained in the first text into each word category and provide them as recommended words for each word category.

[0130] Next, the processor 110 determines caption data using the second text contained in the word bag and the predetermined sentence structure via the content generation module 120 (330).

[0131] At this time, the processor 110 can output, via an output unit (not shown), a predetermined sentence structure including a plurality of blanks and connective words between the plurality of blanks, to which word categories are matched to describe the input image through the image captioning model of the content generation module 120.

[0132] The processor 110 provides recommended words that can be input into each of a plurality of blank spaces, and can provide the recommended words taking into consideration the relevance to the input image and the presence or absence of a match with a word category.

[0133] In detail, the server 100 provides the user terminal with recommended words for the first text and recommended words for the input image so as to set categories for each word item of caption data having a predetermined sentence structure, and based on this, the server 100 can receive second text determined for each word item category in the user terminal according to user input, and determine the second text of the caption data.

[0134] The processor 110 can then automatically set the determined second text, connectives, and particles to determine text for the caption data. Next, the processor 110 inputs the text of the caption data into the first learning model via the content generation module 120 to generate at least one reconstructed content corresponding to the text represented by the caption data (340).

[0135] Next, the processor 110 provides the at least one reconstructed content via the content generation module 120 and the input content based on the caption data input process that generates the reconstructed content to the user terminal via the communication processor 150 for output (350).

[0136] The processor 110 may also control the process of inputting feedback for a reconstructed image selected by a user from among the at least one reconstructed image via the feedback processing module 130 to provide the process to the user terminal.

[0137] When the user terminal provides feedback on the reconstructed image to the feedback processing module 130, the user terminal can receive feedback opinions on the reconstructed image and pinpoints within the reconstructed image that match the feedback opinions in accordance with the user's operation.

[0138] When the processor 110 receives feedback from the user terminal via the communication processor 150, the processor 110 inputs the reconstructed image into the learning model, identifies common morpheme tokens associated with the reconstructed image, reconstructs a sentence of the caption data using the identified common morpheme tokens, and outputs the reconstructed sentence through the output unit. To this end, the processor 110 can read from the memory 140 and use a third learning model (not shown), which is a morpheme analyzer trained to preprocess and analyze text and separate it into morpheme units.

[0139] The processor 110 can modify the reconstructed sentences according to user actions via the feedback processing module 130 .

[0140] The processor 110 can then update the word bag based on feedback of the reconstructed image via the feedback processing module 130 to generate additional second text.

[0141] The processor 110 can then reconstruct the sentences of the caption data based on the feedback via the feedback processing module 130 .

[0142] The processor 110 may additionally generate second text according to the word category of each item in the word bag based on the feedback via the feedback processing module 130.

[0143] The processor 110 can generate, via the content generation module 120, an archive of each of the reconstructed images generated by the initial generation of the reconstructed images and the iterative execution of the feedback.

[0144] The processor 110 can then update the first learning model through the archive thus generated, and since the first learning model is trained using text based on the predetermined sentence structure of the caption data and a reconstructed image (e.g., an input image) that matches it, it can be trained with a high matching rate to the predetermined sentence structure.

[0145] The processor 110 may search for and provide real images having a similarity degree equal to or greater than a predetermined reference value based on at least one reconstructed image via the content generation module 120.

[0146] The processor 110 extracts caption data including sentences or words by captioning at least one reconstructed image through the content generation module 120, and can recommend new caption data by comparing the extracted caption data with the actual image.

[0147] Meanwhile, the method according to the present disclosure can be implemented as a program (or application) and stored on a medium to be executed in combination with a server that is hardware.

[0148] The disclosed embodiments may be realized in the form of a recording medium storing computer-executable instructions. The instructions may be stored in the form of program code, which, when executed by a processor, may generate program modules to perform the operations of the disclosed embodiments. The recording medium may be realized as a computer-readable medium.

[0149] Computer-readable recording media include all types of recording media that store computer-readable instructions, such as ROM (Read Only Memory), RAM (Random Access Memory), magnetic tape, magnetic disk, flash memory, and optical data storage devices.

[0150] The disclosed embodiments have been described above with reference to the accompanying drawings. Those skilled in the art will understand that the present disclosure can be implemented in forms different from the disclosed embodiments without changing the technical idea or essential features of the present disclosure. The disclosed embodiments are illustrative and should not be construed as limiting. [Industrial Applicability]

[0151] The present invention has industrial applicability because it is a method for generating user experience-based content using a learning model through processing by a server and a terminal.

Claims

1. a memory including a first learning model trained to generate reconstructed content based on text; and a processor that communicates with the memory and that, when text is input from a user terminal, controls the first learning model to output at least one reconstructed content corresponding to the text; The processor: input content and a first text that matches the input content from the user terminal; generating a bag of words based on the first text; determining caption data using a second text determined based on the first text included in the word bag and a predetermined sentence structure; inputting a sentence represented by the caption data into the first learning model to generate the at least one reconstructed content corresponding to the sentence; Concatenating the at least one reconstructed content and the input content and outputting the combined content to the user terminal; Content generation platform server.

2. the memory includes a second learning model that outputs at least one text representing an input image; If the input content is an input image, The processor: When the first text that matches the input content is input, generating a recommendation text including sentences or words including objects, appearances, and backgrounds by captioning the input image using the second learning model, and providing the recommendation text to the user terminal; At least one or more final texts determined or revised by the user based on the recommended texts are input as the first texts. The content creation platform server of claim 1 .

3. If the input content is an input image, The processor: providing the user terminal with a process for generating the caption data including the predetermined sentence structure constituting the caption data and word categories for each item constituting the sentence structure; Each of the itemized word categories is composed of at least one blank that is filled in by a user setting, A linking word is formed between the blanks of each word category. The content creation platform server of claim 1 .

4. The processor: providing the user terminal with recommended words that can be input into each of the plurality of blank spaces, and providing the recommended words taking into consideration their relevance to the input image and their matching with the word category; The content creation platform server of claim 3 .

5. When the reconstructed content is a reconstructed image, The processor: a feedback for a reconstructed image selected by the user from the at least one reconstructed image is input from the user terminal; additionally generating the second text for the bag of words based on the feedback; and reconstructing the text of the caption data based on the feedback; The content creation platform server of claim 1 .

6. The processor: generating additional second text for at least one word category of a word bag based on the feedback; The content creation platform server of claim 5 .

7. The processor: When feedback is input from the user terminal, A user interface is provided to the user terminal so that a feedback opinion on the reconstructed image and a pinpoint in the reconstructed image that matches the feedback opinion can be input by the user's operation. The content creation platform server of claim 5 .

8. the memory includes a third learning model that is a morphological analyzer trained to preprocess text and separate it into morphemes; The processor: After the feedback is entered, inputting the reconstructed image into the third learning model to determine common morpheme tokens associated with the reconstructed image, reconstructing sentences of the caption data using the determined common morpheme tokens, and providing the reconstructed sentences to the user terminal; The content creation platform server of claim 5 .

9. The user terminal modifying the reconstructed sentence according to the user's operation and transmitting the final reconstructed sentence to the content generation platform server; The processor: When the final reconstructed sentence is received, it is input to the first learning model, and the reconstructed image is output again and provided to the user terminal.

9. The content creation platform server of claim 8.

10. The processor: generating an archive of the initial generation of the reconstructed image and each of the reconstructed images generated by the repeated execution of the feedback; The content creation platform server of claim 5 .

11. The processor: searching for real images having a similarity equal to or greater than a predetermined reference value based on the at least one reconstructed image, and providing the real images to the user terminal; The content creation platform server of claim 5 .

12. The processor: extracting caption data including sentences or words from captioning each of the at least one reconstructed image, and comparing the extracted caption data with the actual image to recommend new caption data; 12. The content creation platform server of claim 11.

13. A method for generating user experience-based content by a processor of a user terminal in cooperation with a content generation platform server including at least one learning model, comprising: determining an input image and first text for the input image in response to a user input; image captioning the input image to provide a recommendation text including sentences or words for at least one category of object, appearance, or background; determining the first text for the input image according to a user input for the recommendation text; generating a bag of words based on the determined first text; providing the user with a process for setting caption data based on the bag of words; determining second text of the caption data in response to the user's input, and setting caption data in response to the determined second text and a predetermined sentence structure; a step of inputting a sentence represented by the set caption data, generating a reconstructed image through the at least one learning model, and outputting the generated reconstructed image. How to provide a content generation platform.

14. providing the user with a process for setting caption data based on the bag of words, and outputting the process for generating the caption data, the process including the predetermined sentence structure constituting the caption data and the word category for each item constituting the sentence structure.

14. The method of claim 13 for providing a content creation platform.

15. each of the itemized word categories in the process for generating caption data comprises at least one blank that is filled in by a user setting; A linking word is formed between the blanks of each word category.

15. The method of claim 14 for providing a content creation platform.

16. receiving, from the user, a feedback regarding a reconstructed image selected by the user from the at least one output reconstructed image; and reconstructing the text of the caption data based on the feedback.

14. The method of claim 13 for providing a content creation platform.

17. The step of inputting feedback to the reconstructed image includes: and inputting a feedback opinion on the reconstructed image and a pinpoint in the reconstructed image that matches the feedback opinion through the user's operation.

17. The method of claim 16 for providing a content creation platform.

18. The step of reconstructing the sentences of the caption data based on the feedback includes: inputting the reconstructed image into the third learning model to determine common morpheme tokens associated with the reconstructed image; reconstructing sentences of the caption data using the determined common morpheme tokens; and providing the reconstructed sentences to the user terminal.

17. The method of claim 16 for providing a content creation platform.

19. The step of reconstructing the sentences of the caption data based on the feedback includes: generating and providing the second text for at least one word category of the word bag based on the feedback; 17. The method of claim 16 for providing a content creation platform.

20. A program stored on a computer-readable recording medium, which, when combined with a computer, causes the program to execute the content creation platform providing method of claim 13.

Citation Information

Patent Citations

  • Translation device, translation method and translation program

    JP2012064059A

  • Image processor, image processing system, image processing method, program and recording medium

    JP2014171153A

  • Image processing apparatus, image processing method, and image processing program

    JP2020052947A

  • Image / text-based design generation device and method

    JP2021513181A

  • Image processing apparatus, image processing method, and program

    JP2023096759A