User experience-based content generation platform server and platform providing method

By designing a content generation platform server based on user experience, using learning models and bag of words technology, the problem of the existing technology being difficult to generate creative content that conforms to the user's will is solved, and extremely creative content generation is achieved.

CN119998800APending Publication Date: 2025-05-13LG MANAGEMENT & DEVELOPMENT INSTITUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380069940.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-26
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing artificial intelligence-based content generation methods are difficult to accurately output creative content when processing unfiltered languages ​​or information different from user needs.

Method used

A content generation platform server based on user experience is designed to receive text or images input by users through learning models in memory and output recombinant content or images. The platform uses bags of words and a pre-determined literary and scientific structure to generate descriptions and generate data that conforms to the user's will, and ultimately provides extremely creative content.

Benefits of technology

Through user experience and concepts, content reflects the user's own ideas is formed, and images and texts with extremely creative nature are provided, solving the problem that the existing technology is difficult to generate creative content that conforms to the user's will.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998800A_ABST
    Figure CN119998800A_ABST
Patent Text Reader

Abstract

The present invention provides a program that repeatedly implements an archive formation process that reflects the user's own idea on the basis of the user's experience and idea, thereby causing a user to generate an extremely creative image by means of a learning model for generating an image in accordance with the desired direction and on the basis of a text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a content generation platform server based on user experience and a platform providing method. Background Art

[0002] Currently, in an environment where a variety of content is provided, creative content generated by users is not created out of nothing, but can be generated by reflecting the user's personal inclinations such as experience, ideas, preferences, etc. on existing content.

[0003] Specifically, creative content is generated by connecting several existing contents in various ways. To this end, it is actually recommended to adopt a content generation method based on artificial intelligence.

[0004] However, when learning for the purpose of AI-based content generation, it is difficult to output creative content based on user intentions as unfiltered language or information that is different from user needs is input. Summary of the invention

[0005] Issues to be solved

[0006] The purpose of the embodiments disclosed herein is to provide a content generation platform server based on user experience and a platform providing method, so as to form the user's own ideas based on the user's experience and ideas and provide highly creative content.

[0007] The technical problems disclosed in the present invention are not limited to the technical problems mentioned above. Generally, a technician can clearly understand another technical problem not mentioned from the following content.

[0008] Problem Solution

[0009] The content generation platform server described in the embodiment of the present invention includes: a memory including a first learning model, the first learning model generates reorganized content based on text after learning; and a processor, which communicates with the memory and controls the first learning model to output at least one reorganized content corresponding to the text when receiving the text from the user terminal.

[0010] The processor receives input content and a first text matching the input content from the user terminal, generates a bag of words based on the first text, determines description generation data using a second text determined based on the first text included in the bag of words and a predetermined context structure, inputs a sentence showing the description generation data into the first learning model, generates the at least one reorganized content corresponding to the sentence, and enables the user terminal to connect the at least one reorganized content and the input content and output them.

[0011] At this time, the memory includes a second learning model, and the second learning model outputs at least one text representing the input image.

[0012] When the input content is an input image,

[0013] If the first text matching the input content is received, the processor can generate a recommended text including sentences or words of things, shapes and backgrounds by using the description generation (captioning) of the input image using the second learning model, provide it to the user terminal, and receive at least one final text determined or modified by the user based on the recommended text as the first text.

[0014] Furthermore, when the input content is an input image,

[0015] The processor provides a program for generating the description generation data to the user terminal, and the description generation data includes the predetermined textual structure constituting the description generation data and various word classes constituting the textual structure, wherein each word class is composed of at least one space filled by user setting, and the spaces between each word class can form conjunctions.

[0016] Furthermore, the processor provides recommended words that can be respectively input into the plurality of spaces to the user terminal, and provides the recommended words in consideration of the correlation between the input images and whether they are consistent with the word classes.

[0017] Furthermore, when the reorganized content is a reorganized image,

[0018] The processor may receive feedback of the recombined image selected by the user from the at least one recombined image from the user terminal, and further generate the second text of the word bag based on the feedback,

[0019] The sentences describing the generated data are reorganized based on the feedback.

[0020] Furthermore, the processor may further generate the second text for at least one word class in the word bag based on the feedback.

[0021] Furthermore, when receiving feedback from the user terminal, the processor may provide a user interface for the user terminal so as to receive feedback of the reconstructed image and a pinpoint in the reconstructed image matching the feedback according to the user's operation.

[0022] Furthermore, the memory includes a third learning model, which is used as a morpheme analyzer. The morpheme analyzer is trained to pre-process the text and separate it into morphemes.

[0023] Furthermore, after receiving the feedback, the processor can input the reorganized image into the third learning model, determine the common morpheme tokens related to the reorganized image, and use the determined common morpheme tokens to reorganize the sentences describing the generated data, and provide the reorganized sentences to the user terminal.

[0024] Furthermore, the user terminal modifies the reorganized sentence according to the user's operation, and transmits the finalized reorganized sentence to the content generation platform server.

[0025] When the finally determined reorganized sentence is received, the processor inputs it into the first learning model, outputs the reorganized image again, and provides it to the user terminal.

[0026] Furthermore, the processor may generate an archive of each of the reconstructed images, wherein the archive of each of the reconstructed images is generated following initial generation of the reconstructed images and repeated implementation of the feedback.

[0027] Furthermore, the processor may search for real images having a similarity greater than a preset standard value based on the at least one reconstructed image, and provide the real images to the user terminal.

[0028] Furthermore, the processor may extract description generation data including sentences or words through description generation of each of the at least one reorganized image, compare the extracted description generation data with the actual image, and recommend new description generation data.

[0029] The content generation platform providing method described in the embodiment of the present invention is a method for generating content based on user experience by a processor of a user terminal in linkage with a content generation platform server including at least one learning model, and includes the following steps: determining an input image and a first text corresponding to the input image in response to user input; realizing image description generation (captioning) of the input image, and providing a recommended text, which includes a sentence or word in at least one category of things, appearance or background; further determining the first text of the input image in response to user input of the recommended text; generating a bag of words based on the determined first text; providing a program for setting description generation data based on the bag of words to the user; determining a second text of the description generation data in response to user input, and setting the description generation data according to the determined second text and a predetermined context structure; inputting a sentence displayed by the set description generation data, generating a recombined image through the at least one learning model, and outputting the generated recombined image.

[0030] At this time, the step of providing to the user a program for setting description generation data based on the word bag may include the following steps: outputting a program for generating the description generation data, wherein the description generation data includes: the predetermined textual structure constituting the description generation data and the word classes of the various items constituting the textual structure.

[0031] Furthermore, the word classes of the various items generating the program describing the data generation are composed of at least one space filled by user setting, and the spaces between the word classes of the various items can constitute conjunctions.

[0032] Furthermore, the method may include the following steps: receiving feedback from the user on a reorganized image selected by the user from among the at least one reorganized image outputted; and reorganizing the sentences of the description generation data based on the feedback.

[0033] Furthermore, the step of receiving feedback of the recombined image may include the following steps: receiving feedback of the recombined image and a pinpoint in the recombined image matching the feedback according to the user's operation.

[0034] Furthermore, based on the feedback, the step of reorganizing the sentence describing the generated data may include the following steps: inputting the reorganized image into the third learning model, determining the common morpheme tokens related to the reorganized image, using the determined common morpheme tokens to reorganize the sentence describing the generated data, and providing the reorganized sentence to the user terminal.

[0035] Furthermore, the step of reorganizing the text describing the generated data based on the feedback may include the following steps:

[0036] Based on the feedback, further generate the second text corresponding to at least one word class of the word bag and provide

[0037] Effects of the Invention

[0038] According to the technical solution disclosed in the present invention, an archive reflecting the user's own ideas is formed based on the user's experience and ideas, and is used to provide highly creative images.

[0039] According to the technical solution disclosed in the present invention, the learning model is learned by using the text in the bag of words and the predetermined text structure, so that an image that meets the user's wishes can be generated.

[0040] The effects of the present disclosure are not limited to the effects mentioned above, and a person skilled in the art can clearly understand another effect not mentioned from the following contents. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a block diagram of the content generation platform server disclosed in the present invention.

[0042] Figures 2 to 29 is a schematic diagram of the content generation method disclosed herein.

[0043] Fig.30 is a flow chart of the platform providing method disclosed herein. DETAILED DESCRIPTION

[0044] Throughout this disclosure, the same reference symbols refer to the same components. This disclosure does not describe all the factors of the embodiments, but omits the conventional content of the technical field to which this disclosure belongs or the repetitive content between the embodiments. The terms "part, module, component, block" used in the specification can be implemented as software or hardware. According to the embodiment, multiple "parts, modules, components, blocks" can be implemented as one component, or one "part, module, component, block" can include multiple components.

[0045] Throughout the specification, when it is mentioned that a certain part is “connected” to other parts, it includes not only the case of direct connection but also the case of indirect connection. The indirect connection includes the connection formed through a wireless communication network connection.

[0046] Furthermore, when it is mentioned that a certain part “includes” a certain component, unless there is a particular description to the contrary, it means that other components are not excluded but further included.

[0047] Throughout the specification, when it is mentioned that a certain component is “on” other components, it includes not only a case where the certain component contacts other components but also a case where another component exists between the two components.

[0048] The terms first, second, etc. are used to distinguish a component from other components, and the component is not limited by the above terms.

[0049] Unless there is a clear exception in the context, expressions in the singular include expressions in the plural.

[0050] In each step, identification codes are used for convenience of description. The identification codes do not describe the order of the steps. Unless the context clearly states a specific order, the implementation of the steps may be different from the marked order.

[0051] Hereinafter, the working principle and embodiments of the present disclosure will be described with reference to the accompanying drawings.

[0052] In this specification, the "content generation platform server described in this disclosure" includes various devices that perform calculations and provide results to users. For example, the content generation platform server described in this disclosure may include a computer, a server device, and a portable terminal or a certain form thereof.

[0053] For example, the computer may include a notebook computer equipped with a web browser, a desktop computer, a laptop computer, a tablet computer, a touch-screen tablet computer, etc.

[0054] The server device, as a server that communicates with external devices and processes information, may include an application server, a computer server, a database server, a file server, a game server, an email server, a proxy server, and a web server, etc.

[0055] For example, the portable terminal, as a wireless communication device that can ensure portability and mobility, may include: Personal Communication System (PCS), Global System for Mobile communications (GSM), Personal Digital Cellular (PDC), Personal Handyphone System (PHS), Personal Digital Assistant (PDA), International Mobile Telecommunication (IMT)-2000, Code Division Multiple Access (CDMA)-2000, Wideband Code Division Multiple Access (W-CDMA), Wireless Broadband Internet (WiBro) terminals, smart phones, and other types of portable (handheld) wireless communication devices and wearable devices such as watches, rings, bracelets, anklets, necklaces, glasses, contact lenses or head-mounted devices (HMD).

[0056] The user experience-based content generation platform may include a content generation platform server and a terminal (not shown) that generates user experience-based content according to user input.

[0057] In an embodiment, the content generation terminal drives the content generation program to communicate with the content generation platform server. As the user inputs and transmits data, the data transmitted by the content generation platform server can be analyzed, content can be generated based on a learning model, and then, content based on user experience can be generated again by transmitting it to the terminal.

[0058] In another embodiment, the content generation terminal can implement user input and partial data processing, pass the analysis using the learning model to the content generation platform server for implementation, and generate content based on user experience by receiving the data output by the learning model and providing it to the user's program.

[0059] In yet another embodiment, after receiving a content generation program including a learning model from a content generation platform server, the content generation terminal may implement the received content generation program, analyze data directly input by the user, and implement a content generation method based on user experience.

[0060] The basis for the following description is that the content generation terminal receives user input, receives data analyzed by the learning model from the content generation platform server, and generates content based on user experience. Of course, according to another embodiment, the content generation program implemented by the content generation terminal and recorded as a part of the content generation method implemented by the content generation platform server described below is also included in the embodiments of the present invention. As mentioned above, the terminal may include a computer, a portable terminal, etc.

[0061] Figure 1 It is a block diagram of the content generation platform server disclosed in the present invention.

[0062] Below, reference Figures 2 to 29 Describe, Figures 2 to 29 is a schematic diagram of the content generation method disclosed herein.

[0063] like Figure 1 As shown, the content generation platform server 100 includes a processor 110, a memory 140 and a communication processor 150. At this time, the processor 110 can run a content generation program based on text to generate content and implement a content generation method. This content generation program includes a content generation module 120 and a feedback processing module 130, which can be run by the processor 110. The content generation module 120 and the feedback processing module 130 can control the actions of the various components. When implementing the content generation platform server 100 shown in the present disclosure, Figure 1 The components shown are not essential, and the content generation platform server 100 described in this specification may have more or less than the components described above.

[0064] The processor 110 communicates with the memory 140, and the content generation program may include a first learning model. When a text based on user input is received from the terminal, the first learning model can output at least one reorganized content corresponding to the text after learning.

[0065] In addition, it may be difficult for the user to write a text for generating the reorganized content desired by the user on such a first learning model, and when the reorganized content is not generated in the direction desired by the user, it may be difficult to find a solution.

[0066] The present invention archives the process of users directly inputting content and the text matching it through the content generation process, and based on this, assists users to compose text, thereby providing a content generation method based on user experience.

[0067] To this end, the processor 110 may enable the user to input the input content and the first text matching the input content.

[0068] The input content and the first text are directly selected by the user through archive input, and the user directly inputs (types) the first text matching the selected input content.

[0069] Furthermore, the first text may be a text formed by user input inputted into the first learning model, so as to grasp the initial concept of the reorganized content to be generated, and the reorganized content generated based on the first text may be the input content. That is, in order to grasp the initial concept, the user may input the first text into the reorganized content generated by the first learning model and set it as the input content. In the following, the input image of the input content and the reorganized content are exemplified as the reorganized image.

[0070] The first text may include personal reasons and opinions related to the input image, but is not limited thereto, as long as it is information that can reflect the user's experience. For example, when the input images are all fox images, the processor 110 may register Silly fox face and Strong touching texture related to the fox images input by the user as the first text. In this way, the first text does not only describe the object of the input image, but includes a variety of opinions such as the drawing style of the input image.

[0071] In order to generate an archive to set an initial concept reflecting the user's experience, the processor 110 may receive input content and a first text through the user's operation. The archive may match a variety of content (e.g., images) and related first texts for storage based on the user's experience. Then, the processor 110 may be applied to the generation of reorganized content that meets the user's will based on the archive including at least one input content and the first text matching it.

[0072] like Figure 2 , Figures 6 to 9 As shown, the processor 110 provides an item for registering an input image, according to the item selected by the user operation (Experience Archive->Image Archive->Upload input image), as shown in Fig. 9 As shown, a fox image can be registered. At this time, the processor 110 can receive the first text corresponding to the input image input with the user's operation. For example, the processor 110 can obtain the fox image and the first text related thereto such as Silly fox face and Strong touching texture with the user's operation.

[0073] like Fig.10 As shown, the processor 110 can match and display the input image and the first text. At this time, the first text can be at least one.

[0074] In addition, the present disclosure can provide the following recommendation service for the input of the first text through the learning model.

[0075] Figure 3 Can be Figure 2 The detailed diagram of the Documentation (Media to Text) process is shown in FIG. Figure 3 As shown, when receiving the first text matching the input content, the processor 110 can generate and output the recommended text including sentences or words by using the description generation (captioning) of the input image using the second learning model, and the sentences or words include things, appearance and background. At this time, the processor 110 can obtain the recommended text through the bag of words search after performing the description generation of the input image.

[0076] For example, the second learning model may include a first image description generation learning model, which, when input with an image, generates description text for each target object and situation within the image.

[0077] Furthermore, the second learning model may include a second image description generation learning model which generates text for describing other attribute word classes, ie, categories, characteristics, etc., when an input of an image and an input of an image are input.

[0078] Therefore, if Figure 3 As shown, when the input image is input into the image description generation learning model, the processor 110 can provide at least one recommended text to the user. In addition, the processor 110 can provide a user interface, which re-edits the first text of the input image on the display screen of the text input by the user and at least one recommended text (or recommended words) to enable the user to set an initial concept.

[0079] More specifically, in order to set the initial concept, the processor 110 may filter the text input by the user and the recommended text into basic text for generating the first text through a bow filter. Furthermore, the processor 110 may input the filtered basic text into a sentence construction learning model to generate a recommended first text of the input image and provide it to the user.

[0080] The user may modify and edit the provided first text to generate a final first text of the input image.

[0081] That is, the processor 110 may receive the final text determined or modified as the first text based on the user's input test and the recommended text. That is, the recommended text may be selected as it is or modified to be reflected as the first text according to the user's operation.

[0082] like Figures 11 to 13 As shown, the processor 110 can repeatedly receive multiple input images of the user's concept to be expressed and the first text matching the input image. And, as the user inputs, the processor 110 can connect and store in the archive according to the correlation between the input image, the first text, other input images and the first text of other input images. In this way, the state of connecting multiple input images and the first text can be expressed as follows: Fig.13 When storing the archive including the input image and the first text in the memory 140, the processor 110 may reflect this connection relationship and store it.

[0083] Then, the processor 110 can generate a bag of words based on the first text. After the above repeated process, there can be multiple first texts, and the processor 110 can generate a bag of words based on them. At this time, when forming a bag of words, the first text input by the user is used as the basis to reflect the characteristics of each user. Further, the processor 110 can recommend the first text for generating a bag of words by learning the input image. .

[0084] The processor 110 may determine the caption data using the second text included in the word bag and the predetermined text structure. At this time, the caption data may be a sentence indicating what content the image contains. The second text may mean the text included in the word bag formed based on the plurality of first texts.

[0085] The processor 110 may output a predetermined context structure through an output unit (not shown), the predetermined context structure comprising: a plurality of spaces matching word classes for describing the input image and connecting words between the plurality of spaces. The word classes may mean that words that should be input into the spaces are classified according to different attributes.

[0086] like Figures 14 to 20As shown, when the recombinant image generation program is run, the processor 110 generates description generation data (e.g., a sentence) for generating the desired recombinant image, and based on the generated description generation data, generates and provides the recombinant image through the first learning model. In order to assist in the generation of the description generation data, the processor 110 provides various different word class items as the basic sentences for recombinant image generation and at least one or more first texts or / and words included in the first texts that match each item, and determines the second text based on the first text (or the words belonging to) of each item selected by the user's operation (AI Image Generation->Sentence Builder->Select Word), thereby determining the description generation data.

[0087] At this time, the processor 110 can provide recommended words based on the first text, and the recommended words can be respectively input into each space corresponding to a plurality of different word class items. To this end, the processor 110 can consider the correlation between the first text and the input image and whether it is consistent with each word class, and provide the first text or / and the recommended words included in the first text.

[0088] For example, the processor 110 may provide a predetermined text structure, such as A'TYPE'of'BASE'that is'DETAIL'in the'STYLE'. In this case, 'TYPE', 'BASE', 'DETAIL', and 'STYLE' may be implemented as space items that require the input of words. Among them, words may mean sentences and fragments composed of multiple words, which may be the first text input by the user relative to the input image or / and the words recommended relative to the input image and included in the first text. Furthermore, each space may represent the first text matching each word class required to describe the image, such as type, base, detail, and style.

[0089] Furthermore, the processor 110 may further provide a reference word representing the attributes of each space item. Specifically, the processor 110 may recommend reference words that the user has not entered into the archive, but are representatively widely used and represent the attributes of each word class item, based on different type, base, detail, and style items. Furthermore, when recommending a reference word, the processor 110 may provide an image that matches the reference word, so that the user can intuitively understand the meaning of the reference word.

[0090] Furthermore, in the description generation data, the A, of, that is, in the as the conjunctions between the multiple spaces may include articles, particles, etc. that complete the sentence when the word is input in the space. That is, the conjunction disclosed in the present invention may mean a word that completes the sentence when the word is input in the space in a predetermined grammatical structure.

[0091] Specifically, in order to prevent the user from inputting a sentence whose form or content does not meet the standard value, or not knowing which sentence to input, or inputting a sentence that does not accurately reflect the user's intention, the processor 110 described in the present disclosure can allow the user to input a high-quality sentence that is helpful for learning the learning model. That is, the processor 110 guides the user on which words to include to combine sentences.

[0092] like Figures 16 to 20 As shown, the processor 110 can display a predetermined text structure, such as ATYPE of BASE that is DETAIL in the STYLE, through the output unit, and recommend displaying at least one word that is consistent with the word class in each space, and complete the input through the user.

[0093] For example, when the input image is a fox, and a predetermined text structure, i.e., A'TYPE'of'BASE'that is'DETAIL'in the'STYLE', is output on the interface, the processor 110 can provide a plurality of recommended words that can be input into the 'TYPE', 'BASE', 'DETAIL' and 'STYLE' blanks related to the fox through a word bag search. The processor 110 can input the oil-color painting, a fox, sitting in a field at sunrise, and realism art corresponding to 'TYPE', 'BASE', 'DETAIL' and 'STYLE' and selected by the user from the plurality of recommended words into each blank. The processor 110 can determine the description generation data based on the input words, such as the description generation data of "A oil-color painting of a fox that is sitting in a field at sunrise in the style of realism art."

[0094] That is, in order to enable the user to select a word corresponding to a style item, the processor 110 may provide a word belonging to a first text or a recommended word corresponding to an input image, the first text corresponding to a previously recorded archived style item. Furthermore, the processor 110 may provide a reference word representatively used for the style and a related reference image (Related Reference) that helps to understand the meaning of the reference word, so that the user can select a word corresponding to the style item.

[0095] like Figures 19 to 20As shown, the processor 110 can provide the words corresponding to the first text of the style (STYLE) item or the recommended words corresponding to the input image in the archive of previous records, so that the user can select the words corresponding to the style (STYLE) item. In addition, the processor 110 can provide reference words representatively used for the style (STYLE) and related reference images (Related Reference) that help understand the meaning of the reference words, so that the user can select the words corresponding to the style (STYLE) item. At this time, it will depend on the time consumed when generating the recombined image according to each word of the specified style (STYLE). Therefore, when determining the word corresponding to the style (STYLE), the processor 110 can pre-indicate the estimated time for image generation of the word according to the number of images generated.

[0096] like Figure 2 and Fig.21 As shown, the categories of the word bag may include type, figuration, base, description, action, and style.

[0097] At this time, type, shape, reference, description, action and style can be derived from object, expression, atmosphere, style, background and others codes.

[0098] The type may be a final displayed result (eg, a pattern or illustration, an operation, a photo), and the shape may be simplified, visualized, abstract, and the like.

[0099] Furthermore, the reference may be a core object and the description may be a description of the object.

[0100] Also, the action may be an action taken by the object.

[0101] The first codes of type, shape, reference, description, action and style mentioned above can be specifically classified and applied as the second code, thereby obtaining the prototype code of type, reference, specific and style.

[0102] Such a prototype code may be a word class that guides input into a space of a predetermined text structure. The scope of the word class and word bag is not limited to the above description, but may be changed according to the needs of the operator.

[0103] In addition, in order to improve the sentence completion of the determined description generation data, the processor 110 described in the present disclosure can perform the following processing.

[0104] In one example, the processor 110 described in the present disclosure includes a grammar checking function, so that a grammar check can be performed on the determined description generation data. The processor 110 can check the sentences that may have grammatical errors based only on the words input from the word bag search and the predetermined conjunctions, and modify the form of the conjunctions or the words input into the space, so as to accurately form the combination of the words input into the space. For example, the processor 110 can modify the part of speech of the word input into the space or modify the conjunctions to make it conform to the text structure (e.g., add, delete, change).

[0105] In another example, when prompting a recommended word searched from a word bag, the processor 110 of the present disclosure may modify the part of speech of the recommended word and prompt it out, taking into account the space where the recommended word is input and the contextual structure between the preset connectives. For example, the processor 110 may modify the part of speech of A or add connectives such as articles and particles to A based on the adjacent connectives of the space where the word A is input, and prompt it out as a recommended word.

[0106] The processor 110 may provide a tool to allow the user to freely input sentences based on the word bag to generate description generation data. Fig. 22 and Fig.23 As shown, with the selection of the Direct Generator item, the processor 110 can select the words in the word bag. In this process, the grammar checking function can be applied.

[0107] The processor 110 may input the description generation data into the learning model to generate at least one reorganized content corresponding to the first text. The reorganized content may be a reorganized image. At this time, the processor 110 may use multiple mode AIs that simultaneously input multiple program learning to generate the reorganized content corresponding to the first text.

[0108] like Figure 24 to Figure 26 As shown, the processor 110 can generate 16 reconstructed images based on the description generation data of “A painting of a fox sitting in a field at sunrise in the style of realism art.” generated by the description generation data generation program for generating reconstructed images.

[0109] like Fig. 27 As shown, the processor 110 may receive feedback of the reconstructed image selected by the user in at least one reconstructed image, and the processor 110 stores the feedback of the reconstructed image as the input image-first text in the archive, so as to facilitate its use again when generating the reconstructed image.

[0110] The feedback process may be performed after generating the at least one reorganized content and before connecting the at least one reorganized content and the input content and outputting the content. The present invention is not limited thereto. The process may also be performed after connecting the reorganized content and the input content and outputting the content.

[0111] The processor 110 can further generate a second text of the bag of words based on the feedback and based on the archive, the archive including the first text with the reorganized image as the input image. That is, the processor 110 receives feedback of the reorganized image, receives the text and accurately locates the reorganized image as feedback, generates the reorganized image as the input image-first text, and stores it in the archive. In addition, a procedure is provided for the processor 110 to generate a description generation data based on the additionally stored input image-first text, and the generation process of the reorganized image that gradually conforms to the user's will can be repeatedly implemented.

[0112] To this end, the processor 110 may further generate a second text based on the feedback and according to the category of the word bag.

[0113] like Fig.28 As shown, when receiving feedback, the processor 110 may receive feedback (O) of the reconstructed image and a pinpoint (P) in the reconstructed image matching the feedback according to the user's operation.

[0114] In addition, if Figure 4 As shown, when receiving feedback, the processor 110 can input the reorganized image into the learning model, grasp the common morpheme tokens related to the reorganized image, use the grasped common morpheme tokens to reorganize the sentence describing the generated data, and output the reorganized sentence through the output unit. At this time, the processor 110 can modify the reorganized sentence according to the user's operation.

[0115] like Figure 2 As shown, the processor 110 can generate archives of each reconstructed image (1st ImageArchiving ... Nth ... Nth Image Archiving), and the archives of each reconstructed image (1st ImageArchiving ... Nth ... Nth Image Archiving) are generated as the initial generation of the reconstructed image and the feedback are repeatedly implemented. The multiple archives generated in this way are stored in the memory 140 and can be used to generate images more in line with the user's will.

[0116] like Figure 5 As shown, the processor 110 searches for and provides real images having a similarity greater than a preset standard value based on at least one reconstructed group image.

[0117] And, if Figure 5As shown, the processor 110 generates descriptions of at least one of the reorganized images, extracts description generation data including sentences or words, compares the extracted description generation data with the actual image, and recommends new description generation data.

[0118] like Fig.29 As shown, the processor 110 can connect at least one reorganized content (A-1) and the input content (A) and output them.

[0119] The processor 110 disclosed in the present invention may be composed of more than one core, and may include a data analysis and deep learning processor such as a central processing unit (CPU), a general purpose graphics processing unit (GPGPU), and a tensor processing unit (Tensor processing unit) of a computer device. The processor 110 reads the computer program stored in the memory 140 to implement data processing for the machine learning described in the present invention. The processor 110 described in the present invention may implement operations for neural network learning. The processor 110 may perform calculations for neural network learning, such as processing input data for deep learning, extracting features from input data, calculating errors, and updating weighted values ​​of a neural network using backpropagation. Although not shown, the processor 110 described in the present invention inputs noisy training data including cleared label data and label noise data into a neural network model, selects label noise, mixes label noise data and cleared label data, and enables the diverter to achieve learning.

[0120] The neural network model can be a deep neural network. In this disclosure, neural networks, network functions, and neural networks (Neural networks) can be used as the same meaning. A deep neural network (DNN: deep neural network, deep neural network) means a neural network with multiple hidden layers in addition to an input layer and an output layer. When using a deep neural network, the latent structure of the data can be grasped. That is, the latent structure of photos, articles, videos, sounds, and music (for example, whether an object is in a photo, the content and feelings of an article, the content and feelings of a sound, etc.) can be grasped. A deep neural network can include a convolutional neural network (CNN: convolutional neural network), a recurrent neural network (RNN: recurrent neural network), a restricted Boltzmann machine (RBM: restricted boltzmannmachine), a deep belief network (DBN: deep belief network), a Q network, a U network, a siamese network, etc.

[0121] As a type of deep neural network, a convolutional neural network includes a neural network with a convolutional layer. A convolutional neural network is a type of multilayer perceptrons, which is designed to use minimal preprocessing. A CNN can be composed of one or more convolutional layers and an artificial neural network layer combined therewith. A CNN can further apply weighted values ​​and several pooling layers. Due to this structure, a CNN can fully utilize input data of a two-dimensional structure. To identify a target object from an image, a convolutional neural network can be used. A convolutional neural network can display image data as a matrix with dimensions for processing. For example, image data encoded into red, green, and blue (RGB, red-green-blue) can be displayed as a two-dimensional (e.g., a two-dimensional image) matrix according to different colors of R, G, and B. That is, the color value of each pixel of the image data can become a component of the matrix, and the matrix size can be equal to the image size. Therefore, the image data can be displayed as three two-dimensional matrices (three-dimensional data arrays).

[0122] In a convolutional neural network, the convolution process (input and output of the convolution layer) is performed by moving the convolution filter while multiplying the matrix components of the convolution filter and each position of the image. The convolution filter can be composed of a matrix of n*n shape. The convolution filter can usually be composed of a solid filter that is less than the total number of pixels in the image. That is, when an m*m image is input to a convolution layer (for example, a convolution layer with a convolution filter of size n*n), a matrix of n*n pixels including each pixel in the image becomes the product of the convolution filter and the components (that is, the product between the components in the matrix). By the product between the convolution filter, the components that match the convolution filter can be extracted from the image. For example, the composition of a 3*3 convolution filter is such as [[0,1,0], [0,1,0], [0,1,0]], so that the upper and lower straight line components can be extracted from the image. In order to extract the upper and lower straight line components from the image, when a 3*3 convolution filter is applied to the input image, the upper and lower straight line components that match the convolution filter can be extracted from the image and output. The convolutional layer can apply convolutional filters to each matrix of each channel of the displayed image (i.e., R, G, B colors when encoding an image). The convolutional layer can apply convolutional filters to the input image and extract features from the input image that match the convolutional filter. The filter values ​​of the convolutional filter (i.e., the values ​​of each component of the matrix) can be updated by back-propagation during the learning process of the convolutional neural network.

[0123] The output of the convolution layer can be connected to the subsampling layer to simplify the output of the convolution layer and reduce the amount of memory and computation. For example, when the output of the convolution layer is input to the pooling layer with a 2*2 maximum pooling filter, each 2*2 patch of each pixel in the image outputs the maximum value included in each patch, thereby compressing the image. The above pooling can also be a way of outputting the minimum value of the patch or outputting the average value of the patch, and any pooling method can be included in the present disclosure.

[0124] A convolutional neural network may include more than one convolutional layer and a subsampling layer. A convolutional neural network may repeatedly implement a convolution process and a subsampling process (e.g., the above-mentioned maximum pooling, etc.) to extract features from an image. A neural network may extract overall features of an image through repeated convolution and subsampling processes.

[0125] The output of the convolutional layer or the subsampling layer can be input to a fully connected layer. A fully connected layer is a layer in which all neurons in one layer are connected to all neurons in the adjacent layers. A fully connected layer can mean a structure in which all nodes in each layer are connected to all nodes in other layers in a neural network.

[0126] At least one of the CPU, GPGPU, and TPU of the processor 110 can process the learning of the network function. For example, the CPU and GPGPU can process the learning of the network function and the data classification using the network function together. In addition, in one embodiment of the present disclosure, the processors of multiple computer devices can be used together to process the learning of the network function and the data classification using the network function. In addition, the computer program implemented by the computer device described in one embodiment of the present disclosure can be a CPU, GPGPU, or TPU executable program.

[0127] The memory 140 may store a computer program for providing a platform providing method, and the stored computer program may be read and driven by the processor 150. The memory 170 may store any form of information generated or determined by the processor 110 and any form of information received by the communication processor 150.

[0128] The memory 140 can store data for supporting various functions of the content generation platform server 100 and programs for running the processor 110, can store input / output data (e.g., input image, first text, reorganized image, second text of word bag, etc.), can store multiple application programs (application programs or applications) driven by the content generation platform server 100, and several data and instructions for making the content generation platform server 100 work. At least part of such application programs can be downloaded from an external server through wireless communication.

[0129] Such memory 140 may include at least one storage medium in the form of flash memory type, hard disk type, solid state disk type (SSD type), silicon disk drive type (SDD type, Silicon Disk Drive type), multimedia card micro type, card memory (e.g., SD or XD memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk and optical disk. In addition, the memory is separated from the device, but it can also be a database connected by wire or wirelessly.

[0130] The communication processor 150 may include one or more components capable of communicating with an external device, for example, may include at least one of a broadcast receiving module, a wired communication module, a wireless communication module, a local area network module, and a location information module.

[0131] Although not shown, the content generation platform server 100 described in the present disclosure may further include an output unit and an input unit.

[0132] The output unit may display a user interface (UI) for providing label noise selection results and learning results, etc. The output unit may output any form of information generated or determined by the processor 110 and any form of information received by the communication processor 150 .

[0133] The output unit may include at least one of a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT LCD, a thin film transistor-liquid crystal display, an organic light-emitting diode (OLED), a flexible display, and a 3D display. Among them, part of the display module may be composed of a transparent type or a light-transmitting type so that the outside can be observed through it. This is called a transparent display module, and representative examples of the transparent display module include a transparent organic screen (TOLED, Transparent OLED) and the like.

[0134] The input unit can receive information input by the user. The input unit can have a plurality of keyboards and / or keys on a user interface or a mechanical keyboard and / or keys for receiving user input information. With the user input implemented through the input unit, a computer program for controlling the display described in the embodiment of the present disclosure can be run.

[0135] Fig.30 is a schematic flow chart of the platform providing method disclosed herein.

[0136] The following describes a method for providing a platform for providing a learning model, wherein the learning model is learned so that when the content generation platform server 100 receives text from a user terminal, it can output at least one reconstructed image corresponding to the text.

[0137] The processor 110 of the content generation platform server 100 may receive input content and first text matching the input content from the user terminal through the content generation module 120 (310). The input content may be an input image.

[0138] Furthermore, the input content may be a reconstructed image previously generated by a user through the first learning model, and the first text may be a text matching the reconstructed image.

[0139] In step 310, the processor 110 generates a recommended text including sentences or words by using the description generation (captioning) of the input image of the second learning model of the content generation module 120. The recommended text includes things, shapes and backgrounds, and transmits it to the user terminal and outputs it.

[0140] Furthermore, the processor 110 may generate a model through the content generation module 120 based on the text or / and image description input by the user in the past to output the recommended text based on the input image, and input the final text determined or modified as the first text.

[0141] Then, the processor 110 may generate a bag of words based on the first text through the content generation module 120 ( 320 ).

[0142] Specifically, the processor 110 may classify the words included in the first text (wherein the words are words constituting the first text, meaning sentences and fragments) into different word classes, and provide recommended words of different word classes.

[0143] Then, the processor 110 determines description generation data ( 330 ) by using the second text included in the word bag and the predetermined context structure through the content generation module 120 .

[0144] At this time, the processor 110 can output a predetermined contextual structure through an output unit (not shown) by using the image description generation model of the content generation module 120. The predetermined contextual structure includes: a plurality of spaces that match the word classes describing the input image and connecting words between the plurality of spaces.

[0145] The processor 110 provides recommended words that can be input into a plurality of different spaces, and provides the recommended words in consideration of relevance to the input image and whether or not the word class is consistent.

[0146] Specifically, the server 100 provides the user terminal with recommended words corresponding to the first text and recommended words corresponding to the input image, so as to set the categories of various different word items of the description generation data of a predetermined text structure, and on this basis, the content of the second text can be determined based on the categories of various different word items received from the user terminal as the user inputs, so as to determine the second text of the description generation data.

[0147] Furthermore, the processor 110 may automatically set the determined second text and the conjunctions and auxiliary words to determine the text describing the generated data. Then, the processor 110 may input the text describing the generated data into the first learning model through the content generation module 120 to generate at least one reorganized content corresponding to the text represented by the described generated data (340).

[0148] Secondly, the processor 110 can provide the following scheme to the user terminal through the communication processor 150, that is, through the content generation module 120, connect at least one reorganized content and generate the basic input content of the description generation data input program of the reorganized content and output it (350).

[0149] Furthermore, the processor 110 may provide a program for receiving feedback to the terminal through control, and the program receives feedback of a reconstructed image selected by a user among at least one reconstructed image through the program feedback processing module 130 .

[0150] When providing feedback of the reconstructed image to the feedback processing module 130 , the user terminal can receive the feedback of the reconstructed image and the pinpoint in the reconstructed image matching the feedback according to the user's operation.

[0151] Furthermore, the processor 110 receives feedback from the user terminal through the communication processor 150. When receiving feedback through the feedback processing module 130, the reorganized image can be input into the learning model, the common morpheme tokens related to the reorganized image can be mastered, and the sentences describing the generated data can be reorganized using the mastered common morpheme tokens, and the reorganized sentences can be output through the output unit. To this end, the processor 110 can read the morpheme analyzer, that is, the third learning model (not shown) from the memory 140 and apply it. The third learning model as the morpheme analyzer can be learned to pre-process and analyze the text and separate it in units of morphemes.

[0152] The processor 110 can modify and reorganize the sentence according to the user's operation through the feedback processing module 130.

[0153] Secondly, the processor 110 can update the word bag based on the feedback of the reorganized image through the feedback processing module 130 to further generate the second text.

[0154] Secondly, the processor 110 can reorganize the sentences describing the generated data based on the feedback through the feedback processing module 130.

[0155] The processor 110 may further generate a second text based on the feedback and the word classes of each item in the word bag through the feedback processing module 130 .

[0156] The processor 110 may generate an archive of each reconstructed image, which is generated by the content generation module 120 as the reconstructed image is initially generated and feedback is iteratively implemented.

[0157] Furthermore, the processor 110 can update the first learning model through the archive generated as described above. The learning basis of the first learning model is the text describing the predetermined context structure of the generated data and the recombinant image (for example, the input image) matching it. Therefore, during learning, the matching rate of the predetermined context structure is high.

[0158] The processor 110 may search for and provide real images having a similarity greater than a preset standard value based on at least one reconstructed image through the content generation module 120 .

[0159] The processor 110 performs description generation for at least one of the reconstructed images through the content generation module 120, extracts description generation data including sentences or words, compares the extracted description generation data with the actual image, and recommends new description generation data.

[0160] In addition, the above method disclosed in the present invention can be implemented as a program (or application program) and stored in a medium so as to be implemented in combination with a hardware server.

[0161] The disclosed embodiments may be implemented as a recording medium storing computer-implementable instructions. The instructions may be stored in a program code form, which, when implemented by a processor, may generate a program module to implement the actions of the disclosed embodiments. The recording medium may be implemented as a computer-readable recording medium.

[0162] Computer-readable recording media include all types of recording media that store computer-readable instructions, such as read-only memory (ROM), random access memory (RAM), magnetic tape, magnetic disk, flash memory, optical data storage device, etc.

[0163] The above content describes the disclosed embodiments with reference to the accompanying drawings. It should be understood that a person skilled in the art of the present disclosure may implement the present disclosure in a form different from the disclosed embodiments without changing the technical ideas or essential features of the present disclosure. The disclosed embodiments are illustrative and should not be interpreted as being limited.

[0164] Industrial Application Possibility

[0165] The present invention implements a method for generating content based on user experience using a learning model through processing of a server and a terminal, and has industrial applicability.

Claims

1. A content generation platform server, characterized in that: include: A memory including a first learning model, wherein the first learning model is learned to generate reorganized content based on text; as well as a processor, which communicates with the memory and controls the first learning model to output at least one reorganized content corresponding to the text when receiving the text from the user terminal; The processor receives input content and a first text matching the input content from the user terminal, generates a word bag based on the first text, determines description generation data using a second text determined based on the first text included in the word bag and a predetermined context structure, inputs a sentence displaying the description generation data into the first learning model, generates the at least one reorganized content corresponding to the sentence, and enables the user terminal to connect the at least one reorganized content and the input content and output them.

2. The content generation platform server according to claim 1, characterized in that: The memory includes a second learning model, the second learning model outputting at least one text representing an input image, When the input content is an input image, If the first text matching the input content is received, the processor can generate a recommended text including sentences or words of things, shapes and backgrounds by utilizing the description of the input image using the second learning model, provide the text to the user terminal, and receive at least one final text determined or modified by the user based on the recommended text as the first text.

3. The content generation platform server according to claim 1, characterized in that: When the input content is an input image, the processor provides a program for generating the description generation data to the user terminal, and the description generation data includes the predetermined text structure constituting the description generation data and various word classes constituting the text structure, wherein each word class is composed of at least one space filled by user setting, and the spaces between each word class can form a conjunction.

4. The content generation platform server according to claim 3, characterized in that: The processor provides recommended words that can be respectively input into the plurality of spaces to the user terminal, and provides the recommended words in consideration of the correlation between the input images and whether they are consistent with the word classes.

5. The content generation platform server according to claim 1, characterized in that: When the reorganized content is a reorganized image, the processor may receive feedback of a reorganized image selected by the user from the at least one reorganized image from the user terminal, and further generate the second text of the word bag based on the feedback; and The sentences describing the generated data are reorganized based on the feedback.

6. The content generation platform server according to claim 5, characterized in that: The processor may further generate the second text for at least one word class of the bag-of-words based on the feedback.

7. The content generation platform server according to claim 5, characterized in that: When receiving feedback from the user terminal, the processor may provide a user interface for the user terminal, so as to receive feedback of the reconstructed image and precise positioning within the reconstructed image matching the feedback according to the user's operation.

8. The content generation platform server according to claim 5, characterized in that: The memory includes a third learning model, which is used as a morpheme analyzer. The morpheme analyzer is learned to pre-process the text to separate morphemes. Furthermore, after receiving the feedback, the processor can input the reorganized image into the third learning model, determine the common morpheme tokens related to the reorganized image, and use the determined common morpheme tokens to reorganize the sentences describing the generated data, and provide the reorganized sentences to the user terminal.

9. The content generation platform server according to claim 8, characterized in that: The user terminal modifies the reorganized sentence according to the user's operation, and transmits the finalized reorganized sentence to the content generation platform server. When the finally determined reorganized sentence is received, the processor inputs it into the first learning model, outputs the reorganized image again, and provides it to the user terminal.

10. The content generation platform server according to claim 5, characterized in that: The processor may generate an archive of each of the reconstructed images, the archive of each of the reconstructed images being generated following initial generation of the reconstructed images and repeated implementation of the feedback.

11. The content generation platform server according to claim 5, characterized in that: The processor may search for an actual image having a similarity greater than a preset standard value based on the at least one reconstructed image, and provide the actual image to the user terminal.

12. The content generation platform server according to claim 11, characterized in that: The processor may extract description generation data including sentences or words through description generation of each of the at least one reorganized image, compare the extracted description generation data with the actual image, and recommend new description generation data.

13. A method for providing a content generation platform, characterized in that: A method for generating content based on user experience by a processor of a user terminal in conjunction with a content generation platform server including at least one learning model comprises the following steps: Following the user's input, determining an input image and a first text corresponding to the input image; A method for generating an image description of the input image is provided, wherein the recommended text includes a sentence or a word in at least one of the following categories: object, appearance or background; Following user input of the recommended text, further determining the first text of the input image; generating a word bag based on the determined first text; providing the user with a program for setting description generation data based on the bag of words; Determining a second text of the description generation data according to the user's input, and setting the description generation data according to the determined second text and a predetermined text structure; The set sentence describing the generated data display is input, a reconstructed image is generated by the at least one learning model, and the generated reconstructed image is output.

14. The method for providing a content generation platform according to claim 13, characterized in that: The step of providing the user with a program for setting description generation data based on the word bag includes the following steps: outputting the program for generating the description generation data, wherein the description generation data includes: the predetermined textual structure constituting the description generation data and the word classes of the various items constituting the textual structure.

15. The method for providing a content generation platform according to claim 14, characterized in that: The word classes of the different items generating the program describing the data generation are composed of at least one space filled by a user setting, and the spaces between the word classes of the different items constitute conjunctions.

16. The method for providing a content generation platform according to claim 13, characterized in that: The following steps are involved: receiving feedback of a reconstructed image selected by the user from the at least one reconstructed image outputted; Based on the feedback, the sentences describing the generated data are reorganized.

17. The method for providing a content generation platform according to claim 16, characterized in that: The step of receiving feedback of the recombined image may include the following steps: receiving feedback of the recombined image and a precise positioning within the recombined image matching the feedback in response to the user's operation.

18. The method for providing a content generation platform according to claim 16, characterized in that: Based on the feedback, the step of reorganizing the sentence describing the generated data includes the following steps: inputting the reorganized image into the third learning model, determining the common morpheme tokens related to the reorganized image, using the determined common morpheme tokens to reorganize the sentence describing the generated data, and providing the reorganized sentence to the user terminal.

19. The method for providing a content generation platform according to claim 16, characterized in that: The step of reorganizing the sentence describing the generated data based on the feedback comprises the following steps: Based on the feedback, the second text corresponding to at least one word class in the word bag is further generated and provided.

20. A program combined with a computer and stored in a computer-readable recording medium for implementing the content generation platform providing method according to claim 13.