System
The system addresses inefficiencies in affiliate article generation by preprocessing data, training a generative AI model, and performing internal checks to produce high-quality articles with reduced legal risks and manual effort.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2026-03-05
AI Technical Summary
Existing systems face inefficiencies and legal risks in generating affiliate articles due to errors, misrepresentations, and the need for significant manual work in data preprocessing, style and tone adjustments, and internal checks, which affect the quality and consistency of the articles.
A system that preprocesses data, trains a generative AI model to match specific styles and tones, automatically generates articles with legal annotations, and performs internal checks to ensure accuracy, then publishes them efficiently.
Enables the efficient generation and publication of high-quality affiliate articles with reduced legal risks and manual workload, ensuring consistency in style and tone.
Smart Images

Figure 2026036119000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Errors, misrepresentations, and missing notes in affiliate articles pose legal risks, and quality issues with these articles reduce a company's competitive advantage in the market. Furthermore, checking and correcting these issues requires a significant amount of man-hours, increasing the workload. A comprehensive solution to these issues is needed. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for preprocessing collected data and training a generative AI model, a means for extracting style and tone from advertiser websites and affiliate sites and training the generative AI model, a means for automatically generating affiliate articles based on the collected and learned data, a means for appropriately inserting advertising notices and annotations into the generated articles, a means for internally checking the generated articles, and a means for publishing the generated articles on a website after the internal check. This reduces the risk of errors, misrepresentations, and omitted notes, enabling the efficient generation of high-quality affiliate articles. It also reduces the workload.
[0006] "Data preprocessing" is the process of removing unnecessary information from collected data and shaping it into a form suitable for analysis and learning.
[0007] A "generative artificial intelligence model" is an artificial intelligence model that has the ability to generate language and sentences based on input data.
[0008] "Training" is the process of using collected data to teach a generative artificial intelligence model to perform a specific task.
[0009] "Style" refers to characteristics of a piece of writing, such as its form, structure, style, and tone.
[0010] "Tone" refers to the emotion or attitude conveyed in a piece of writing, and influences the impression it leaves on the reader.
[0011] An "affiliate article" is an article that introduces a specific product or service and encourages readers to make a purchase or application through a link.
[0012] "Advertising notices" refers to statements regarding warnings and legal restrictions regarding the information provided in the advertisement.
[0013] "Notes" refers to supplementary explanations, cautions, and legal requirements inserted into the text.
[0014] "Internal check" is the process of reviewing the content of generated articles to verify that there are no errors, misrepresentations, or missing notes.
[0015] "Publishing on a website" means uploading the generated article to a website on the Internet and making it available for general users to view. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] The following description is provided as an example of an embodiment of the present invention: The system collects data from advertisers and affiliate sites and uses a generative artificial intelligence model (e.g., ChatGPT®) to generate highly accurate affiliate articles.
[0038] System Configuration
[0039] The system mainly consists of the following components:
[0040] 1. Data Collection Module
[0041] 2. Data Preprocessing Module
[0042] 3. Generative AI Model Training Module
[0043] 4. Style and Tone Extraction Module
[0044] 5. Annotation Learning Module
[0045] 6. Article Generation Module
[0046] 7. Internal Check Module
[0047] 8. Article Publishing Module
[0048] Data collection
[0049] First, the server collects the necessary data from the advertiser and affiliate sites. The collected data includes detailed product information, review articles, and past posting data from the affiliate sites. The server obtains the data from each site using web crawling technology.
[0050] Data Preprocessing
[0051] Next, the server preprocesses the collected data. This step involves data cleaning to remove unnecessary information, such as removing HTML tags, deleting unnecessary strings using regular expressions, tokenizing the text, and extracting important key phrases and features.
[0052] Training generative artificial intelligence models
[0053] Using the preprocessed data, the server trains a generative AI model (ChatGPT), which fine-tunes the model based on the initial dataset to learn specific writing styles and expressions.
[0054] Style and Tone Extraction
[0055] The server extracts the style and tone of the advertiser's and affiliate's sites and trains the generative AI model. During this process, the server analyzes each site's writing style, tone, and punctuation usage. This data is then fed back into the generative model to generate articles that match the characteristics of each site.
[0056] Annotation Learning
[0057] The server trains a generative AI model to create appropriate ad descriptions and annotations. This involves collecting standard annotations based on legal requirements and teaching the model necessary annotation information, such as "This product is not a medicine."
[0058] Article Generation
[0059] The server automatically generates affiliate articles based on the collected and learned data. For example, when generating an article about "Healthcare Product A," the generative AI model emphasizes the product's effects and features and generates text that matches the tone and style. Annotations based on legal requirements are also automatically inserted into the article.
[0060] Internal Check
[0061] The generated articles are then checked internally by a user (an internal checker). A checklist is used to check for typographical errors, misrepresentations, and missing notes. During this process, the quality can be checked again using the generative AI model. If necessary, the user can make manual corrections.
[0062] Article Publication
[0063] Finally, articles that have passed the check are published on the website via the terminal, which automatically posts the articles based on the publication schedule and records the publication log.
[0064] Specific examples
[0065] For example, generating an affiliate article for healthcare product A involves the following process:
[0066] 1. The server collects information about the ingredients, efficacy, and reviews of healthcare product A.
[0067] 2. The server preprocesses the collected data, removing HTML tags, deleting unnecessary strings, and tokenizing the text.
[0068] 3. The server trains the ChatGPT model to learn specific writing styles and expressions.
[0069] 4. The server extracts the advertiser's formal tone and the affiliate site's friendly tone and trains the model accordingly.
[0070] 5. The server learns the advertising and annotations and applies them to the generated articles.
[0071] 6. The server generates an article such as, "Healthcare product A contains ingredient B and is expected to be effective in improving immunity."
[0072] 7. The user checks the generated article for errors, misrepresentations, and missing notes, and makes corrections as necessary.
[0073] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[0074] The above system allows high-quality affiliate articles to be generated efficiently and published with reduced legal risks.
[0075] The processing flow will be explained below.
[0076] Step 1:
[0077] The server collects data from advertiser and affiliate sites. Specifically, it uses a crawler to retrieve the HTML of web pages from the specified URLs and stores it in a database. The retrieved data includes titles, body text, meta tags, image URLs, etc.
[0078] Step 2:
[0079] The server preprocesses the collected data, removing HTML tags and extracting only the text data. Regular expressions are also used to remove unnecessary symbols and spaces. Next, the text is tokenized and techniques such as TF-IDF and Word2Vec are used to extract important key phrases and features.
[0080] Step 3:
[0081] The server uses the preprocessed data to train a generative AI model, based on the initial ChatGPT model, which is then fine-tuned to fit specific writing styles and themes. This process uses high-quality samples from the collected data as training data.
[0082] Step 4:
[0083] The server extracts the style and tone of the advertiser and affiliate sites, uses analysis algorithms to identify the unique writing style and tone patterns of each site, and feeds these features back into the generative AI model.
[0084] Step 5:
[0085] The server trains a generative AI model to create appropriate annotations. Standard annotations and warnings based on legal requirements are collected and used as training data for the model. This training gives the model the ability to automatically insert appropriate annotations when generating articles.
[0086] Step 6:
[0087] The server automatically generates affiliate articles based on the data it has collected and learned. For example, it generates articles that emphasize the effectiveness and features of specific products and services using information about those products and services. The generated articles also include necessary annotations.
[0088] Step 7:
[0089] User-generated articles are internally checked using a dedicated checklist to ensure there are no errors, misrepresentations, or missing notes. This process can also be repeated using a generative AI model to check quality.
[0090] Step 8:
[0091] The device publishes the checked articles on the website. Articles are automatically posted according to the publication schedule. The URL and related information after publication are recorded in a database and stored as data for analysis.
[0092] Through the above steps, high-quality affiliate articles can be efficiently generated and published with reduced legal risks.
[0093] Example 1
[0094] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0095] In conventional affiliate article generation systems, a lot of manual work is required to efficiently generate a large amount of content while improving the quality of the articles, which leads to issues such as inconsistency in the articles and inconsistency in the quality of the articles.In addition, inserting accurate annotations based on legal requirements and internal article checking processes require a great deal of effort.
[0096] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0097] In this invention, the server includes means for preprocessing the collected data and training the generative AI model, means for extracting style and tone from advertiser information sources and websites and training the generative AI model, and means for automatically generating text based on the collected and learned data, thereby enabling the efficient generation of high-quality affiliate articles and the provision of content with a consistent style and tone.
[0098] "Collected data" refers to detailed product information, review articles, past posting data, etc. obtained from advertisers and affiliate sites.
[0099] "Preprocessing" refers to a series of processes that remove HTML tags and unnecessary strings from collected data, tokenize the text, and extract important key phrases and features.
[0100] A "generative artificial intelligence model" refers to an artificial intelligence that learns a specific writing style or expression method based on an initial dataset.
[0101] "Training methods" refers to the use of collected data to fine-tune generative AI models to learn specific writing styles and idiomatic expressions.
[0102] "Advertiser Source" refers to a website or database that contains information about an advertiser.
[0103] "Style and tone extraction" refers to the process of analyzing each site's writing style, tone, and punctuation usage and training the model.
[0104] "Means for automatic generation" refers to a method of automatically generating sentences based on collected and learned data using a generative artificial intelligence model.
[0105] "Means for appropriately inserting annotations" refers to a method for automatically inserting standard legal annotations into the generated text.
[0106] "Means for internal checking" refers to a method in which an in-house checker reviews the generated text to check for any typos, misrepresentations, or missing notes.
[0107] "Means of disclosing to the information source" refers to a method in which the generated text is automatically posted on a website after being verified and a disclosure log is recorded.
[0108] "Crawling" refers to the act of collecting data from a website using an automated crawler program.
[0109] "HTML tag stripping" refers to the process of removing unnecessary HTML code from text data obtained from a web page.
[0110] "Extracting text information" refers to the method of extracting the actual required text parts from the preprocessed data.
[0111] "Methods for checking for errors, misrepresentations, and missing notes" refers to methods for checking the content of generated articles and checking for typos, inaccurate expressions, and missing necessary notes.
[0112] This invention relates to a system that collects data from advertisers and affiliate sites and generates highly accurate affiliate articles using a generative artificial intelligence model (e.g., a generative AI model). The system automatically inserts annotations based on legal requirements into the generated articles, and then publishes them on a website after internal checks.
[0113] Hardware and Software
[0114] The system requires the following hardware and software:
[0115] Server: Collects data, preprocesses it, trains the generative AI model, generates articles, and publishes them.
[0116] Terminal: Serves as an interface for publishing checked articles on a website.
[0117] User: Internal checks and corrections of generated articles.
[0118] The main software used includes:
[0119] Web crawler tools: Used to collect data from advertiser and affiliate sites.
[0120] Data preprocessing library: Used to remove HTML tags, tokenize text, and extract key phrases.
[0121] Generative artificial intelligence model (generative AI model): Refers to a generative artificial intelligence platform used to generate articles, such as a generative AI model or ChatGPT.
[0122] Linguistic Style Analysis Library: Used to extract style and tone.
[0123] Text review tools: Used for internal checks and to identify errors, misstatements, and missing notes.
[0124] Specific processing and calculation of data
[0125] 1. Data Collection:
[0126] The server uses a web crawler tool to automatically retrieve product details, reviews, and past posting data from advertiser and affiliate sites, and stores this data in an internal database.
[0127] 2. Data Preprocessing:
[0128] The server performs preprocessing on the collected data, such as removing HTML tags, deleting unnecessary text using regular expressions, tokenizing, and extracting important key phrases.
[0129] 3. Training the generative AI model:
[0130] The server uses the preprocessed data to train a generative artificial intelligence model, a process that customizes the model by teaching it specific writing styles and idioms.
[0131] 4. Extracting style and tone:
[0132] The server uses a language style analysis library to analyze the style and tone of advertiser and affiliate sites, and the analysis data is fed back to a generative AI model.
[0133] 5. Annotation Learning:
[0134] The server collects annotations based on legal requirements and trains a generative artificial intelligence model based on them.
[0135] 6. Article Generation:
[0136] The server uses a generative artificial intelligence model to generate affiliate articles based on the collected and learned data, and the generated articles are automatically supplemented with the necessary legal annotations.
[0137] 7. Internal check:
[0138] The generated articles are then internally checked by users who use text review tools to identify errors, misstatements, and omissions, and make manual corrections as needed.
[0139] 8. Article Publication:
[0140] Finally, articles that have passed the check are published on the website via the terminal, which posts the articles based on the publication schedule and records the publication log.
[0141] Examples of concrete examples and prompts
[0142] For example, let's say you want to generate an affiliate article for "Healthcare Product A":
[0143] 1. The server collects information about the ingredients, efficacy, and reviews of healthcare product A.
[0144] 2. The server preprocesses the collected data and removes HTML tags and unnecessary strings.
[0145] 3. The server trains the generative AI model to learn specific writing styles and expressions.
[0146] 4. The server extracts the advertiser's formal tone and the affiliate site's friendly tone and trains the model.
[0147] 5. The server learns the advertising notations and annotations and applies them to the generated article.
[0148] 6. The server generates an article stating, "Healthcare product A contains ingredient B and is expected to be effective in improving immunity."
[0149] 7. The user checks the generated article for errors and missing notes.
[0150] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[0151] Example prompt sentence:
[0152] "Use a generative AI model to create an affiliate ad article about healthcare product A based on its features, ingredients, and user reviews. Focus on its immune-boosting benefits. Also, be sure to include legal annotations."
[0153] With the above configuration, the present invention makes it possible to efficiently generate high-quality affiliate articles, reduce legal risks, and reliably publish them.
[0154] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0155] Step 1:
[0156] Data collection
[0157] The server collects the necessary data from advertisers and affiliate sites.
[0158] Input: URL list, crawling settings
[0159] Processing: Use a web crawler tool to retrieve detailed product information, reviews, and past posting data from the specified URL.
[0160] Output: Collected data (HTML file, text data, etc.)
[0161] What it does: The server runs a crawler tool that parses the HTML content of each website to extract data and store it in an internal database.
[0162] Step 2:
[0163] Data Preprocessing
[0164] The server preprocesses the collected data.
[0165] Input: Collected data (HTML files, text data, etc.)
[0166] Processing: Strip HTML tags, remove unnecessary strings, tokenize text, and extract important key phrases.
[0167] Output: Preprocessed text data
[0168] What it does: It uses a regular expression library to remove HTML tags and unnecessary text, a morphological analyzer to break down text into words, and a keyphrase extraction algorithm to extract important keywords.
[0169] Step 3:
[0170] Training generative artificial intelligence models
[0171] The server uses the preprocessed data to train a generative artificial intelligence model (ChatGPT).
[0172] Input: Preprocessed text data
[0173] Processing: Fine-tuning the model to learn specific writing styles and phrasing.
[0174] Output: A trained generative artificial intelligence model
[0175] What it does: It processes batches of data, feeds it into the model, and fine-tunes it by setting optimal hyperparameters.
[0176] Step 4:
[0177] Style and Tone Extraction
[0178] The server extracts the style and tone of the advertiser and affiliate sites.
[0179] Input: Text data from advertiser and affiliate sites
[0180] Processing: Uses a linguistic style analysis library to analyze writing style, tone, and punctuation usage.
[0181] Output: Extracted style and tone data
[0182] Specific operations: Analyzes text data from each site, extracts elements that characterize writing style and tone, and feeds this back to a generative AI model.
[0183] Step 5:
[0184] Annotation Learning
[0185] The server trains a generative artificial intelligence model to generate annotations based on legal requirements.
[0186] Input: Legal annotation database
[0187] Processing: Train the model with standard annotations based on legal requirements.
[0188] Output: A generative AI model that can apply annotations
[0189] What it does: Legal annotations are fed into the model as a dataset, and the model is trained to accurately insert annotations under specific conditions.
[0190] Step 6:
[0191] Article Generation
[0192] The server generates affiliate articles using a generative artificial intelligence model.
[0193] Input: Trained generative artificial intelligence model, collected and preprocessed data
[0194] Processing: Using data, we automatically generate affiliate articles that match your tone and style.
[0195] Output: Generated affiliate article
[0196] What it does: The model generates sentences based on the prompts you provide and automatically adds any necessary legal annotations.
[0197] Step 7:
[0198] Internal Check
[0199] Internally check user-generated articles.
[0200] Input: Generated affiliate article
[0201] Action: Check for clerical errors, misrepresentations, and missing notes.
[0202] Output: Affiliate articles checked or corrected
[0203] What you'll do: Proofread articles using text review tools, conduct team reviews, and manually correct as needed.
[0204] Step 8:
[0205] Article published
[0206] The device publishes the checked articles on the website.
[0207] Input: Checked affiliate article
[0208] Processing: Post articles based on the publishing schedule and record publishing logs.
[0209] Output: Published affiliate articles
[0210] What it does: Uses a scheduler tool to automatically post articles at specified times and records the publishing history.
[0211] (Application example 1)
[0212] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0213] Conventional ad article generation systems have problems with inefficiency and time because the data collection and preprocessing process is complicated and style and tone adjustments are often done manually. Other issues include difficulty in previewing and editing generated articles, and difficulty in automatically posting articles to content management systems. Furthermore, there is a lack of a mechanism for quickly generating articles using smartphones.
[0214] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0215] In this invention, the server includes means for preprocessing collected data and training a generative AI model, means for extracting style and tone from the advertiser's website and affiliated sites and training the generative AI model, means for automatically generating advertising articles based on the collected and learned data, means for appropriately inserting advertising notices and annotations into the generated articles, means for internally checking the generated articles, means for publishing the generated articles on the website after the internal check, means for quickly and easily generating articles using a smartphone, means for displaying the generated articles during preview and editing and enabling editing, and means for automatically posting the generated articles via the API of a content management system, thereby enabling efficient and highly accurate generation and publication of advertising articles.
[0216] "Collected data" refers to information such as product information, reviews, and past posting data obtained from advertisers and affiliated sites.
[0217] "Preprocessing" is the process of removing unnecessary information from collected data and shaping the data.
[0218] A "generative artificial intelligence model" is an artificial intelligence algorithm that can generate new text based on input data.
[0219] "Training" is the learning process used to improve the accuracy of a generative artificial intelligence model based on specific data.
[0220] "Style" refers to characteristics such as writing style, tone, and format.
[0221] "Tone" describes the mood or emotional atmosphere of a piece of writing.
[0222] An "advertising piece" is a piece of writing written to promote a particular product or service.
[0223] "Advertising notation" refers to text or notes that clearly indicate that the item is an advertisement.
[0224] An "annotation" is an explanatory text based on supplementary information or legal requirements inserted within an article.
[0225] "Internal checking" is the process of checking the content of generated articles and correcting errors and inappropriate expressions.
[0226] "Publishing on a website" means posting the generated article on the web so that it can be viewed by general internet users.
[0227] "Using a smartphone" means performing various operations and processes using a smartphone.
[0228] "Preview" is a function that allows you to check the state before the final release.
[0229] "Editing" is the process of correcting or changing the content of the generated article.
[0230] A "content management system" is software for organizing, managing, and delivering content.
[0231] An "API" is an interface that allows applications and services to communicate with each other and utilize their functionality.
[0232] "Automatically posting" means that the system automatically uploads content to a website based on a set schedule or conditions.
[0233] As an embodiment of the present invention, a novel advertising article generation system is described, which preprocesses collected data, trains a generative artificial intelligence model, and ultimately automatically generates and publishes highly accurate advertising articles.
[0234] System Configuration
[0235] The system of the present invention includes the following major components:
[0236] 1. Data Collection Module
[0237] 2. Data Preprocessing Module
[0238] 3. Generative AI Model Training Module
[0239] 4. Style and Tone Extraction Module
[0240] 5. Annotation Learning Module
[0241] 6. Article Generation Module
[0242] 7. Internal Check Module
[0243] 8. Article Publishing Module
[0244] 9. Smartphone Application Module
[0245] 10. Preview and Edit Module
[0246] 11. Auto-posting module
[0247] Hardware and Software Used
[0248] The main hardware and software components of this system include:
[0249] Hardware: High-performance servers and smartphones.
[0250] Software: Python, BeautifulSoup, Scrapy, pandas, transformers, React Native, Content Management Systems (CMS) WordPress and Wix, ChatGPT by OpenAI®.
[0251] Processing Description
[0252] 1. Data Collection
[0253] The server collects the necessary data from advertisers and partner sites using APIs and web crawling technologies (BeautifulSoup and Scrapy).
[0254] 2. Data Preprocessing
[0255] The collected data is cleaned using a data preprocessing module. Specifically, HTML tags are removed, unnecessary strings are deleted using regular expressions, and important key phrases and features are extracted. The data is formatted and processed using the pandas library.
[0256] 3. Generative AI Model Training
[0257] The server uses the preprocessed data to train a generative AI model (ChatGPT), which learns specific writing styles and expressions.
[0258] 4. Extracting Style and Tone
[0259] The server extracts style and tone from the text of advertisers and partner sites and feeds this back into a generative artificial intelligence model.
[0260] 5. Annotation Learning
[0261] The server trains the model with annotations based on legal requirements and applies them to the generated advertising articles.
[0262] 6. Article Generation
[0263] Using a generative artificial intelligence model, advertising articles are automatically generated based on collected and learned data.
[0264] 7. Internal Checks
[0265] The generated articles are then internally checked by the user to ensure there are no errors, misrepresentations, or missing notes, and any necessary corrections are made. The quality can also be checked again using the generative AI model.
[0266] 8. Public
[0267] Once the article has been checked, it will be published on the website via the terminal.
[0268] 9. Smartphone use
[0269] Users can quickly and easily create, preview, edit and publish articles using their smartphones.
[0270] 10. Preview and Edit
[0271] It displays a preview of the generated article and allows users to easily make any necessary corrections. It is provided with an interface that uses React Native.
[0272] 11. Auto-posting
[0273] Finally, the generated articles are automatically posted to the website via the content management system's API.
[0274] Specific examples
[0275] For example, generating an advertisement for a health device A involves the following process:
[0276] 1. The server collects information about the ingredients, efficacy, and reviews of health device A.
[0277] 2. The server preprocesses the collected data, removing HTML tags, deleting unnecessary strings, and tokenizing the text.
[0278] 3. The server trains the ChatGPT model to learn specific writing styles and expressions.
[0279] 4. The server extracts the advertiser's formal tone and the partner site's friendly tone and trains the model accordingly.
[0280] 5. The server learns the advertising and annotations and applies them to the generated articles.
[0281] 6. The server generates an article such as "Health device A contains ingredient B and is expected to be effective in improving immunity."
[0282] 7. The user checks the generated article for errors, misrepresentations, and missing notes, and makes corrections as necessary.
[0283] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[0284] Prompt Sentence Examples
[0285] Product name: Health equipment A
[0286] Features: Quiet design, foldable, multi-function display
[0287] As described above, the present invention enables efficient and highly accurate generation and publication of advertising articles.
[0288] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0289] Step 1:
[0290] Data collection
[0291] The server uses APIs and web crawling technologies (BeautifulSoup and Scrapy) to retrieve product details, reviews, and past postings from advertisers and partner sites. The input is the URL or API endpoint of each site, and the output is the retrieved raw data. This allows a wealth of information to be stored in the database used.
[0292] Step 2:
[0293] Data Preprocessing
[0294] The server preprocesses the collected data. Specifically, it uses the pandas library to remove HTML tags, delete unnecessary strings, and extract important key phrases and features. The input is the raw data collected in step 1, and the output is cleansed and formatted data. This process processes the data into a form that is easy for the generative artificial intelligence model to process.
[0295] Step 3:
[0296] Generative AI model training
[0297] The server uses the preprocessed data to train a generative AI model (ChatGPT). In the process, it learns specific writing styles and expressions. Specifically, it fine-tunes the model using the transformers library. The input is the preprocessed data, and the output is a trained generative AI model. This allows the model to generate sentences with even greater accuracy.
[0298] Step 4:
[0299] Style and Tone Extraction
[0300] The server performs text analysis to extract style and tone from the text of advertisers' and partner sites. Specifically, it uses machine learning algorithms and natural language processing (NLP) technology to extract text characteristics. The input is text data from advertisers' and partner sites, and the output is the extracted style and tone features. This ensures that the generated articles are tailored to the characteristics of each site.
[0301] Step 5:
[0302] Annotation Learning
[0303] The server collects specific texts (legal annotations) and trains a generative AI model to learn annotations based on legal requirements. Specifically, it identifies patterns in the annotations and incorporates them into the model. The input is a dataset of legal annotations, and the output is a generative AI model that can insert annotations appropriately. This results in articles that reduce legal risk.
[0304] Step 6:
[0305] Article Generation
[0306] The server uses a generative artificial intelligence model to automatically generate advertising articles based on collected and learned data. Specifically, when a prompt is entered, the corresponding trained model generates an article. The input is the prompt and past posting data, and the output is an automatically generated advertising article. This enables fast and highly accurate article generation.
[0307] Example prompt sentence:
[0308] Product name: Health equipment A
[0309] Features: Quiet design, foldable, multi-function display
[0310] Step 7:
[0311] Internal Check
[0312] The generated article undergoes an internal check by the user. Specifically, it checks for typographical errors, misrepresentations, and missing notes, and makes any necessary corrections. The input is the automatically generated advertising article, and the output is the corrected final version of the article. This allows the article to be published with guaranteed quality.
[0313] Step 8:
[0314] Article published
[0315] Once the article has been checked, it is published on the website via the terminal. Specifically, the article is automatically posted via the API of the content management system (CMS). The input is the final, corrected version of the article, and the output is the published article. This allows the article to be published efficiently on the web, saving the user time and effort.
[0316] Step 9:
[0317] Smartphone use
[0318] Users use their smartphones to create, preview, edit, and publish articles. Specifically, articles are created and edited using a dedicated application. The input is the requirements and prompts specified by the user, and the output is the edited article. This allows article creation and management to be done anywhere, anytime.
[0319] Step 10:
[0320] Preview and Edit
[0321] A preview of the generated article is displayed, allowing the user to easily make any necessary corrections. Specifically, the interface uses React Native, and edits are reflected in real time. The input is the automatically generated article, and the output is the article corrected by the user. This further improves the quality of the final article.
[0322] Step 11:
[0323] Auto-post
[0324] Finally, the generated articles are automatically posted to the website via the content management system's API. The server uploads the articles to the website based on the set schedule and conditions. The input is the checked, final version of the article, and the output is the published article and its publication log. This minimizes user effort and efficiently publishes articles through an automated process.
[0325] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0326] The following description is provided as an embodiment of the present invention. This system collects data from advertisers and affiliate sites and generates highly accurate affiliate articles using a generative artificial intelligence model (e.g., ChatGPT). Furthermore, by combining it with an emotion engine that recognizes user emotions, the system can be customized according to the user's emotions.
[0327] System Configuration
[0328] The system mainly consists of the following components:
[0329] 1. Data Collection Module
[0330] 2. Data Preprocessing Module
[0331] 3. Generative AI Model Training Module
[0332] 4. Style and Tone Extraction Module
[0333] 5. Annotation Learning Module
[0334] 6. Emotion Engine Module
[0335] 7. Article Generation Module
[0336] 8. Internal Check Module
[0337] 9. Article Publishing Module
[0338] Data collection
[0339] First, the server collects the necessary data from the advertiser and affiliate sites. The collected data includes detailed product information, review articles, and past posting data from the affiliate sites. The server obtains the data from each site using web crawling technology.
[0340] Data Preprocessing
[0341] Next, the server preprocesses the collected data. This step involves data cleaning to remove unnecessary information, such as removing HTML tags, deleting unnecessary strings using regular expressions, tokenizing the text, and extracting important key phrases and features.
[0342] Training generative artificial intelligence models
[0343] Using the preprocessed data, the server trains a generative AI model (ChatGPT), which fine-tunes the model based on the initial dataset to learn specific writing styles and expressions.
[0344] Style and Tone Extraction
[0345] The server extracts the style and tone of the advertiser's and affiliate's sites and trains the generative AI model. During this process, the server analyzes each site's writing style, tone, and punctuation usage. This data is then fed back into the generative model to generate articles that match the characteristics of each site.
[0346] Annotation Learning
[0347] The server trains a generative AI model to create appropriate ad descriptions and annotations. This involves collecting standard annotations based on legal requirements and teaching the model necessary annotation information, such as "This product is not a medicine."
[0348] Emotion Engine
[0349] The server uses an emotion engine to analyze the user's emotions in real time. The emotion engine recognizes the user's emotional state based on user input and various sensor data (e.g., facial recognition and voice analysis). Based on the results, the style and tone of the generated articles can be adjusted.
[0350] Article Generation
[0351] The server automatically generates affiliate articles based on the collected and learned data. The generative AI model emphasizes the product's benefits and features and generates text that matches the tone and style. For example, if the user is in a positive emotional state, text with a more optimistic tone will be generated. Annotations based on legal requirements are also automatically inserted into the generated articles.
[0352] Internal Check
[0353] The generated articles are then checked internally by a user (an internal checker). A checklist is used to check for typographical errors, misrepresentations, and missing notes. During this process, the quality can be checked again using the generative AI model. If necessary, the user can make manual corrections.
[0354] Article Publication
[0355] Finally, articles that have passed the check are published on the website via the terminal, which automatically posts the articles based on the publication schedule and records the publication log.
[0356] Specific examples
[0357] For example, generating an affiliate article for healthcare product A involves the following process:
[0358] 1. The server collects information about the ingredients, efficacy, and reviews of healthcare product A.
[0359] 2. The server preprocesses the collected data, removing HTML tags, deleting unnecessary strings, and tokenizing the text.
[0360] 3. The server trains the ChatGPT model to learn specific writing styles and expressions.
[0361] 4. The server extracts the advertiser's formal tone and the affiliate site's friendly tone and trains the model accordingly.
[0362] 5. The server learns the advertising and annotations and applies them to the generated articles.
[0363] 6. The server uses an emotion engine to analyze the user's emotions and adjust the tone of the article accordingly. For example, if the user is excited, the article will be generated with a more enthusiastic tone.
[0364] 7. The user reviews the generated article, checking for errors, misrepresentations, and missing notes, and makes corrections as necessary.
[0365] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[0366] The above system makes it possible to efficiently generate high-quality affiliate articles that respond to user sentiment and publish them while reducing legal risks.
[0367] The processing flow will be explained below.
[0368] Step 1:
[0369] The server collects data from advertiser and affiliate sites. Specifically, it uses a crawler to retrieve the HTML of web pages from the specified URLs and stores it in a database. The retrieved data includes titles, body text, meta tags, image URLs, etc.
[0370] Step 2:
[0371] The server preprocesses the collected data, removing HTML tags and extracting only the text data. Regular expressions are also used to remove unnecessary symbols and spaces. Next, the text is tokenized and techniques such as TF-IDF and Word2Vec are used to extract important key phrases and features.
[0372] Step 3:
[0373] The server uses the preprocessed data to train a generative AI model, based on the initial ChatGPT model, which is then fine-tuned to fit specific writing styles and themes. This process uses high-quality samples from the collected data as training data.
[0374] Step 4:
[0375] The server extracts the style and tone of the advertiser and affiliate sites, uses analysis algorithms to identify the unique writing style and tone patterns of each site, and feeds these features back into the generative AI model.
[0376] Step 5:
[0377] The server trains a generative AI model to create appropriate annotations. Standard annotations and warnings based on legal requirements are collected and used as training data for the model. This training gives the model the ability to automatically insert appropriate annotations when generating articles.
[0378] Step 6:
[0379] The server starts an emotion engine that analyzes the user's emotions in real time. The emotion engine detects emotions from the user's input (text, voice, facial recognition data, etc.) and records the results in a database.
[0380] Step 7:
[0381] The server adjusts the style and tone of the generated article based on the user's emotional data analyzed by the emotion engine. For example, if the user is in a positive emotional state, the server generates an optimistic tone of text.
[0382] Step 8:
[0383] Affiliate articles are automatically generated based on the data collected and learned by the server. As a specific example, when generating an article for "Healthcare Product A," the article is generated by incorporating product features and user reviews and adjusting the tone.
[0384] Step 9:
[0385] User-generated articles are internally checked using a dedicated checklist to ensure there are no errors, misrepresentations, or missing notes. Again, during this process, the generative AI model can be used to check the quality.
[0386] Step 10:
[0387] The device publishes the checked articles on the website. Articles are automatically posted according to the publication schedule. The URL and related information after publication are recorded in a database and stored as data for analysis.
[0388] Through the above steps, high-quality affiliate articles that respond to user sentiment can be efficiently generated and published with reduced legal risks.
[0389] Example 2
[0390] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0391] Conventional affiliate article generation systems have had issues with preprocessing collected data, matching the style and tone of generated articles, appropriately inserting advertisements and annotations, and customizing articles to reflect user sentiment. In particular, generating content based on user sentiment is difficult, and generated articles often do not reflect the user's sentiment. There is also a need for more efficient internal checks of generated articles.
[0392] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0393] In this invention, the server includes means for preprocessing collected data and training a generative artificial intelligence model, means for extracting style and tone from the advertiser's site and affiliate site and training the generative artificial intelligence model, and means for the generative artificial intelligence model to analyze user emotions and adjust the style and tone of articles. This enables the generation of more personalized affiliate articles in response to user emotions. Furthermore, using efficiently preprocessed data allows for the efficient generation of high-quality articles, contributing to the efficiency of internal checks.
[0394] "Collected data" refers to information such as detailed product information, review articles, and past posting data obtained from advertiser and affiliate sites using web crawling technology.
[0395] "Preprocessing" refers to a series of processes performed on collected data, such as removing HTML tags, deleting unnecessary strings, tokenizing text, and extracting key phrases and features.
[0396] A "generative artificial intelligence model" is an artificial intelligence algorithm that generates sentences based on training data, and a specific example is ChatGPT.
[0397] "Training" is the process of using preprocessed data to teach a generative artificial intelligence model a particular style or method of expression.
[0398] "Advertiser's Site and Affiliate Site" refers to a website that provides detailed product information and review articles, and includes affiliate marketing content.
[0399] "Style and tone" refers to the writing style, tone, punctuation, and other characteristics of a particular site or document.
[0400] "Learning" is the process by which a generative AI model analyzes data to acquire a particular style or tone, and then applies the results to the model.
[0401] "User emotion" refers to the emotional state analyzed based on the user's input information (text, voice, facial image, etc.), and includes positive, negative, excited, etc.
[0402] "Internal checking" is a process in which the generated articles are checked for errors, misrepresentations, missing notes, etc., and is carried out by in-house checkers.
[0403] "Advertising statements and annotations" refers to standard legally required annotations and advertising statements, such as warnings such as "This product is not a medicine."
[0404] "Publishing on a website" means posting the checked article on a website based on a specific publishing schedule and recording a publishing log.
[0405] To implement this invention, the following system must be constructed. This system uses data collected from advertisers and affiliate sites and a generative artificial intelligence model to automatically generate highly accurate affiliate articles. Furthermore, it is possible to analyze user sentiment and adjust the style and tone of the resulting articles.
[0406] System configuration
[0407] The system consists of the following main components:
[0408] 1. Data Collection Module
[0409] 2. Data Preprocessing Module
[0410] 3. Generative AI Model Training Module
[0411] 4. Style and Tone Extraction Module
[0412] 5. Annotation Learning Module
[0413] 6. Emotion Engine Module
[0414] 7. Article Generation Module
[0415] 8. Internal Check Module
[0416] 9. Article Publishing Module
[0417] Hardware and software used
[0418] The implementation of this system uses the following hardware and software:
[0419] Server: A server with high-performance processing power is used to process collected data, train generative AI models, analyze emotions, and perform other heavy workloads.
[0420] Generative AI models, such as ChatGPT, are used to generate text and adjust style and tone.
[0421] Emotion engine: An engine for analyzing user emotions, using sensors such as facial recognition and voice analysis.
[0422] Program processing
[0423] The server first collects data from advertiser and affiliate sites, including product information, reviews, and past posting data, and the collected data is obtained using web crawling technology.
[0424] The server then preprocesses the collected data, removing HTML tags, eliminating unnecessary strings, tokenizing the text, and extracting important key phrases and features.
[0425] Based on the preprocessed data, the server trains a generative AI model (ChatGPT), fine-tuning the model based on the initial dataset to learn specific writing styles and expressions.
[0426] The server then extracts the style and tone of the advertiser and affiliate sites and trains them into a generative AI model. Each site's writing style, tone, and punctuation usage are analyzed and fed back to the model.
[0427] The server also trains a generative AI model to create appropriate ad labeling and annotations, collecting standard annotations based on legal requirements and training the model to apply them to the generated articles.
[0428] Additionally, the server uses an emotion engine to analyze users' emotions in real time, recognizing their emotional state based on user input and various sensor data (facial recognition and voice analysis), and adjusting the style and tone of the generated articles accordingly.
[0429] As a concrete example, the prompt text to automatically generate an affiliate article for healthcare product A is as follows:
[0430] "Generate an article with an optimistic tone based on the ingredients, efficacy, and reviews of healthcare product A when the user is in a positive emotional state."
[0431] Finally, the generated article is checked internally by the user (an in-house checker). A checklist is used to check for typographical errors, misrepresentations, and missing notes, and manual corrections are made as necessary. The completed article is published on the website via the terminal, and a publication log is recorded.
[0432] summary
[0433] This system allows us to efficiently generate high-quality affiliate articles that reflect user sentiment and publish them while reducing legal risks.
[0434] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0435] Step 1:
[0436] Data collection
[0437] The server collects data from advertiser and affiliate sites. It uses web crawling technology to obtain HTML content based on the URL list specified as input. Specifically, it collects detailed product information, review articles, and past posting data from advertiser and affiliate sites. The output is raw data, which can be in a wide variety of formats.
[0438] Step 2:
[0439] Data Preprocessing
[0440] The data collected by the server is preprocessed. The input is the raw data collected in step 1. Specifically, regular expressions are used to remove HTML tags and extract only the text. Unnecessary strings (e.g., advertising banner text) are also filtered. The text is then tokenized and split into words. Important key phrases are then extracted using algorithms such as TF-IDF and Word2Vec. The output is preprocessed, clean text data.
[0441] Step 3:
[0442] Training generative artificial intelligence models
[0443] The server uses the preprocessed data to train a generative AI model (ChatGPT). The input is preprocessed, clean text data. Specifically, the initial dataset is fed into the generative AI model, and fine-tuned to learn specific writing styles and expressions. The output is a trained generative AI model.
[0444] Step 4:
[0445] Style and Tone Extraction
[0446] The server analyzes the style and tone of advertiser and affiliate sites and trains a generative AI model. The input is the text data of each site. Specific processing involves analyzing writing style (frequency of use of punctuation at the end of sentences, use of honorifics, etc.) and tone (formal or informal, optimistic or pessimistic, etc.). The results of these analyses are then fed back to the generative AI model, enabling it to generate text that matches the characteristics of each site. The output is a model that has learned the characteristics of style and tone.
[0447] Step 5:
[0448] Annotation Learning
[0449] The server trains a generative AI model on appropriate methods for advertising labeling and annotations. The input is standard annotations based on legal requirements (e.g., "This product is not a medicine"). Specific processing involves collecting standard annotations and training the generative AI model so that they can be applied when needed. The output is a model that can insert annotations.
[0450] Step 6:
[0451] Emotion Engine
[0452] The server uses an emotion engine to analyze the user's emotions in real time. The input is the user's input information (text, voice, facial image, etc.). Specifically, the process uses sensors such as facial recognition and voice analysis to recognize the user's emotional state (positive, negative, excited, etc.). Based on the results, the style and tone of the generated article are adjusted. The output is text whose tone and style have been adjusted to reflect the user's emotional state.
[0453] Step 7:
[0454] Article Generation
[0455] The server automatically generates affiliate articles based on collected and learned data. The input is a model that has learned style and tone characteristics that reflect emotional states, as well as requirements such as product information. A generative AI model is used to generate text that highlights the product's features and benefits. Annotations based on legal requirements are also automatically inserted. The output is the completed affiliate article.
[0456] Step 8:
[0457] Internal Check
[0458] The user (an internal reviewer) checks the generated article. The input is the generated affiliate article. Specifically, a checklist is used to check for errors, misrepresentations, and missing notes. Manual corrections may be made as necessary. The output is the final version of the article after corrections have been completed.
[0459] Step 9:
[0460] Article Publication
[0461] The terminal publishes the checked article on the website. The input is the final version of the article. Specific processing involves setting a publishing schedule and automatically posting the article based on the specified date and time. A publishing log is also recorded. The output is the published article and its publishing log.
[0462] (Application example 2)
[0463] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0464] Conventional affiliate article generation systems have difficulty generating advertising content that takes user emotions into account, making it impossible to provide ads that are adapted to user emotions. Furthermore, it is difficult to properly reflect the style and tone of the advertiser or affiliate site, making it impossible to provide an effective advertising experience for users.
[0465] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for preprocessing collected data and training a generative AI model, means for extracting style and tone from the advertiser's website and affiliate site and training the generative AI model, means for automatically generating affiliate articles based on the collected and learned data, means for appropriately inserting advertising notices and annotations into the generated articles, means for internally checking the generated articles, means for publishing the articles generated after the internal check on the website, and means for recognizing user emotions and adjusting the tone and style of the generated advertising content based on the emotions. This makes it possible to provide effective advertising content that corresponds to the user's emotions and generate and publish high-quality affiliate articles while reflecting the style and tone of the advertiser and affiliate site.
[0466] "Collected data" refers to information such as detailed product information, review articles, and past posting data obtained from advertisers and affiliate sites.
[0467] "Preprocessing" refers to the process of cleaning the collected data, removing unnecessary information, and extracting important key phrases and features using regular expressions.
[0468] A "generative artificial intelligence model" is an artificial intelligence model that learns specific writing styles and methods of expression based on given data, and generates new sentences based on the results of that learning.
[0469] "Training" refers to the learning process of using preprocessed data to improve the performance of a generative artificial intelligence model.
[0470] "Style and tone extraction" refers to the process of analyzing characteristics such as writing style and punctuation usage from advertiser websites and affiliate sites, and training a generative AI model to learn these characteristics.
[0471] An "affiliate article" refers to a piece of writing that provides information about a specific product or service and is intended to encourage users to purchase that product or service.
[0472] "Blank and annotation insertion" refers to the process of automatically adding standard legally required annotations to generated articles.
[0473] "Internal Check" refers to the internal review process to ensure that the content of the generated article is accurate and that there are no misrepresentations or omissions.
[0474] "Publication" refers to making the generated article publicly available on a website or other media.
[0475] "Emotion recognition" refers to the process of analyzing a user's emotional state using user input and sensor data (e.g., facial recognition and voice analysis).
[0476] "Tone and style adjustment" refers to the process of changing the tone and presentation of generated text based on the perceived user sentiment.
[0477] The present invention provides an emotion-based advertising content generation system, which includes means for preprocessing collected data and training a generative AI model, means for extracting style and tone from advertiser and affiliate sites and training the generative AI model, means for automatically generating affiliate articles based on the collected and learned data, means for appropriately inserting advertising notices and annotations into the generated articles, means for internally checking the generated articles, means for publishing the generated articles on a website after the internal check, and means for recognizing user emotions and adjusting the tone and style of the generated advertising content based on the emotions.
[0478] Program processing explanation
[0479] The server uses an emotion recognition module to analyze the user's emotions in real time. This module utilizes facial recognition technology using the smartphone camera and TENSORFLOW (registered trademark). Specifically, it captures the user's face and analyzes their emotions using a trained emotion recognition model (e.g., CNN model).
[0480] The server then uses web crawling techniques to collect the required data from the advertiser's website and affiliate sites, using Python web scraping libraries (e.g., BeautifulSoup and Requests), and preprocesses the collected data by removing HTML tags and extracting text information using regular expressions.
[0481] Based on the preprocessed data, the server trains a generative AI model (e.g., GPT-2). During this process, the collected data is analyzed for its writing style and tone, and the model is trained to learn from it. In particular, the model extracts features of the advertiser's formal writing style and the affiliate site's friendly tone.
[0482] The generated advertisements are automatically inserted with ad captions and annotations to comply with legal requirements. In this process, a generative model is trained on standard annotations collected in advance, and the resulting advertisement captions and annotations are added to the generated articles.
[0483] After generation is complete, the server performs an internal check to ensure there are no errors, misrepresentations, or missing notes. This includes a second quality check using the generative AI model. Once the check is complete, the article is published to the website via the terminal. Here, posts are automatically made based on the publishing schedule, and a publishing log is recorded.
[0484] Examples of concrete examples and prompts
[0485] For example, generating an affiliate article for healthcare product A involves the following process:
[0486] The server uses an emotion recognition module to analyze whether the user is in a positive emotional state and generates an affiliate article with a positive tone. The generated article uses the following prompt:
[0487] "Preprocess this data, normalize it, filter it, tokenize it."
[0488] "Generate ad text based on sentiment and preprocessed data."
[0489] "Please display the generated ad text on your smartphone screen."
[0490] In this way, it becomes possible to provide effective advertising content that matches the user's emotions, and high-quality affiliate articles that reflect the style and tone of the advertiser and affiliate site are generated.
[0491] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0492] Step 1:
[0493] Recognize user emotions.
[0494] The server captures the user's face in real time using the smartphone camera. The captured image is input to an emotion recognition model using TensorFlow to analyze the user's emotional state (e.g., positive, negative, neutral). The analysis result is passed to the next processing step.
[0495] Step 2:
[0496] Data collection.
[0497] The server uses web scraping technology to collect the necessary data from advertiser and affiliate sites. The collected data includes product details, reviews, past posting data, etc. This data is pre-processed in the next step.
[0498] Step 3:
[0499] Data preprocessing.
[0500] The server preprocesses the collected data. Specifically, it uses Python's BeautifulSoup and Requests to remove HTML tags and regular expressions to remove unnecessary characters. It also tokenizes the text and extracts important key phrases. The preprocessed data is used to train generative artificial intelligence models.
[0501] Step 4:
[0502] Training generative artificial intelligence models.
[0503] The server uses the preprocessed data to train a generative artificial intelligence model (e.g., GPT-2). During this training, the model learns the style and tone of the advertiser's website and affiliate site. The trained model is then used in the next step.
[0504] Step 5:
[0505] Generating advertising articles.
[0506] The server uses the trained generative artificial intelligence model to generate advertising articles based on the user's emotional state. For example, if the user is in a positive emotional state, an article with a positive tone is generated. The generated advertising article is then processed in the following steps.
[0507] Step 6:
[0508] Inserting advertising labels and annotations.
[0509] The server automatically inserts appropriate advertisements and annotations into the generated articles by training a generative AI model with standard annotations based on legal requirements, resulting in articles that mitigate legal risks.
[0510] Step 7:
[0511] Internal check.
[0512] The generated articles are then subjected to an internal check by the server, where a quality check is carried out again using the generative AI model to check for errors, misrepresentations, and missing notes. If necessary, an in-house checker manually corrects the errors.
[0513] Step 8:
[0514] Publication of article.
[0515] The device publishes articles on the website after internal checks are complete. Publication is done automatically based on a pre-set schedule and a publication log is recorded, streamlining the article publishing process.
[0516] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0517] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0518] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0519] [Second embodiment]
[0520] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0521] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0522] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0523] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0524] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0525] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0526] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0527] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0528] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0529] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0530] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0531] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0532] The following description is provided as an example of an embodiment of the present invention: The system collects data from advertisers and affiliate sites and uses a generative artificial intelligence model (e.g., ChatGPT) to generate highly accurate affiliate articles.
[0533] System Configuration
[0534] The system mainly consists of the following components:
[0535] 1. Data Collection Module
[0536] 2. Data Preprocessing Module
[0537] 3. Generative AI Model Training Module
[0538] 4. Style and Tone Extraction Module
[0539] 5. Annotation Learning Module
[0540] 6. Article Generation Module
[0541] 7. Internal Check Module
[0542] 8. Article Publishing Module
[0543] Data collection
[0544] First, the server collects the necessary data from the advertiser and affiliate sites. The collected data includes detailed product information, review articles, and past posting data from the affiliate sites. The server obtains the data from each site using web crawling technology.
[0545] Data Preprocessing
[0546] Next, the server preprocesses the collected data. This step involves data cleaning to remove unnecessary information, such as removing HTML tags, deleting unnecessary strings using regular expressions, tokenizing the text, and extracting important key phrases and features.
[0547] Training generative artificial intelligence models
[0548] Using the preprocessed data, the server trains a generative AI model (ChatGPT), which fine-tunes the model based on the initial dataset to learn specific writing styles and expressions.
[0549] Style and Tone Extraction
[0550] The server extracts the style and tone of the advertiser's and affiliate's sites and trains the generative AI model. During this process, the server analyzes each site's writing style, tone, and punctuation usage. This data is then fed back into the generative model to generate articles that match the characteristics of each site.
[0551] Annotation Learning
[0552] The server trains a generative AI model to create appropriate ad descriptions and annotations. This involves collecting standard annotations based on legal requirements and teaching the model necessary annotation information, such as "This product is not a medicine."
[0553] Article Generation
[0554] The server automatically generates affiliate articles based on the collected and learned data. For example, when generating an article about "Healthcare Product A," the generative AI model emphasizes the product's effects and features and generates text that matches the tone and style. Annotations based on legal requirements are also automatically inserted into the article.
[0555] Internal Check
[0556] The generated articles are then checked internally by a user (an internal checker). A checklist is used to check for typographical errors, misrepresentations, and missing notes. During this process, the quality can be checked again using the generative AI model. If necessary, the user can make manual corrections.
[0557] Article Publication
[0558] Finally, articles that have passed the check are published on the website via the terminal, which automatically posts the articles based on the publication schedule and records the publication log.
[0559] Specific examples
[0560] For example, generating an affiliate article for healthcare product A involves the following process:
[0561] 1. The server collects information about the ingredients, efficacy, and reviews of healthcare product A.
[0562] 2. The server preprocesses the collected data, removing HTML tags, deleting unnecessary strings, and tokenizing the text.
[0563] 3. The server trains the ChatGPT model to learn specific writing styles and expressions.
[0564] 4. The server extracts the advertiser's formal tone and the affiliate site's friendly tone and trains the model accordingly.
[0565] 5. The server learns the advertising and annotations and applies them to the generated articles.
[0566] 6. The server generates an article such as, "Healthcare product A contains ingredient B and is expected to be effective in improving immunity."
[0567] 7. The user checks the generated article for errors, misrepresentations, and missing notes, and makes corrections as necessary.
[0568] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[0569] The above system allows high-quality affiliate articles to be generated efficiently and published with reduced legal risks.
[0570] The processing flow will be explained below.
[0571] Step 1:
[0572] The server collects data from advertiser and affiliate sites. Specifically, it uses a crawler to retrieve the HTML of web pages from the specified URLs and stores it in a database. The retrieved data includes titles, body text, meta tags, image URLs, etc.
[0573] Step 2:
[0574] The server preprocesses the collected data, removing HTML tags and extracting only the text data. Regular expressions are also used to remove unnecessary symbols and spaces. Next, the text is tokenized and techniques such as TF-IDF and Word2Vec are used to extract important key phrases and features.
[0575] Step 3:
[0576] The server uses the preprocessed data to train a generative AI model, based on the initial ChatGPT model, which is then fine-tuned to fit specific writing styles and themes. This process uses high-quality samples from the collected data as training data.
[0577] Step 4:
[0578] The server extracts the style and tone of the advertiser and affiliate sites, uses analysis algorithms to identify the unique writing style and tone patterns of each site, and feeds these features back into the generative AI model.
[0579] Step 5:
[0580] The server trains a generative AI model to create appropriate annotations. Standard annotations and warnings based on legal requirements are collected and used as training data for the model. This training gives the model the ability to automatically insert appropriate annotations when generating articles.
[0581] Step 6:
[0582] The server automatically generates affiliate articles based on the data it has collected and learned. For example, it generates articles that emphasize the effectiveness and features of specific products and services using information about those products and services. The generated articles also include necessary annotations.
[0583] Step 7:
[0584] User-generated articles are internally checked using a dedicated checklist to ensure there are no errors, misrepresentations, or missing notes. This process can also be repeated using a generative AI model to check quality.
[0585] Step 8:
[0586] The device publishes the checked articles on the website. Articles are automatically posted according to the publication schedule. The URL and related information after publication are recorded in a database and stored as data for analysis.
[0587] Through the above steps, high-quality affiliate articles can be efficiently generated and published with reduced legal risks.
[0588] Example 1
[0589] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0590] In conventional affiliate article generation systems, a lot of manual work is required to efficiently generate a large amount of content while improving the quality of the articles, which leads to issues such as inconsistency in the articles and inconsistency in the quality of the articles.In addition, inserting accurate annotations based on legal requirements and internal article checking processes require a great deal of effort.
[0591] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0592] In this invention, the server includes means for preprocessing the collected data and training the generative AI model, means for extracting style and tone from advertiser information sources and websites and training the generative AI model, and means for automatically generating text based on the collected and learned data, thereby enabling the efficient generation of high-quality affiliate articles and the provision of content with a consistent style and tone.
[0593] "Collected data" refers to detailed product information, review articles, past posting data, etc. obtained from advertisers and affiliate sites.
[0594] "Preprocessing" refers to a series of processes that remove HTML tags and unnecessary strings from collected data, tokenize the text, and extract important key phrases and features.
[0595] A "generative artificial intelligence model" refers to an artificial intelligence that learns a specific writing style or expression method based on an initial dataset.
[0596] "Training methods" refers to the use of collected data to fine-tune generative AI models to learn specific writing styles and idiomatic expressions.
[0597] "Advertiser Source" refers to a website or database that contains information about an advertiser.
[0598] "Style and tone extraction" refers to the process of analyzing each site's writing style, tone, and punctuation usage and training the model.
[0599] "Means for automatic generation" refers to a method of automatically generating sentences based on collected and learned data using a generative artificial intelligence model.
[0600] "Means for appropriately inserting annotations" refers to a method for automatically inserting standard legal annotations into the generated text.
[0601] "Means for internal checking" refers to a method in which an in-house checker reviews the generated text to check for any typos, misrepresentations, or missing notes.
[0602] "Means of disclosing to the information source" refers to a method in which the generated text is automatically posted on a website after being verified and a disclosure log is recorded.
[0603] "Crawling" refers to the act of collecting data from a website using an automated crawler program.
[0604] "HTML tag stripping" refers to the process of removing unnecessary HTML code from text data obtained from a web page.
[0605] "Extracting text information" refers to the method of extracting the actual required text parts from the preprocessed data.
[0606] "Methods for checking for errors, misrepresentations, and missing notes" refers to methods for checking the content of generated articles and checking for typos, inaccurate expressions, and missing necessary notes.
[0607] This invention relates to a system that collects data from advertisers and affiliate sites and generates highly accurate affiliate articles using a generative artificial intelligence model (e.g., a generative AI model). The system automatically inserts annotations based on legal requirements into the generated articles, and then publishes them on a website after internal checks.
[0608] Hardware and Software
[0609] The system requires the following hardware and software:
[0610] Server: Collects data, preprocesses it, trains the generative AI model, generates articles, and publishes them.
[0611] Terminal: Serves as an interface for publishing checked articles on a website.
[0612] User: Internal checks and corrections of generated articles.
[0613] The main software used includes:
[0614] Web crawler tools: Used to collect data from advertiser and affiliate sites.
[0615] Data preprocessing library: Used to remove HTML tags, tokenize text, and extract key phrases.
[0616] Generative artificial intelligence model (generative AI model): Refers to a generative artificial intelligence platform used to generate articles, such as a generative AI model or ChatGPT.
[0617] Linguistic Style Analysis Library: Used to extract style and tone.
[0618] Text review tools: Used for internal checks and to identify errors, misstatements, and missing notes.
[0619] Specific processing and calculation of data
[0620] 1. Data Collection:
[0621] The server uses a web crawler tool to automatically retrieve product details, reviews, and past posting data from advertiser and affiliate sites, and stores this data in an internal database.
[0622] 2. Data Preprocessing:
[0623] The server performs preprocessing on the collected data, such as removing HTML tags, deleting unnecessary text using regular expressions, tokenizing, and extracting important key phrases.
[0624] 3. Training the generative AI model:
[0625] The server uses the preprocessed data to train a generative artificial intelligence model, a process that customizes the model by teaching it specific writing styles and idioms.
[0626] 4. Extracting style and tone:
[0627] The server uses a language style analysis library to analyze the style and tone of advertiser and affiliate sites, and the analysis data is fed back to a generative AI model.
[0628] 5. Annotation Learning:
[0629] The server collects annotations based on legal requirements and trains a generative artificial intelligence model based on them.
[0630] 6. Article Generation:
[0631] The server uses a generative artificial intelligence model to generate affiliate articles based on the collected and learned data, and the generated articles are automatically supplemented with the necessary legal annotations.
[0632] 7. Internal check:
[0633] The generated articles are then internally checked by users who use text review tools to identify errors, misstatements, and omissions, and make manual corrections as needed.
[0634] 8. Article Publication:
[0635] Finally, articles that have passed the check are published on the website via the terminal, which posts the articles based on the publication schedule and records the publication log.
[0636] Examples of concrete examples and prompts
[0637] For example, let's say you want to generate an affiliate article for "Healthcare Product A":
[0638] 1. The server collects information about the ingredients, efficacy, and reviews of healthcare product A.
[0639] 2. The server preprocesses the collected data and removes HTML tags and unnecessary strings.
[0640] 3. The server trains the generative AI model to learn specific writing styles and expressions.
[0641] 4. The server extracts the advertiser's formal tone and the affiliate site's friendly tone and trains the model.
[0642] 5. The server learns the advertising notations and annotations and applies them to the generated article.
[0643] 6. The server generates an article stating, "Healthcare product A contains ingredient B and is expected to be effective in improving immunity."
[0644] 7. The user checks the generated article for errors and missing notes.
[0645] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[0646] Example prompt sentence:
[0647] "Use a generative AI model to create an affiliate ad article about healthcare product A based on its features, ingredients, and user reviews. Focus on its immune-boosting benefits. Also, be sure to include legal annotations."
[0648] With the above configuration, the present invention makes it possible to efficiently generate high-quality affiliate articles, reduce legal risks, and reliably publish them.
[0649] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0650] Step 1:
[0651] Data collection
[0652] The server collects the necessary data from advertisers and affiliate sites.
[0653] Input: URL list, crawling settings
[0654] Processing: Use a web crawler tool to retrieve detailed product information, reviews, and past posting data from the specified URL.
[0655] Output: Collected data (HTML file, text data, etc.)
[0656] What it does: The server runs a crawler tool that parses the HTML content of each website to extract data and store it in an internal database.
[0657] Step 2:
[0658] Data Preprocessing
[0659] The server preprocesses the collected data.
[0660] Input: Collected data (HTML files, text data, etc.)
[0661] Processing: Strip HTML tags, remove unnecessary strings, tokenize text, and extract important key phrases.
[0662] Output: Preprocessed text data
[0663] What it does: It uses a regular expression library to remove HTML tags and unnecessary text, a morphological analyzer to break down text into words, and a keyphrase extraction algorithm to extract important keywords.
[0664] Step 3:
[0665] Training generative artificial intelligence models
[0666] The server uses the preprocessed data to train a generative artificial intelligence model (ChatGPT).
[0667] Input: Preprocessed text data
[0668] Processing: Fine-tuning the model to learn specific writing styles and phrasing.
[0669] Output: A trained generative artificial intelligence model
[0670] What it does: It processes batches of data, feeds it into the model, and fine-tunes it by setting optimal hyperparameters.
[0671] Step 4:
[0672] Style and Tone Extraction
[0673] The server extracts the style and tone of the advertiser and affiliate sites.
[0674] Input: Text data from advertiser and affiliate sites
[0675] Processing: Uses a linguistic style analysis library to analyze writing style, tone, and punctuation usage.
[0676] Output: Extracted style and tone data
[0677] Specific operations: Analyzes text data from each site, extracts elements that characterize writing style and tone, and feeds this back to a generative AI model.
[0678] Step 5:
[0679] Annotation Learning
[0680] The server trains a generative artificial intelligence model to generate annotations based on legal requirements.
[0681] Input: Legal annotation database
[0682] Processing: Train the model with standard annotations based on legal requirements.
[0683] Output: A generative AI model that can apply annotations
[0684] What it does: Legal annotations are fed into the model as a dataset, and the model is trained to accurately insert annotations under specific conditions.
[0685] Step 6:
[0686] Article Generation
[0687] The server generates affiliate articles using a generative artificial intelligence model.
[0688] Input: Trained generative artificial intelligence model, collected and preprocessed data
[0689] Processing: Using data, we automatically generate affiliate articles that match your tone and style.
[0690] Output: Generated affiliate article
[0691] What it does: The model generates sentences based on the prompts you provide and automatically adds any necessary legal annotations.
[0692] Step 7:
[0693] Internal Check
[0694] Internally check user-generated articles.
[0695] Input: Generated affiliate article
[0696] Action: Check for clerical errors, misrepresentations, and missing notes.
[0697] Output: Affiliate articles checked or corrected
[0698] What you'll do: Proofread articles using text review tools, conduct team reviews, and manually correct as needed.
[0699] Step 8:
[0700] Article published
[0701] The device publishes the checked articles on the website.
[0702] Input: Checked affiliate article
[0703] Processing: Post articles based on the publishing schedule and record publishing logs.
[0704] Output: Published affiliate articles
[0705] What it does: Uses a scheduler tool to automatically post articles at specified times and records the publishing history.
[0706] (Application example 1)
[0707] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0708] Conventional ad article generation systems have problems with inefficiency and time because the data collection and preprocessing process is complicated and style and tone adjustments are often done manually. Other issues include difficulty in previewing and editing generated articles, and difficulty in automatically posting articles to content management systems. Furthermore, there is a lack of a mechanism for quickly generating articles using smartphones.
[0709] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0710] In this invention, the server includes means for preprocessing collected data and training a generative AI model, means for extracting style and tone from the advertiser's website and affiliated sites and training the generative AI model, means for automatically generating advertising articles based on the collected and learned data, means for appropriately inserting advertising notices and annotations into the generated articles, means for internally checking the generated articles, means for publishing the generated articles on the website after the internal check, means for quickly and easily generating articles using a smartphone, means for displaying the generated articles during preview and editing and enabling editing, and means for automatically posting the generated articles via the API of a content management system, thereby enabling efficient and highly accurate generation and publication of advertising articles.
[0711] "Collected data" refers to information such as product information, reviews, and past posting data obtained from advertisers and affiliated sites.
[0712] "Preprocessing" is the process of removing unnecessary information from collected data and shaping the data.
[0713] A "generative artificial intelligence model" is an artificial intelligence algorithm that can generate new text based on input data.
[0714] "Training" is the learning process used to improve the accuracy of a generative artificial intelligence model based on specific data.
[0715] "Style" refers to characteristics such as writing style, tone, and format.
[0716] "Tone" describes the mood or emotional atmosphere of a piece of writing.
[0717] An "advertising piece" is a piece of writing written to promote a particular product or service.
[0718] "Advertising notation" refers to text or notes that clearly indicate that the item is an advertisement.
[0719] An "annotation" is an explanatory text based on supplementary information or legal requirements inserted within an article.
[0720] "Internal checking" is the process of checking the content of generated articles and correcting errors and inappropriate expressions.
[0721] "Publishing on a website" means posting the generated article on the web so that it can be viewed by general internet users.
[0722] "Using a smartphone" means performing various operations and processes using a smartphone.
[0723] "Preview" is a function that allows you to check the state before the final release.
[0724] "Editing" is the process of correcting or changing the content of the generated article.
[0725] A "content management system" is software for organizing, managing, and delivering content.
[0726] An "API" is an interface that allows applications and services to communicate with each other and utilize their functionality.
[0727] "Automatically posting" means that the system automatically uploads content to a website based on a set schedule or conditions.
[0728] As an embodiment of the present invention, a novel advertising article generation system is described, which preprocesses collected data, trains a generative artificial intelligence model, and ultimately automatically generates and publishes highly accurate advertising articles.
[0729] System Configuration
[0730] The system of the present invention includes the following major components:
[0731] 1. Data Collection Module
[0732] 2. Data Preprocessing Module
[0733] 3. Generative AI Model Training Module
[0734] 4. Style and Tone Extraction Module
[0735] 5. Annotation Learning Module
[0736] 6. Article Generation Module
[0737] 7. Internal Check Module
[0738] 8. Article Publishing Module
[0739] 9. Smartphone Application Module
[0740] 10. Preview and Edit Module
[0741] 11. Auto-posting module
[0742] Hardware and Software Used
[0743] The main hardware and software components of this system include:
[0744] Hardware: High-performance servers and smartphones.
[0745] Software: Python, BeautifulSoup, Scrapy, pandas, transformers, React Native, Content Management Systems (CMS) WordPress and Wix, OpenAI's ChatGPT.
[0746] Processing Description
[0747] 1. Data Collection
[0748] The server collects the necessary data from advertisers and partner sites using APIs and web crawling technologies (BeautifulSoup and Scrapy).
[0749] 2. Data Preprocessing
[0750] The collected data is cleaned using a data preprocessing module. Specifically, HTML tags are removed, unnecessary strings are deleted using regular expressions, and important key phrases and features are extracted. The data is formatted and processed using the pandas library.
[0751] 3. Generative AI Model Training
[0752] The server uses the preprocessed data to train a generative AI model (ChatGPT), which learns specific writing styles and expressions.
[0753] 4. Extracting Style and Tone
[0754] The server extracts style and tone from the text of advertisers and partner sites and feeds this back into a generative artificial intelligence model.
[0755] 5. Annotation Learning
[0756] The server trains the model with annotations based on legal requirements and applies them to the generated advertising articles.
[0757] 6. Article Generation
[0758] Using a generative artificial intelligence model, advertising articles are automatically generated based on collected and learned data.
[0759] 7. Internal Checks
[0760] The generated articles are then internally checked by the user to ensure there are no errors, misrepresentations, or missing notes, and any necessary corrections are made. The quality can also be checked again using the generative AI model.
[0761] 8. Public
[0762] Once the article has been checked, it will be published on the website via the terminal.
[0763] 9. Smartphone use
[0764] Users can quickly and easily create, preview, edit and publish articles using their smartphones.
[0765] 10. Preview and Edit
[0766] It displays a preview of the generated article and allows users to easily make any necessary corrections. It is provided with an interface that uses React Native.
[0767] 11. Auto-posting
[0768] Finally, the generated articles are automatically posted to the website via the content management system's API.
[0769] Specific examples
[0770] For example, generating an advertisement for a health device A involves the following process:
[0771] 1. The server collects information about the ingredients, efficacy, and reviews of health device A.
[0772] 2. The server preprocesses the collected data, removing HTML tags, deleting unnecessary strings, and tokenizing the text.
[0773] 3. The server trains the ChatGPT model to learn specific writing styles and expressions.
[0774] 4. The server extracts the advertiser's formal tone and the partner site's friendly tone and trains the model accordingly.
[0775] 5. The server learns the advertising and annotations and applies them to the generated articles.
[0776] 6. The server generates an article such as "Health device A contains ingredient B and is expected to be effective in improving immunity."
[0777] 7. The user checks the generated article for errors, misrepresentations, and missing notes, and makes corrections as necessary.
[0778] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[0779] Prompt Sentence Examples
[0780] Product name: Health equipment A
[0781] Features: Quiet design, foldable, multi-function display
[0782] As described above, the present invention enables efficient and highly accurate generation and publication of advertising articles.
[0783] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0784] Step 1:
[0785] Data collection
[0786] The server uses APIs and web crawling technologies (BeautifulSoup and Scrapy) to retrieve product details, reviews, and past postings from advertisers and partner sites. The input is the URL or API endpoint of each site, and the output is the retrieved raw data. This allows a wealth of information to be stored in the database used.
[0787] Step 2:
[0788] Data Preprocessing
[0789] The server preprocesses the collected data. Specifically, it uses the pandas library to remove HTML tags, delete unnecessary strings, and extract important key phrases and features. The input is the raw data collected in step 1, and the output is cleansed and formatted data. This process processes the data into a form that is easy for the generative artificial intelligence model to process.
[0790] Step 3:
[0791] Generative AI model training
[0792] The server uses the preprocessed data to train a generative AI model (ChatGPT). In the process, it learns specific writing styles and expressions. Specifically, it fine-tunes the model using the transformers library. The input is the preprocessed data, and the output is a trained generative AI model. This allows the model to generate sentences with even greater accuracy.
[0793] Step 4:
[0794] Style and Tone Extraction
[0795] The server performs text analysis to extract style and tone from the text of advertisers' and partner sites. Specifically, it uses machine learning algorithms and natural language processing (NLP) technology to extract text characteristics. The input is text data from advertisers' and partner sites, and the output is the extracted style and tone features. This ensures that the generated articles are tailored to the characteristics of each site.
[0796] Step 5:
[0797] Annotation Learning
[0798] The server collects specific texts (legal annotations) and trains a generative AI model to learn annotations based on legal requirements. Specifically, it identifies patterns in the annotations and incorporates them into the model. The input is a dataset of legal annotations, and the output is a generative AI model that can insert annotations appropriately. This results in articles that reduce legal risk.
[0799] Step 6:
[0800] Article Generation
[0801] The server uses a generative artificial intelligence model to automatically generate advertising articles based on collected and learned data. Specifically, when a prompt is entered, the corresponding trained model generates an article. The input is the prompt and past posting data, and the output is an automatically generated advertising article. This enables fast and highly accurate article generation.
[0802] Example prompt sentence:
[0803] Product name: Health equipment A
[0804] Features: Quiet design, foldable, multi-function display
[0805] Step 7:
[0806] Internal Check
[0807] The generated article undergoes an internal check by the user. Specifically, it checks for typographical errors, misrepresentations, and missing notes, and makes any necessary corrections. The input is the automatically generated advertising article, and the output is the corrected final version of the article. This allows the article to be published with guaranteed quality.
[0808] Step 8:
[0809] Article published
[0810] Once the article has been checked, it is published on the website via the terminal. Specifically, the article is automatically posted via the API of the content management system (CMS). The input is the final, corrected version of the article, and the output is the published article. This allows the article to be published efficiently on the web, saving the user time and effort.
[0811] Step 9:
[0812] Smartphone use
[0813] Users use their smartphones to create, preview, edit, and publish articles. Specifically, articles are created and edited using a dedicated application. The input is the requirements and prompts specified by the user, and the output is the edited article. This allows article creation and management to be done anywhere, anytime.
[0814] Step 10:
[0815] Preview and Edit
[0816] A preview of the generated article is displayed, allowing the user to easily make any necessary corrections. Specifically, the interface uses React Native, and edits are reflected in real time. The input is the automatically generated article, and the output is the article corrected by the user. This further improves the quality of the final article.
[0817] Step 11:
[0818] Auto-post
[0819] Finally, the generated articles are automatically posted to the website via the content management system's API. The server uploads the articles to the website based on the set schedule and conditions. The input is the checked, final version of the article, and the output is the published article and its publication log. This minimizes user effort and efficiently publishes articles through an automated process.
[0820] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0821] The following description is provided as an embodiment of the present invention. This system collects data from advertisers and affiliate sites and generates highly accurate affiliate articles using a generative artificial intelligence model (e.g., ChatGPT). Furthermore, by combining it with an emotion engine that recognizes user emotions, the system can be customized according to the user's emotions.
[0822] System Configuration
[0823] The system mainly consists of the following components:
[0824] 1. Data Collection Module
[0825] 2. Data Preprocessing Module
[0826] 3. Generative AI Model Training Module
[0827] 4. Style and Tone Extraction Module
[0828] 5. Annotation Learning Module
[0829] 6. Emotion Engine Module
[0830] 7. Article Generation Module
[0831] 8. Internal Check Module
[0832] 9. Article Publishing Module
[0833] Data collection
[0834] First, the server collects the necessary data from the advertiser and affiliate sites. The collected data includes detailed product information, review articles, and past posting data from the affiliate sites. The server obtains the data from each site using web crawling technology.
[0835] Data Preprocessing
[0836] Next, the server preprocesses the collected data. This step involves data cleaning to remove unnecessary information, such as removing HTML tags, deleting unnecessary strings using regular expressions, tokenizing the text, and extracting important key phrases and features.
[0837] Training generative artificial intelligence models
[0838] Using the preprocessed data, the server trains a generative AI model (ChatGPT), which fine-tunes the model based on the initial dataset to learn specific writing styles and expressions.
[0839] Style and Tone Extraction
[0840] The server extracts the style and tone of the advertiser's and affiliate's sites and trains the generative AI model. During this process, the server analyzes each site's writing style, tone, and punctuation usage. This data is then fed back into the generative model to generate articles that match the characteristics of each site.
[0841] Annotation Learning
[0842] The server trains a generative AI model to create appropriate ad descriptions and annotations. This involves collecting standard annotations based on legal requirements and teaching the model necessary annotation information, such as "This product is not a medicine."
[0843] Emotion Engine
[0844] The server uses an emotion engine to analyze the user's emotions in real time. The emotion engine recognizes the user's emotional state based on user input and various sensor data (e.g., facial recognition and voice analysis). Based on the results, the style and tone of the generated articles can be adjusted.
[0845] Article Generation
[0846] The server automatically generates affiliate articles based on the collected and learned data. The generative AI model emphasizes the product's benefits and features and generates text that matches the tone and style. For example, if the user is in a positive emotional state, text with a more optimistic tone will be generated. Annotations based on legal requirements are also automatically inserted into the generated articles.
[0847] Internal Check
[0848] The generated articles are then checked internally by a user (an internal checker). A checklist is used to check for typographical errors, misrepresentations, and missing notes. During this process, the quality can be checked again using the generative AI model. If necessary, the user can make manual corrections.
[0849] Article Publication
[0850] Finally, articles that have passed the check are published on the website via the terminal, which automatically posts the articles based on the publication schedule and records the publication log.
[0851] Specific examples
[0852] For example, generating an affiliate article for healthcare product A involves the following process:
[0853] 1. The server collects information about the ingredients, efficacy, and reviews of healthcare product A.
[0854] 2. The server preprocesses the collected data, removing HTML tags, deleting unnecessary strings, and tokenizing the text.
[0855] 3. The server trains the ChatGPT model to learn specific writing styles and expressions.
[0856] 4. The server extracts the advertiser's formal tone and the affiliate site's friendly tone and trains the model accordingly.
[0857] 5. The server learns the advertising and annotations and applies them to the generated articles.
[0858] 6. The server uses an emotion engine to analyze the user's emotions and adjust the tone of the article accordingly. For example, if the user is excited, the article will be generated with a more enthusiastic tone.
[0859] 7. The user reviews the generated article, checking for errors, misrepresentations, and missing notes, and makes corrections as necessary.
[0860] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[0861] The above system makes it possible to efficiently generate high-quality affiliate articles that respond to user sentiment and publish them while reducing legal risks.
[0862] The processing flow will be explained below.
[0863] Step 1:
[0864] The server collects data from advertiser and affiliate sites. Specifically, it uses a crawler to retrieve the HTML of web pages from the specified URLs and stores it in a database. The retrieved data includes titles, body text, meta tags, image URLs, etc.
[0865] Step 2:
[0866] The server preprocesses the collected data, removing HTML tags and extracting only the text data. Regular expressions are also used to remove unnecessary symbols and spaces. Next, the text is tokenized and techniques such as TF-IDF and Word2Vec are used to extract important key phrases and features.
[0867] Step 3:
[0868] The server uses the preprocessed data to train a generative AI model, based on the initial ChatGPT model, which is then fine-tuned to fit specific writing styles and themes. This process uses high-quality samples from the collected data as training data.
[0869] Step 4:
[0870] The server extracts the style and tone of the advertiser and affiliate sites, uses analysis algorithms to identify the unique writing style and tone patterns of each site, and feeds these features back into the generative AI model.
[0871] Step 5:
[0872] The server trains a generative AI model to create appropriate annotations. Standard annotations and warnings based on legal requirements are collected and used as training data for the model. This training gives the model the ability to automatically insert appropriate annotations when generating articles.
[0873] Step 6:
[0874] The server starts an emotion engine that analyzes the user's emotions in real time. The emotion engine detects emotions from the user's input (text, voice, facial recognition data, etc.) and records the results in a database.
[0875] Step 7:
[0876] The server adjusts the style and tone of the generated article based on the user's emotional data analyzed by the emotion engine. For example, if the user is in a positive emotional state, the server generates an optimistic tone of text.
[0877] Step 8:
[0878] Affiliate articles are automatically generated based on the data collected and learned by the server. As a specific example, when generating an article for "Healthcare Product A," the article is generated by incorporating product features and user reviews and adjusting the tone.
[0879] Step 9:
[0880] User-generated articles are internally checked using a dedicated checklist to ensure there are no errors, misrepresentations, or missing notes. Again, during this process, the generative AI model can be used to check the quality.
[0881] Step 10:
[0882] The device publishes the checked articles on the website. Articles are automatically posted according to the publication schedule. The URL and related information after publication are recorded in a database and stored as data for analysis.
[0883] Through the above steps, high-quality affiliate articles that respond to user sentiment can be efficiently generated and published with reduced legal risks.
[0884] Example 2
[0885] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0886] Conventional affiliate article generation systems have had issues with preprocessing collected data, matching the style and tone of generated articles, appropriately inserting advertisements and annotations, and customizing articles to reflect user sentiment. In particular, generating content based on user sentiment is difficult, and generated articles often do not reflect the user's sentiment. There is also a need for more efficient internal checks of generated articles.
[0887] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0888] In this invention, the server includes means for preprocessing collected data and training a generative artificial intelligence model, means for extracting style and tone from the advertiser's site and affiliate site and training the generative artificial intelligence model, and means for the generative artificial intelligence model to analyze user emotions and adjust the style and tone of articles. This enables the generation of more personalized affiliate articles in response to user emotions. Furthermore, using efficiently preprocessed data allows for the efficient generation of high-quality articles, contributing to the efficiency of internal checks.
[0889] "Collected data" refers to information such as detailed product information, review articles, and past posting data obtained from advertiser and affiliate sites using web crawling technology.
[0890] "Preprocessing" refers to a series of processes performed on collected data, such as removing HTML tags, deleting unnecessary strings, tokenizing text, and extracting key phrases and features.
[0891] A "generative artificial intelligence model" is an artificial intelligence algorithm that generates sentences based on training data, and a specific example is ChatGPT.
[0892] "Training" is the process of using preprocessed data to teach a generative artificial intelligence model a particular style or method of expression.
[0893] "Advertiser's Site and Affiliate Site" refers to a website that provides detailed product information and review articles, and includes affiliate marketing content.
[0894] "Style and tone" refers to the writing style, tone, punctuation, and other characteristics of a particular site or document.
[0895] "Learning" is the process by which a generative AI model analyzes data to acquire a particular style or tone, and then applies the results to the model.
[0896] "User emotion" refers to the emotional state analyzed based on the user's input information (text, voice, facial image, etc.), and includes positive, negative, excited, etc.
[0897] "Internal checking" is a process in which the generated articles are checked for errors, misrepresentations, missing notes, etc., and is carried out by in-house checkers.
[0898] "Advertising statements and annotations" refers to standard legally required annotations and advertising statements, such as warnings such as "This product is not a medicine."
[0899] "Publishing on a website" means posting the checked article on a website based on a specific publishing schedule and recording a publishing log.
[0900] To implement this invention, the following system must be constructed. This system uses data collected from advertisers and affiliate sites and a generative artificial intelligence model to automatically generate highly accurate affiliate articles. Furthermore, it is possible to analyze user sentiment and adjust the style and tone of the resulting articles.
[0901] System configuration
[0902] The system consists of the following main components:
[0903] 1. Data Collection Module
[0904] 2. Data Preprocessing Module
[0905] 3. Generative AI Model Training Module
[0906] 4. Style and Tone Extraction Module
[0907] 5. Annotation Learning Module
[0908] 6. Emotion Engine Module
[0909] 7. Article Generation Module
[0910] 8. Internal Check Module
[0911] 9. Article Publishing Module
[0912] Hardware and software used
[0913] The implementation of this system uses the following hardware and software:
[0914] Server: A server with high-performance processing power is used to process collected data, train generative AI models, analyze emotions, and perform other heavy workloads.
[0915] Generative AI models, such as ChatGPT, are used to generate text and adjust style and tone.
[0916] Emotion engine: An engine for analyzing user emotions, using sensors such as facial recognition and voice analysis.
[0917] Program processing
[0918] The server first collects data from advertiser and affiliate sites, including product information, reviews, and past posting data, and the collected data is obtained using web crawling technology.
[0919] The server then preprocesses the collected data, removing HTML tags, eliminating unnecessary strings, tokenizing the text, and extracting important key phrases and features.
[0920] Based on the preprocessed data, the server trains a generative AI model (ChatGPT), fine-tuning the model based on the initial dataset to learn specific writing styles and expressions.
[0921] The server then extracts the style and tone of the advertiser and affiliate sites and trains them into a generative AI model. Each site's writing style, tone, and punctuation usage are analyzed and fed back to the model.
[0922] The server also trains a generative AI model to create appropriate ad labeling and annotations, collecting standard annotations based on legal requirements and training the model to apply them to the generated articles.
[0923] Additionally, the server uses an emotion engine to analyze users' emotions in real time, recognizing their emotional state based on user input and various sensor data (facial recognition and voice analysis), and adjusting the style and tone of the generated articles accordingly.
[0924] As a concrete example, the prompt text to automatically generate an affiliate article for healthcare product A is as follows:
[0925] "Generate an article with an optimistic tone based on the ingredients, efficacy, and reviews of healthcare product A when the user is in a positive emotional state."
[0926] Finally, the generated article is checked internally by the user (an in-house checker). A checklist is used to check for typographical errors, misrepresentations, and missing notes, and manual corrections are made as necessary. The completed article is published on the website via the terminal, and a publication log is recorded.
[0927] summary
[0928] This system allows us to efficiently generate high-quality affiliate articles that reflect user sentiment and publish them while reducing legal risks.
[0929] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0930] Step 1:
[0931] Data collection
[0932] The server collects data from advertiser and affiliate sites. It uses web crawling technology to obtain HTML content based on the URL list specified as input. Specifically, it collects detailed product information, review articles, and past posting data from advertiser and affiliate sites. The output is raw data, which can be in a wide variety of formats.
[0933] Step 2:
[0934] Data Preprocessing
[0935] The data collected by the server is preprocessed. The input is the raw data collected in step 1. Specifically, regular expressions are used to remove HTML tags and extract only the text. Unnecessary strings (e.g., advertising banner text) are also filtered. The text is then tokenized and split into words. Important key phrases are then extracted using algorithms such as TF-IDF and Word2Vec. The output is preprocessed, clean text data.
[0936] Step 3:
[0937] Training generative artificial intelligence models
[0938] The server uses the preprocessed data to train a generative AI model (ChatGPT). The input is preprocessed, clean text data. Specifically, the initial dataset is fed into the generative AI model, and fine-tuned to learn specific writing styles and expressions. The output is a trained generative AI model.
[0939] Step 4:
[0940] Style and Tone Extraction
[0941] The server analyzes the style and tone of advertiser and affiliate sites and trains a generative AI model. The input is the text data of each site. Specific processing involves analyzing writing style (frequency of use of punctuation at the end of sentences, use of honorifics, etc.) and tone (formal or informal, optimistic or pessimistic, etc.). The results of these analyses are then fed back to the generative AI model, enabling it to generate text that matches the characteristics of each site. The output is a model that has learned the characteristics of style and tone.
[0942] Step 5:
[0943] Annotation Learning
[0944] The server trains a generative AI model on appropriate methods for advertising labeling and annotations. The input is standard annotations based on legal requirements (e.g., "This product is not a medicine"). Specific processing involves collecting standard annotations and training the generative AI model so that they can be applied when needed. The output is a model that can insert annotations.
[0945] Step 6:
[0946] Emotion Engine
[0947] The server uses an emotion engine to analyze the user's emotions in real time. The input is the user's input information (text, voice, facial image, etc.). Specifically, the process uses sensors such as facial recognition and voice analysis to recognize the user's emotional state (positive, negative, excited, etc.). Based on the results, the style and tone of the generated article are adjusted. The output is text whose tone and style have been adjusted to reflect the user's emotional state.
[0948] Step 7:
[0949] Article Generation
[0950] The server automatically generates affiliate articles based on collected and learned data. The input is a model that has learned style and tone characteristics that reflect emotional states, as well as requirements such as product information. A generative AI model is used to generate text that highlights the product's features and benefits. Annotations based on legal requirements are also automatically inserted. The output is the completed affiliate article.
[0951] Step 8:
[0952] Internal Check
[0953] The user (an internal reviewer) checks the generated article. The input is the generated affiliate article. Specifically, a checklist is used to check for errors, misrepresentations, and missing notes. Manual corrections may be made as necessary. The output is the final version of the article after corrections have been completed.
[0954] Step 9:
[0955] Article Publication
[0956] The terminal publishes the checked article on the website. The input is the final version of the article. Specific processing involves setting a publishing schedule and automatically posting the article based on the specified date and time. A publishing log is also recorded. The output is the published article and its publishing log.
[0957] (Application example 2)
[0958] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0959] Conventional affiliate article generation systems have difficulty generating advertising content that takes user emotions into account, making it impossible to provide ads that are adapted to user emotions. Furthermore, it is difficult to properly reflect the style and tone of the advertiser or affiliate site, making it impossible to provide an effective advertising experience for users.
[0960] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for preprocessing collected data and training a generative AI model, means for extracting style and tone from the advertiser's website and affiliate site and training the generative AI model, means for automatically generating affiliate articles based on the collected and learned data, means for appropriately inserting advertising notices and annotations into the generated articles, means for internally checking the generated articles, means for publishing the articles generated after the internal check on the website, and means for recognizing user emotions and adjusting the tone and style of the generated advertising content based on the emotions. This makes it possible to provide effective advertising content that corresponds to the user's emotions and generate and publish high-quality affiliate articles while reflecting the style and tone of the advertiser and affiliate site.
[0961] "Collected data" refers to information such as detailed product information, review articles, and past posting data obtained from advertisers and affiliate sites.
[0962] "Preprocessing" refers to the process of cleaning the collected data, removing unnecessary information, and extracting important key phrases and features using regular expressions.
[0963] A "generative artificial intelligence model" is an artificial intelligence model that learns specific writing styles and methods of expression based on given data, and generates new sentences based on the results of that learning.
[0964] "Training" refers to the learning process of using preprocessed data to improve the performance of a generative artificial intelligence model.
[0965] "Style and tone extraction" refers to the process of analyzing characteristics such as writing style and punctuation usage from advertiser websites and affiliate sites, and training a generative AI model to learn these characteristics.
[0966] An "affiliate article" refers to a piece of writing that provides information about a specific product or service and is intended to encourage users to purchase that product or service.
[0967] "Blank and annotation insertion" refers to the process of automatically adding standard legally required annotations to generated articles.
[0968] "Internal Check" refers to the internal review process to ensure that the content of the generated article is accurate and that there are no misrepresentations or omissions.
[0969] "Publication" refers to making the generated article publicly available on a website or other media.
[0970] "Emotion recognition" refers to the process of analyzing a user's emotional state using user input and sensor data (e.g., facial recognition and voice analysis).
[0971] "Tone and style adjustment" refers to the process of changing the tone and presentation of generated text based on the perceived user sentiment.
[0972] The present invention provides an emotion-based advertising content generation system, which includes means for preprocessing collected data and training a generative AI model, means for extracting style and tone from advertiser and affiliate sites and training the generative AI model, means for automatically generating affiliate articles based on the collected and learned data, means for appropriately inserting advertising notices and annotations into the generated articles, means for internally checking the generated articles, means for publishing the generated articles on a website after the internal check, and means for recognizing user emotions and adjusting the tone and style of the generated advertising content based on the emotions.
[0973] Program processing explanation
[0974] The server uses an emotion recognition module to analyze the user's emotions in real time. This module utilizes facial recognition technology using the smartphone camera and TensorFlow. Specifically, it captures the user's face and analyzes their emotions using a trained emotion recognition model (e.g., CNN model).
[0975] The server then uses web crawling techniques to collect the required data from the advertiser's website and affiliate sites, using Python web scraping libraries (e.g., BeautifulSoup and Requests), and preprocesses the collected data by removing HTML tags and extracting text information using regular expressions.
[0976] Based on the preprocessed data, the server trains a generative AI model (e.g., GPT-2). During this process, the collected data is analyzed for its writing style and tone, and the model is trained to learn from it. In particular, the model extracts features of the advertiser's formal writing style and the affiliate site's friendly tone.
[0977] The generated advertisements are automatically inserted with ad captions and annotations to comply with legal requirements. In this process, a generative model is trained on standard annotations collected in advance, and the resulting advertisement captions and annotations are added to the generated articles.
[0978] After generation is complete, the server performs an internal check to ensure there are no errors, misrepresentations, or missing notes. This includes a second quality check using the generative AI model. Once the check is complete, the article is published to the website via the terminal. Here, posts are automatically made based on the publishing schedule, and a publishing log is recorded.
[0979] Examples of concrete examples and prompts
[0980] For example, generating an affiliate article for healthcare product A involves the following process:
[0981] The server uses an emotion recognition module to analyze whether the user is in a positive emotional state and generates an affiliate article with a positive tone. The generated article uses the following prompt:
[0982] "Preprocess this data, normalize it, filter it, tokenize it."
[0983] "Generate ad text based on sentiment and preprocessed data."
[0984] "Please display the generated ad text on your smartphone screen."
[0985] In this way, it becomes possible to provide effective advertising content that matches the user's emotions, and high-quality affiliate articles that reflect the style and tone of the advertiser and affiliate site are generated.
[0986] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0987] Step 1:
[0988] Recognize user emotions.
[0989] The server captures the user's face in real time using the smartphone camera. The captured image is input to an emotion recognition model using TensorFlow to analyze the user's emotional state (e.g., positive, negative, neutral). The analysis result is passed to the next processing step.
[0990] Step 2:
[0991] Data collection.
[0992] The server uses web scraping technology to collect the necessary data from advertiser and affiliate sites. The collected data includes product details, reviews, past posting data, etc. This data is pre-processed in the next step.
[0993] Step 3:
[0994] Data preprocessing.
[0995] The server preprocesses the collected data. Specifically, it uses Python's BeautifulSoup and Requests to remove HTML tags and regular expressions to remove unnecessary characters. It also tokenizes the text and extracts important key phrases. The preprocessed data is used to train generative artificial intelligence models.
[0996] Step 4:
[0997] Training generative artificial intelligence models.
[0998] The server uses the preprocessed data to train a generative artificial intelligence model (e.g., GPT-2). During this training, the model learns the style and tone of the advertiser's website and affiliate site. The trained model is then used in the next step.
[0999] Step 5:
[1000] Generating advertising articles.
[1001] The server uses the trained generative artificial intelligence model to generate advertising articles based on the user's emotional state. For example, if the user is in a positive emotional state, an article with a positive tone is generated. The generated advertising article is then processed in the following steps.
[1002] Step 6:
[1003] Inserting advertising labels and annotations.
[1004] The server automatically inserts appropriate advertisements and annotations into the generated articles by training a generative AI model with standard annotations based on legal requirements, resulting in articles that mitigate legal risks.
[1005] Step 7:
[1006] Internal check.
[1007] The generated articles are then subjected to an internal check by the server, where a quality check is carried out again using the generative AI model to check for errors, misrepresentations, and missing notes. If necessary, an in-house checker manually corrects the errors.
[1008] Step 8:
[1009] Publication of article.
[1010] The device publishes articles on the website after internal checks are complete. Publication is done automatically based on a pre-set schedule and a publication log is recorded, streamlining the article publishing process.
[1011] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1012] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1013] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1014] [Third embodiment]
[1015] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1016] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1017] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1018] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1019] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1020] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1021] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1022] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1023] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1024] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1025] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1026] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1027] The following description is provided as an example of an embodiment of the present invention: The system collects data from advertisers and affiliate sites and uses a generative artificial intelligence model (e.g., ChatGPT) to generate highly accurate affiliate articles.
[1028] System Configuration
[1029] The system mainly consists of the following components:
[1030] 1. Data Collection Module
[1031] 2. Data Preprocessing Module
[1032] 3. Generative AI Model Training Module
[1033] 4. Style and Tone Extraction Module
[1034] 5. Annotation Learning Module
[1035] 6. Article Generation Module
[1036] 7. Internal Check Module
[1037] 8. Article Publishing Module
[1038] Data collection
[1039] First, the server collects the necessary data from the advertiser and affiliate sites. The collected data includes detailed product information, review articles, and past posting data from the affiliate sites. The server obtains the data from each site using web crawling technology.
[1040] Data Preprocessing
[1041] Next, the server preprocesses the collected data. This step involves data cleaning to remove unnecessary information, such as removing HTML tags, deleting unnecessary strings using regular expressions, tokenizing the text, and extracting important key phrases and features.
[1042] Training generative artificial intelligence models
[1043] Using the preprocessed data, the server trains a generative AI model (ChatGPT), which fine-tunes the model based on the initial dataset to learn specific writing styles and expressions.
[1044] Style and Tone Extraction
[1045] The server extracts the style and tone of the advertiser's and affiliate's sites and trains the generative AI model. During this process, the server analyzes each site's writing style, tone, and punctuation usage. This data is then fed back into the generative model to generate articles that match the characteristics of each site.
[1046] Annotation Learning
[1047] The server trains a generative AI model to create appropriate ad descriptions and annotations. This involves collecting standard annotations based on legal requirements and teaching the model necessary annotation information, such as "This product is not a medicine."
[1048] Article Generation
[1049] The server automatically generates affiliate articles based on the collected and learned data. For example, when generating an article about "Healthcare Product A," the generative AI model emphasizes the product's effects and features and generates text that matches the tone and style. Annotations based on legal requirements are also automatically inserted into the article.
[1050] Internal Check
[1051] The generated articles are then checked internally by a user (an internal checker). A checklist is used to check for typographical errors, misrepresentations, and missing notes. During this process, the quality can be checked again using the generative AI model. If necessary, the user can make manual corrections.
[1052] Article Publication
[1053] Finally, articles that have passed the check are published on the website via the terminal, which automatically posts the articles based on the publication schedule and records the publication log.
[1054] Specific examples
[1055] For example, generating an affiliate article for healthcare product A involves the following process:
[1056] 1. The server collects information about the ingredients, efficacy, and reviews of healthcare product A.
[1057] 2. The server preprocesses the collected data, removing HTML tags, deleting unnecessary strings, and tokenizing the text.
[1058] 3. The server trains the ChatGPT model to learn specific writing styles and expressions.
[1059] 4. The server extracts the advertiser's formal tone and the affiliate site's friendly tone and trains the model accordingly.
[1060] 5. The server learns the advertising and annotations and applies them to the generated articles.
[1061] 6. The server generates an article such as, "Healthcare product A contains ingredient B and is expected to be effective in improving immunity."
[1062] 7. The user checks the generated article for errors, misrepresentations, and missing notes, and makes corrections as necessary.
[1063] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[1064] The above system allows high-quality affiliate articles to be generated efficiently and published with reduced legal risks.
[1065] The processing flow will be explained below.
[1066] Step 1:
[1067] The server collects data from advertiser and affiliate sites. Specifically, it uses a crawler to retrieve the HTML of web pages from the specified URLs and stores it in a database. The retrieved data includes titles, body text, meta tags, image URLs, etc.
[1068] Step 2:
[1069] The server preprocesses the collected data, removing HTML tags and extracting only the text data. Regular expressions are also used to remove unnecessary symbols and spaces. Next, the text is tokenized and techniques such as TF-IDF and Word2Vec are used to extract important key phrases and features.
[1070] Step 3:
[1071] The server uses the preprocessed data to train a generative AI model, based on the initial ChatGPT model, which is then fine-tuned to fit specific writing styles and themes. This process uses high-quality samples from the collected data as training data.
[1072] Step 4:
[1073] The server extracts the style and tone of the advertiser and affiliate sites, uses analysis algorithms to identify the unique writing style and tone patterns of each site, and feeds these features back into the generative AI model.
[1074] Step 5:
[1075] The server trains a generative AI model to create appropriate annotations. Standard annotations and warnings based on legal requirements are collected and used as training data for the model. This training gives the model the ability to automatically insert appropriate annotations when generating articles.
[1076] Step 6:
[1077] The server automatically generates affiliate articles based on the data it has collected and learned. For example, it generates articles that emphasize the effectiveness and features of specific products and services using information about those products and services. The generated articles also include necessary annotations.
[1078] Step 7:
[1079] User-generated articles are internally checked using a dedicated checklist to ensure there are no errors, misrepresentations, or missing notes. This process can also be repeated using a generative AI model to check quality.
[1080] Step 8:
[1081] The device publishes the checked articles on the website. Articles are automatically posted according to the publication schedule. The URL and related information after publication are recorded in a database and stored as data for analysis.
[1082] Through the above steps, high-quality affiliate articles can be efficiently generated and published with reduced legal risks.
[1083] Example 1
[1084] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1085] In conventional affiliate article generation systems, a lot of manual work is required to efficiently generate a large amount of content while improving the quality of the articles, which leads to issues such as inconsistency in the articles and inconsistency in the quality of the articles.In addition, inserting accurate annotations based on legal requirements and internal article checking processes require a great deal of effort.
[1086] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1087] In this invention, the server includes means for preprocessing the collected data and training the generative AI model, means for extracting style and tone from advertiser information sources and websites and training the generative AI model, and means for automatically generating text based on the collected and learned data, thereby enabling the efficient generation of high-quality affiliate articles and the provision of content with a consistent style and tone.
[1088] "Collected data" refers to detailed product information, review articles, past posting data, etc. obtained from advertisers and affiliate sites.
[1089] "Preprocessing" refers to a series of processes that remove HTML tags and unnecessary strings from collected data, tokenize the text, and extract important key phrases and features.
[1090] A "generative artificial intelligence model" refers to an artificial intelligence that learns a specific writing style or expression method based on an initial dataset.
[1091] "Training methods" refers to the use of collected data to fine-tune generative AI models to learn specific writing styles and idiomatic expressions.
[1092] "Advertiser Source" refers to a website or database that contains information about an advertiser.
[1093] "Style and tone extraction" refers to the process of analyzing each site's writing style, tone, and punctuation usage and training the model.
[1094] "Means for automatic generation" refers to a method of automatically generating sentences based on collected and learned data using a generative artificial intelligence model.
[1095] "Means for appropriately inserting annotations" refers to a method for automatically inserting standard legal annotations into the generated text.
[1096] "Means for internal checking" refers to a method in which an in-house checker reviews the generated text to check for any typos, misrepresentations, or missing notes.
[1097] "Means of disclosing to the information source" refers to a method in which the generated text is automatically posted on a website after being verified and a disclosure log is recorded.
[1098] "Crawling" refers to the act of collecting data from a website using an automated crawler program.
[1099] "HTML tag stripping" refers to the process of removing unnecessary HTML code from text data obtained from a web page.
[1100] "Extracting text information" refers to the method of extracting the actual required text parts from the preprocessed data.
[1101] "Methods for checking for errors, misrepresentations, and missing notes" refers to methods for checking the content of generated articles and checking for typos, inaccurate expressions, and missing necessary notes.
[1102] This invention relates to a system that collects data from advertisers and affiliate sites and generates highly accurate affiliate articles using a generative artificial intelligence model (e.g., a generative AI model). The system automatically inserts annotations based on legal requirements into the generated articles, and then publishes them on a website after internal checks.
[1103] Hardware and Software
[1104] The system requires the following hardware and software:
[1105] Server: Collects data, preprocesses it, trains the generative AI model, generates articles, and publishes them.
[1106] Terminal: Serves as an interface for publishing checked articles on a website.
[1107] User: Internal checks and corrections of generated articles.
[1108] The main software used includes:
[1109] Web crawler tools: Used to collect data from advertiser and affiliate sites.
[1110] Data preprocessing library: Used to remove HTML tags, tokenize text, and extract key phrases.
[1111] Generative artificial intelligence model (generative AI model): Refers to a generative artificial intelligence platform used to generate articles, such as a generative AI model or ChatGPT.
[1112] Linguistic Style Analysis Library: Used to extract style and tone.
[1113] Text review tools: Used for internal checks and to identify errors, misstatements, and missing notes.
[1114] Specific processing and calculation of data
[1115] 1. Data Collection:
[1116] The server uses a web crawler tool to automatically retrieve product details, reviews, and past posting data from advertiser and affiliate sites, and stores this data in an internal database.
[1117] 2. Data Preprocessing:
[1118] The server performs preprocessing on the collected data, such as removing HTML tags, deleting unnecessary text using regular expressions, tokenizing, and extracting important key phrases.
[1119] 3. Training the generative AI model:
[1120] The server uses the preprocessed data to train a generative artificial intelligence model, a process that customizes the model by teaching it specific writing styles and idioms.
[1121] 4. Extracting style and tone:
[1122] The server uses a language style analysis library to analyze the style and tone of advertiser and affiliate sites, and the analysis data is fed back to a generative AI model.
[1123] 5. Annotation Learning:
[1124] The server collects annotations based on legal requirements and trains a generative artificial intelligence model based on them.
[1125] 6. Article Generation:
[1126] The server uses a generative artificial intelligence model to generate affiliate articles based on the collected and learned data, and the generated articles are automatically supplemented with the necessary legal annotations.
[1127] 7. Internal check:
[1128] The generated articles are then internally checked by users who use text review tools to identify errors, misstatements, and omissions, and make manual corrections as needed.
[1129] 8. Article Publication:
[1130] Finally, articles that have passed the check are published on the website via the terminal, which posts the articles based on the publication schedule and records the publication log.
[1131] Examples of concrete examples and prompts
[1132] For example, let's say you want to generate an affiliate article for "Healthcare Product A":
[1133] 1. The server collects information about the ingredients, efficacy, and reviews of healthcare product A.
[1134] 2. The server preprocesses the collected data and removes HTML tags and unnecessary strings.
[1135] 3. The server trains the generative AI model to learn specific writing styles and expressions.
[1136] 4. The server extracts the advertiser's formal tone and the affiliate site's friendly tone and trains the model.
[1137] 5. The server learns the advertising notations and annotations and applies them to the generated article.
[1138] 6. The server generates an article stating, "Healthcare product A contains ingredient B and is expected to be effective in improving immunity."
[1139] 7. The user checks the generated article for errors and missing notes.
[1140] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[1141] Example prompt sentence:
[1142] "Use a generative AI model to create an affiliate ad article about healthcare product A based on its features, ingredients, and user reviews. Focus on its immune-boosting benefits. Also, be sure to include legal annotations."
[1143] With the above configuration, the present invention makes it possible to efficiently generate high-quality affiliate articles, reduce legal risks, and reliably publish them.
[1144] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1145] Step 1:
[1146] Data collection
[1147] The server collects the necessary data from advertisers and affiliate sites.
[1148] Input: URL list, crawling settings
[1149] Processing: Use a web crawler tool to retrieve detailed product information, reviews, and past posting data from the specified URL.
[1150] Output: Collected data (HTML file, text data, etc.)
[1151] What it does: The server runs a crawler tool that parses the HTML content of each website to extract data and store it in an internal database.
[1152] Step 2:
[1153] Data Preprocessing
[1154] The server preprocesses the collected data.
[1155] Input: Collected data (HTML files, text data, etc.)
[1156] Processing: Strip HTML tags, remove unnecessary strings, tokenize text, and extract important key phrases.
[1157] Output: Preprocessed text data
[1158] What it does: It uses a regular expression library to remove HTML tags and unnecessary text, a morphological analyzer to break down text into words, and a keyphrase extraction algorithm to extract important keywords.
[1159] Step 3:
[1160] Training generative artificial intelligence models
[1161] The server uses the preprocessed data to train a generative artificial intelligence model (ChatGPT).
[1162] Input: Preprocessed text data
[1163] Processing: Fine-tuning the model to learn specific writing styles and phrasing.
[1164] Output: A trained generative artificial intelligence model
[1165] What it does: It processes batches of data, feeds it into the model, and fine-tunes it by setting optimal hyperparameters.
[1166] Step 4:
[1167] Style and Tone Extraction
[1168] The server extracts the style and tone of the advertiser and affiliate sites.
[1169] Input: Text data from advertiser and affiliate sites
[1170] Processing: Uses a linguistic style analysis library to analyze writing style, tone, and punctuation usage.
[1171] Output: Extracted style and tone data
[1172] Specific operations: Analyzes text data from each site, extracts elements that characterize writing style and tone, and feeds this back to a generative AI model.
[1173] Step 5:
[1174] Annotation Learning
[1175] The server trains a generative artificial intelligence model to generate annotations based on legal requirements.
[1176] Input: Legal annotation database
[1177] Processing: Train the model with standard annotations based on legal requirements.
[1178] Output: A generative AI model that can apply annotations
[1179] What it does: Legal annotations are fed into the model as a dataset, and the model is trained to accurately insert annotations under specific conditions.
[1180] Step 6:
[1181] Article Generation
[1182] The server generates affiliate articles using a generative artificial intelligence model.
[1183] Input: Trained generative artificial intelligence model, collected and preprocessed data
[1184] Processing: Using data, we automatically generate affiliate articles that match your tone and style.
[1185] Output: Generated affiliate article
[1186] What it does: The model generates sentences based on the prompts you provide and automatically adds any necessary legal annotations.
[1187] Step 7:
[1188] Internal Check
[1189] Internally check user-generated articles.
[1190] Input: Generated affiliate article
[1191] Action: Check for clerical errors, misrepresentations, and missing notes.
[1192] Output: Affiliate articles checked or corrected
[1193] What you'll do: Proofread articles using text review tools, conduct team reviews, and manually correct as needed.
[1194] Step 8:
[1195] Article published
[1196] The device publishes the checked articles on the website.
[1197] Input: Checked affiliate article
[1198] Processing: Post articles based on the publishing schedule and record publishing logs.
[1199] Output: Published affiliate articles
[1200] What it does: Uses a scheduler tool to automatically post articles at specified times and records the publishing history.
[1201] (Application example 1)
[1202] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1203] Conventional ad article generation systems have problems with inefficiency and time because the data collection and preprocessing process is complicated and style and tone adjustments are often done manually. Other issues include difficulty in previewing and editing generated articles, and difficulty in automatically posting articles to content management systems. Furthermore, there is a lack of a mechanism for quickly generating articles using smartphones.
[1204] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1205] In this invention, the server includes means for preprocessing collected data and training a generative AI model, means for extracting style and tone from the advertiser's website and affiliated sites and training the generative AI model, means for automatically generating advertising articles based on the collected and learned data, means for appropriately inserting advertising notices and annotations into the generated articles, means for internally checking the generated articles, means for publishing the generated articles on the website after the internal check, means for quickly and easily generating articles using a smartphone, means for displaying the generated articles during preview and editing and enabling editing, and means for automatically posting the generated articles via the API of a content management system, thereby enabling efficient and highly accurate generation and publication of advertising articles.
[1206] "Collected data" refers to information such as product information, reviews, and past posting data obtained from advertisers and affiliated sites.
[1207] "Preprocessing" is the process of removing unnecessary information from collected data and shaping the data.
[1208] A "generative artificial intelligence model" is an artificial intelligence algorithm that can generate new text based on input data.
[1209] "Training" is the learning process used to improve the accuracy of a generative artificial intelligence model based on specific data.
[1210] "Style" refers to characteristics such as writing style, tone, and format.
[1211] "Tone" describes the mood or emotional atmosphere of a piece of writing.
[1212] An "advertising piece" is a piece of writing written to promote a particular product or service.
[1213] "Advertising notation" refers to text or notes that clearly indicate that the item is an advertisement.
[1214] An "annotation" is an explanatory text based on supplementary information or legal requirements inserted within an article.
[1215] "Internal checking" is the process of checking the content of generated articles and correcting errors and inappropriate expressions.
[1216] "Publishing on a website" means posting the generated article on the web so that it can be viewed by general internet users.
[1217] "Using a smartphone" means performing various operations and processes using a smartphone.
[1218] "Preview" is a function that allows you to check the state before the final release.
[1219] "Editing" is the process of correcting or changing the content of the generated article.
[1220] A "content management system" is software for organizing, managing, and delivering content.
[1221] An "API" is an interface that allows applications and services to communicate with each other and utilize their functionality.
[1222] "Automatically posting" means that the system automatically uploads content to a website based on a set schedule or conditions.
[1223] As an embodiment of the present invention, a novel advertising article generation system is described, which preprocesses collected data, trains a generative artificial intelligence model, and ultimately automatically generates and publishes highly accurate advertising articles.
[1224] System Configuration
[1225] The system of the present invention includes the following major components:
[1226] 1. Data Collection Module
[1227] 2. Data Preprocessing Module
[1228] 3. Generative AI Model Training Module
[1229] 4. Style and Tone Extraction Module
[1230] 5. Annotation Learning Module
[1231] 6. Article Generation Module
[1232] 7. Internal Check Module
[1233] 8. Article Publishing Module
[1234] 9. Smartphone Application Module
[1235] 10. Preview and Edit Module
[1236] 11. Auto-posting module
[1237] Hardware and Software Used
[1238] The main hardware and software components of this system include:
[1239] Hardware: High-performance servers and smartphones.
[1240] Software: Python, BeautifulSoup, Scrapy, pandas, transformers, React Native, Content Management Systems (CMS) WordPress and Wix, OpenAI's ChatGPT.
[1241] Processing Description
[1242] 1. Data Collection
[1243] The server collects the necessary data from advertisers and partner sites using APIs and web crawling technologies (BeautifulSoup and Scrapy).
[1244] 2. Data Preprocessing
[1245] The collected data is cleaned using a data preprocessing module. Specifically, HTML tags are removed, unnecessary strings are deleted using regular expressions, and important key phrases and features are extracted. The data is formatted and processed using the pandas library.
[1246] 3. Generative AI Model Training
[1247] The server uses the preprocessed data to train a generative AI model (ChatGPT), which learns specific writing styles and expressions.
[1248] 4. Extracting Style and Tone
[1249] The server extracts style and tone from the text of advertisers and partner sites and feeds this back into a generative artificial intelligence model.
[1250] 5. Annotation Learning
[1251] The server trains the model with annotations based on legal requirements and applies them to the generated advertising articles.
[1252] 6. Article Generation
[1253] Using a generative artificial intelligence model, advertising articles are automatically generated based on collected and learned data.
[1254] 7. Internal Checks
[1255] The generated articles are then internally checked by the user to ensure there are no errors, misrepresentations, or missing notes, and any necessary corrections are made. The quality can also be checked again using the generative AI model.
[1256] 8. Public
[1257] Once the article has been checked, it will be published on the website via the terminal.
[1258] 9. Smartphone use
[1259] Users can quickly and easily create, preview, edit and publish articles using their smartphones.
[1260] 10. Preview and Edit
[1261] It displays a preview of the generated article and allows users to easily make any necessary corrections. It is provided with an interface that uses React Native.
[1262] 11. Auto-posting
[1263] Finally, the generated articles are automatically posted to the website via the content management system's API.
[1264] Specific examples
[1265] For example, generating an advertisement for a health device A involves the following process:
[1266] 1. The server collects information about the ingredients, efficacy, and reviews of health device A.
[1267] 2. The server preprocesses the collected data, removing HTML tags, deleting unnecessary strings, and tokenizing the text.
[1268] 3. The server trains the ChatGPT model to learn specific writing styles and expressions.
[1269] 4. The server extracts the advertiser's formal tone and the partner site's friendly tone and trains the model accordingly.
[1270] 5. The server learns the advertising and annotations and applies them to the generated articles.
[1271] 6. The server generates an article such as "Health device A contains ingredient B and is expected to be effective in improving immunity."
[1272] 7. The user checks the generated article for errors, misrepresentations, and missing notes, and makes corrections as necessary.
[1273] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[1274] Prompt Sentence Examples
[1275] Product name: Health equipment A
[1276] Features: Quiet design, foldable, multi-function display
[1277] As described above, the present invention enables efficient and highly accurate generation and publication of advertising articles.
[1278] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1279] Step 1:
[1280] Data collection
[1281] The server uses APIs and web crawling technologies (BeautifulSoup and Scrapy) to retrieve product details, reviews, and past postings from advertisers and partner sites. The input is the URL or API endpoint of each site, and the output is the retrieved raw data. This allows a wealth of information to be stored in the database used.
[1282] Step 2:
[1283] Data Preprocessing
[1284] The server preprocesses the collected data. Specifically, it uses the pandas library to remove HTML tags, delete unnecessary strings, and extract important key phrases and features. The input is the raw data collected in step 1, and the output is cleansed and formatted data. This process processes the data into a form that is easy for the generative artificial intelligence model to process.
[1285] Step 3:
[1286] Generative AI model training
[1287] The server uses the preprocessed data to train a generative AI model (ChatGPT). In the process, it learns specific writing styles and expressions. Specifically, it fine-tunes the model using the transformers library. The input is the preprocessed data, and the output is a trained generative AI model. This allows the model to generate sentences with even greater accuracy.
[1288] Step 4:
[1289] Style and Tone Extraction
[1290] The server performs text analysis to extract style and tone from the text of advertisers' and partner sites. Specifically, it uses machine learning algorithms and natural language processing (NLP) technology to extract text characteristics. The input is text data from advertisers' and partner sites, and the output is the extracted style and tone features. This ensures that the generated articles are tailored to the characteristics of each site.
[1291] Step 5:
[1292] Annotation Learning
[1293] The server collects specific texts (legal annotations) and trains a generative AI model to learn annotations based on legal requirements. Specifically, it identifies patterns in the annotations and incorporates them into the model. The input is a dataset of legal annotations, and the output is a generative AI model that can insert annotations appropriately. This results in articles that reduce legal risk.
[1294] Step 6:
[1295] Article Generation
[1296] The server uses a generative artificial intelligence model to automatically generate advertising articles based on collected and learned data. Specifically, when a prompt is entered, the corresponding trained model generates an article. The input is the prompt and past posting data, and the output is an automatically generated advertising article. This enables fast and highly accurate article generation.
[1297] Example prompt sentence:
[1298] Product name: Health equipment A
[1299] Features: Quiet design, foldable, multi-function display
[1300] Step 7:
[1301] Internal Check
[1302] The generated article undergoes an internal check by the user. Specifically, it checks for typographical errors, misrepresentations, and missing notes, and makes any necessary corrections. The input is the automatically generated advertising article, and the output is the corrected final version of the article. This allows the article to be published with guaranteed quality.
[1303] Step 8:
[1304] Article published
[1305] Once the article has been checked, it is published on the website via the terminal. Specifically, the article is automatically posted via the API of the content management system (CMS). The input is the final, corrected version of the article, and the output is the published article. This allows the article to be published efficiently on the web, saving the user time and effort.
[1306] Step 9:
[1307] Smartphone use
[1308] Users use their smartphones to create, preview, edit, and publish articles. Specifically, articles are created and edited using a dedicated application. The input is the requirements and prompts specified by the user, and the output is the edited article. This allows article creation and management to be done anywhere, anytime.
[1309] Step 10:
[1310] Preview and Edit
[1311] A preview of the generated article is displayed, allowing the user to easily make any necessary corrections. Specifically, the interface uses React Native, and edits are reflected in real time. The input is the automatically generated article, and the output is the article corrected by the user. This further improves the quality of the final article.
[1312] Step 11:
[1313] Auto-post
[1314] Finally, the generated articles are automatically posted to the website via the content management system's API. The server uploads the articles to the website based on the set schedule and conditions. The input is the checked, final version of the article, and the output is the published article and its publication log. This minimizes user effort and efficiently publishes articles through an automated process.
[1315] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1316] The following description is provided as an embodiment of the present invention. This system collects data from advertisers and affiliate sites and generates highly accurate affiliate articles using a generative artificial intelligence model (e.g., ChatGPT). Furthermore, by combining it with an emotion engine that recognizes user emotions, the system can be customized according to the user's emotions.
[1317] System Configuration
[1318] The system mainly consists of the following components:
[1319] 1. Data Collection Module
[1320] 2. Data Preprocessing Module
[1321] 3. Generative AI Model Training Module
[1322] 4. Style and Tone Extraction Module
[1323] 5. Annotation Learning Module
[1324] 6. Emotion Engine Module
[1325] 7. Article Generation Module
[1326] 8. Internal Check Module
[1327] 9. Article Publishing Module
[1328] Data collection
[1329] First, the server collects the necessary data from the advertiser and affiliate sites. The collected data includes detailed product information, review articles, and past posting data from the affiliate sites. The server obtains the data from each site using web crawling technology.
[1330] Data Preprocessing
[1331] Next, the server preprocesses the collected data. This step involves data cleaning to remove unnecessary information, such as removing HTML tags, deleting unnecessary strings using regular expressions, tokenizing the text, and extracting important key phrases and features.
[1332] Training generative artificial intelligence models
[1333] Using the preprocessed data, the server trains a generative AI model (ChatGPT), which fine-tunes the model based on the initial dataset to learn specific writing styles and expressions.
[1334] Style and Tone Extraction
[1335] The server extracts the style and tone of the advertiser's and affiliate's sites and trains the generative AI model. During this process, the server analyzes each site's writing style, tone, and punctuation usage. This data is then fed back into the generative model to generate articles that match the characteristics of each site.
[1336] Annotation Learning
[1337] The server trains a generative AI model to create appropriate ad descriptions and annotations. This involves collecting standard annotations based on legal requirements and teaching the model necessary annotation information, such as "This product is not a medicine."
[1338] Emotion Engine
[1339] The server uses an emotion engine to analyze the user's emotions in real time. The emotion engine recognizes the user's emotional state based on user input and various sensor data (e.g., facial recognition and voice analysis). Based on the results, the style and tone of the generated articles can be adjusted.
[1340] Article Generation
[1341] The server automatically generates affiliate articles based on the collected and learned data. The generative AI model emphasizes the product's benefits and features and generates text that matches the tone and style. For example, if the user is in a positive emotional state, text with a more optimistic tone will be generated. Annotations based on legal requirements are also automatically inserted into the generated articles.
[1342] Internal Check
[1343] The generated articles are then checked internally by a user (an internal checker). A checklist is used to check for typographical errors, misrepresentations, and missing notes. During this process, the quality can be checked again using the generative AI model. If necessary, the user can make manual corrections.
[1344] Article Publication
[1345] Finally, articles that have passed the check are published on the website via the terminal, which automatically posts the articles based on the publication schedule and records the publication log.
[1346] Specific examples
[1347] For example, generating an affiliate article for healthcare product A involves the following process:
[1348] 1. The server collects information about the ingredients, efficacy, and reviews of healthcare product A.
[1349] 2. The server preprocesses the collected data, removing HTML tags, deleting unnecessary strings, and tokenizing the text.
[1350] 3. The server trains the ChatGPT model to learn specific writing styles and expressions.
[1351] 4. The server extracts the advertiser's formal tone and the affiliate site's friendly tone and trains the model accordingly.
[1352] 5. The server learns the advertising and annotations and applies them to the generated articles.
[1353] 6. The server uses an emotion engine to analyze the user's emotions and adjust the tone of the article accordingly. For example, if the user is excited, the article will be generated with a more enthusiastic tone.
[1354] 7. The user reviews the generated article, checking for errors, misrepresentations, and missing notes, and makes corrections as necessary.
[1355] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[1356] The above system makes it possible to efficiently generate high-quality affiliate articles that respond to user sentiment and publish them while reducing legal risks.
[1357] The processing flow will be explained below.
[1358] Step 1:
[1359] The server collects data from advertiser and affiliate sites. Specifically, it uses a crawler to retrieve the HTML of web pages from the specified URLs and stores it in a database. The retrieved data includes titles, body text, meta tags, image URLs, etc.
[1360] Step 2:
[1361] The server preprocesses the collected data, removing HTML tags and extracting only the text data. Regular expressions are also used to remove unnecessary symbols and spaces. Next, the text is tokenized and techniques such as TF-IDF and Word2Vec are used to extract important key phrases and features.
[1362] Step 3:
[1363] The server uses the preprocessed data to train a generative AI model, based on the initial ChatGPT model, which is then fine-tuned to fit specific writing styles and themes. This process uses high-quality samples from the collected data as training data.
[1364] Step 4:
[1365] The server extracts the style and tone of the advertiser and affiliate sites, uses analysis algorithms to identify the unique writing style and tone patterns of each site, and feeds these features back into the generative AI model.
[1366] Step 5:
[1367] The server trains a generative AI model to create appropriate annotations. Standard annotations and warnings based on legal requirements are collected and used as training data for the model. This training gives the model the ability to automatically insert appropriate annotations when generating articles.
[1368] Step 6:
[1369] The server starts an emotion engine that analyzes the user's emotions in real time. The emotion engine detects emotions from the user's input (text, voice, facial recognition data, etc.) and records the results in a database.
[1370] Step 7:
[1371] The server adjusts the style and tone of the generated article based on the user's emotional data analyzed by the emotion engine. For example, if the user is in a positive emotional state, the server generates an optimistic tone of text.
[1372] Step 8:
[1373] Affiliate articles are automatically generated based on the data collected and learned by the server. As a specific example, when generating an article for "Healthcare Product A," the article is generated by incorporating product features and user reviews and adjusting the tone.
[1374] Step 9:
[1375] User-generated articles are internally checked using a dedicated checklist to ensure there are no errors, misrepresentations, or missing notes. Again, during this process, the generative AI model can be used to check the quality.
[1376] Step 10:
[1377] The device publishes the checked articles on the website. Articles are automatically posted according to the publication schedule. The URL and related information after publication are recorded in a database and stored as data for analysis.
[1378] Through the above steps, high-quality affiliate articles that respond to user sentiment can be efficiently generated and published with reduced legal risks.
[1379] Example 2
[1380] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1381] Conventional affiliate article generation systems have had issues with preprocessing collected data, matching the style and tone of generated articles, appropriately inserting advertisements and annotations, and customizing articles to reflect user sentiment. In particular, generating content based on user sentiment is difficult, and generated articles often do not reflect the user's sentiment. There is also a need for more efficient internal checks of generated articles.
[1382] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1383] In this invention, the server includes means for preprocessing collected data and training a generative artificial intelligence model, means for extracting style and tone from the advertiser's site and affiliate site and training the generative artificial intelligence model, and means for the generative artificial intelligence model to analyze user emotions and adjust the style and tone of articles. This enables the generation of more personalized affiliate articles in response to user emotions. Furthermore, using efficiently preprocessed data allows for the efficient generation of high-quality articles, contributing to the efficiency of internal checks.
[1384] "Collected data" refers to information such as detailed product information, review articles, and past posting data obtained from advertiser and affiliate sites using web crawling technology.
[1385] "Preprocessing" refers to a series of processes performed on collected data, such as removing HTML tags, deleting unnecessary strings, tokenizing text, and extracting key phrases and features.
[1386] A "generative artificial intelligence model" is an artificial intelligence algorithm that generates sentences based on training data, and a specific example is ChatGPT.
[1387] "Training" is the process of using preprocessed data to teach a generative artificial intelligence model a particular style or method of expression.
[1388] "Advertiser's Site and Affiliate Site" refers to a website that provides detailed product information and review articles, and includes affiliate marketing content.
[1389] "Style and tone" refers to the writing style, tone, punctuation, and other characteristics of a particular site or document.
[1390] "Learning" is the process by which a generative AI model analyzes data to acquire a particular style or tone, and then applies the results to the model.
[1391] "User emotion" refers to the emotional state analyzed based on the user's input information (text, voice, facial image, etc.), and includes positive, negative, excited, etc.
[1392] "Internal checking" is a process in which the generated articles are checked for errors, misrepresentations, missing notes, etc., and is carried out by in-house checkers.
[1393] "Advertising statements and annotations" refers to standard legally required annotations and advertising statements, such as warnings such as "This product is not a medicine."
[1394] "Publishing on a website" means posting the checked article on a website based on a specific publishing schedule and recording a publishing log.
[1395] To implement this invention, the following system must be constructed. This system uses data collected from advertisers and affiliate sites and a generative artificial intelligence model to automatically generate highly accurate affiliate articles. Furthermore, it is possible to analyze user sentiment and adjust the style and tone of the resulting articles.
[1396] System configuration
[1397] The system consists of the following main components:
[1398] 1. Data Collection Module
[1399] 2. Data Preprocessing Module
[1400] 3. Generative AI Model Training Module
[1401] 4. Style and Tone Extraction Module
[1402] 5. Annotation Learning Module
[1403] 6. Emotion Engine Module
[1404] 7. Article Generation Module
[1405] 8. Internal Check Module
[1406] 9. Article Publishing Module
[1407] Hardware and software used
[1408] The implementation of this system uses the following hardware and software:
[1409] Server: A server with high-performance processing power is used to process collected data, train generative AI models, analyze emotions, and perform other heavy workloads.
[1410] Generative AI models, such as ChatGPT, are used to generate text and adjust style and tone.
[1411] Emotion engine: An engine for analyzing user emotions, using sensors such as facial recognition and voice analysis.
[1412] Program processing
[1413] The server first collects data from advertiser and affiliate sites, including product information, reviews, and past posting data, and the collected data is obtained using web crawling technology.
[1414] The server then preprocesses the collected data, removing HTML tags, eliminating unnecessary strings, tokenizing the text, and extracting important key phrases and features.
[1415] Based on the preprocessed data, the server trains a generative AI model (ChatGPT), fine-tuning the model based on the initial dataset to learn specific writing styles and expressions.
[1416] The server then extracts the style and tone of the advertiser and affiliate sites and trains them into a generative AI model. Each site's writing style, tone, and punctuation usage are analyzed and fed back to the model.
[1417] The server also trains a generative AI model to create appropriate ad labeling and annotations, collecting standard annotations based on legal requirements and training the model to apply them to the generated articles.
[1418] Additionally, the server uses an emotion engine to analyze users' emotions in real time, recognizing their emotional state based on user input and various sensor data (facial recognition and voice analysis), and adjusting the style and tone of the generated articles accordingly.
[1419] As a concrete example, the prompt text to automatically generate an affiliate article for healthcare product A is as follows:
[1420] "Generate an article with an optimistic tone based on the ingredients, efficacy, and reviews of healthcare product A when the user is in a positive emotional state."
[1421] Finally, the generated article is checked internally by the user (an in-house checker). A checklist is used to check for typographical errors, misrepresentations, and missing notes, and manual corrections are made as necessary. The completed article is published on the website via the terminal, and a publication log is recorded.
[1422] summary
[1423] This system allows us to efficiently generate high-quality affiliate articles that reflect user sentiment and publish them while reducing legal risks.
[1424] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1425] Step 1:
[1426] Data collection
[1427] The server collects data from advertiser and affiliate sites. It uses web crawling technology to obtain HTML content based on the URL list specified as input. Specifically, it collects detailed product information, review articles, and past posting data from advertiser and affiliate sites. The output is raw data, which can be in a wide variety of formats.
[1428] Step 2:
[1429] Data Preprocessing
[1430] The data collected by the server is preprocessed. The input is the raw data collected in step 1. Specifically, regular expressions are used to remove HTML tags and extract only the text. Unnecessary strings (e.g., advertising banner text) are also filtered. The text is then tokenized and split into words. Important key phrases are then extracted using algorithms such as TF-IDF and Word2Vec. The output is preprocessed, clean text data.
[1431] Step 3:
[1432] Training generative artificial intelligence models
[1433] The server uses the preprocessed data to train a generative AI model (ChatGPT). The input is preprocessed, clean text data. Specifically, the initial dataset is fed into the generative AI model, and fine-tuned to learn specific writing styles and expressions. The output is a trained generative AI model.
[1434] Step 4:
[1435] Style and Tone Extraction
[1436] The server analyzes the style and tone of advertiser and affiliate sites and trains a generative AI model. The input is the text data of each site. Specific processing involves analyzing writing style (frequency of use of punctuation at the end of sentences, use of honorifics, etc.) and tone (formal or informal, optimistic or pessimistic, etc.). The results of these analyses are then fed back to the generative AI model, enabling it to generate text that matches the characteristics of each site. The output is a model that has learned the characteristics of style and tone.
[1437] Step 5:
[1438] Annotation Learning
[1439] The server trains a generative AI model on appropriate methods for advertising labeling and annotations. The input is standard annotations based on legal requirements (e.g., "This product is not a medicine"). Specific processing involves collecting standard annotations and training the generative AI model so that they can be applied when needed. The output is a model that can insert annotations.
[1440] Step 6:
[1441] Emotion Engine
[1442] The server uses an emotion engine to analyze the user's emotions in real time. The input is the user's input information (text, voice, facial image, etc.). Specifically, the process uses sensors such as facial recognition and voice analysis to recognize the user's emotional state (positive, negative, excited, etc.). Based on the results, the style and tone of the generated article are adjusted. The output is text whose tone and style have been adjusted to reflect the user's emotional state.
[1443] Step 7:
[1444] Article Generation
[1445] The server automatically generates affiliate articles based on collected and learned data. The input is a model that has learned style and tone characteristics that reflect emotional states, as well as requirements such as product information. A generative AI model is used to generate text that highlights the product's features and benefits. Annotations based on legal requirements are also automatically inserted. The output is the completed affiliate article.
[1446] Step 8:
[1447] Internal Check
[1448] The user (an internal reviewer) checks the generated article. The input is the generated affiliate article. Specifically, a checklist is used to check for errors, misrepresentations, and missing notes. Manual corrections may be made as necessary. The output is the final version of the article after corrections have been completed.
[1449] Step 9:
[1450] Article Publication
[1451] The terminal publishes the checked article on the website. The input is the final version of the article. Specific processing involves setting a publishing schedule and automatically posting the article based on the specified date and time. A publishing log is also recorded. The output is the published article and its publishing log.
[1452] (Application example 2)
[1453] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1454] Conventional affiliate article generation systems have difficulty generating advertising content that takes user emotions into account, making it impossible to provide ads that are adapted to user emotions. Furthermore, it is difficult to properly reflect the style and tone of the advertiser or affiliate site, making it impossible to provide an effective advertising experience for users.
[1455] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for preprocessing collected data and training a generative AI model, means for extracting style and tone from the advertiser's website and affiliate site and training the generative AI model, means for automatically generating affiliate articles based on the collected and learned data, means for appropriately inserting advertising notices and annotations into the generated articles, means for internally checking the generated articles, means for publishing the articles generated after the internal check on the website, and means for recognizing user emotions and adjusting the tone and style of the generated advertising content based on the emotions. This makes it possible to provide effective advertising content that corresponds to the user's emotions and generate and publish high-quality affiliate articles while reflecting the style and tone of the advertiser and affiliate site.
[1456] "Collected data" refers to information such as detailed product information, review articles, and past posting data obtained from advertisers and affiliate sites.
[1457] "Preprocessing" refers to the process of cleaning the collected data, removing unnecessary information, and extracting important key phrases and features using regular expressions.
[1458] A "generative artificial intelligence model" is an artificial intelligence model that learns specific writing styles and methods of expression based on given data, and generates new sentences based on the results of that learning.
[1459] "Training" refers to the learning process of using preprocessed data to improve the performance of a generative artificial intelligence model.
[1460] "Style and tone extraction" refers to the process of analyzing characteristics such as writing style and punctuation usage from advertiser websites and affiliate sites, and training a generative AI model to learn these characteristics.
[1461] An "affiliate article" refers to a piece of writing that provides information about a specific product or service and is intended to encourage users to purchase that product or service.
[1462] "Blank and annotation insertion" refers to the process of automatically adding standard legally required annotations to generated articles.
[1463] "Internal Check" refers to the internal review process to ensure that the content of the generated article is accurate and that there are no misrepresentations or omissions.
[1464] "Publication" refers to making the generated article publicly available on a website or other media.
[1465] "Emotion recognition" refers to the process of analyzing a user's emotional state using user input and sensor data (e.g., facial recognition and voice analysis).
[1466] "Tone and style adjustment" refers to the process of changing the tone and presentation of generated text based on the perceived user sentiment.
[1467] The present invention provides an emotion-based advertising content generation system, which includes means for preprocessing collected data and training a generative AI model, means for extracting style and tone from advertiser and affiliate sites and training the generative AI model, means for automatically generating affiliate articles based on the collected and learned data, means for appropriately inserting advertising notices and annotations into the generated articles, means for internally checking the generated articles, means for publishing the generated articles on a website after the internal check, and means for recognizing user emotions and adjusting the tone and style of the generated advertising content based on the emotions.
[1468] Program processing explanation
[1469] The server uses an emotion recognition module to analyze the user's emotions in real time. This module utilizes facial recognition technology using the smartphone camera and TensorFlow. Specifically, it captures the user's face and analyzes their emotions using a trained emotion recognition model (e.g., CNN model).
[1470] The server then uses web crawling techniques to collect the required data from the advertiser's website and affiliate sites, using Python web scraping libraries (e.g., BeautifulSoup and Requests), and preprocesses the collected data by removing HTML tags and extracting text information using regular expressions.
[1471] Based on the preprocessed data, the server trains a generative AI model (e.g., GPT-2). During this process, the collected data is analyzed for its writing style and tone, and the model is trained to learn from it. In particular, the model extracts features of the advertiser's formal writing style and the affiliate site's friendly tone.
[1472] The generated advertisements are automatically inserted with ad captions and annotations to comply with legal requirements. In this process, a generative model is trained on standard annotations collected in advance, and the resulting advertisement captions and annotations are added to the generated articles.
[1473] After generation is complete, the server performs an internal check to ensure there are no errors, misrepresentations, or missing notes. This includes a second quality check using the generative AI model. Once the check is complete, the article is published to the website via the terminal. Here, posts are automatically made based on the publishing schedule, and a publishing log is recorded.
[1474] Examples of concrete examples and prompts
[1475] For example, generating an affiliate article for healthcare product A involves the following process:
[1476] The server uses an emotion recognition module to analyze whether the user is in a positive emotional state and generates an affiliate article with a positive tone. The generated article uses the following prompt:
[1477] "Preprocess this data, normalize it, filter it, tokenize it."
[1478] "Generate ad text based on sentiment and preprocessed data."
[1479] "Please display the generated ad text on your smartphone screen."
[1480] In this way, it becomes possible to provide effective advertising content that matches the user's emotions, and high-quality affiliate articles that reflect the style and tone of the advertiser and affiliate site are generated.
[1481] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1482] Step 1:
[1483] Recognize user emotions.
[1484] The server captures the user's face in real time using the smartphone camera. The captured image is input to an emotion recognition model using TensorFlow to analyze the user's emotional state (e.g., positive, negative, neutral). The analysis result is passed to the next processing step.
[1485] Step 2:
[1486] Data collection.
[1487] The server uses web scraping technology to collect the necessary data from advertiser and affiliate sites. The collected data includes product details, reviews, past posting data, etc. This data is pre-processed in the next step.
[1488] Step 3:
[1489] Data preprocessing.
[1490] The server preprocesses the collected data. Specifically, it uses Python's BeautifulSoup and Requests to remove HTML tags and regular expressions to remove unnecessary characters. It also tokenizes the text and extracts important key phrases. The preprocessed data is used to train generative artificial intelligence models.
[1491] Step 4:
[1492] Training generative artificial intelligence models.
[1493] The server uses the preprocessed data to train a generative artificial intelligence model (e.g., GPT-2). During this training, the model learns the style and tone of the advertiser's website and affiliate site. The trained model is then used in the next step.
[1494] Step 5:
[1495] Generating advertising articles.
[1496] The server uses the trained generative artificial intelligence model to generate advertising articles based on the user's emotional state. For example, if the user is in a positive emotional state, an article with a positive tone is generated. The generated advertising article is then processed in the following steps.
[1497] Step 6:
[1498] Inserting advertising labels and annotations.
[1499] The server automatically inserts appropriate advertisements and annotations into the generated articles by training a generative AI model with standard annotations based on legal requirements, resulting in articles that mitigate legal risks.
[1500] Step 7:
[1501] Internal check.
[1502] The generated articles are then subjected to an internal check by the server, where a quality check is carried out again using the generative AI model to check for errors, misrepresentations, and missing notes. If necessary, an in-house checker manually corrects the errors.
[1503] Step 8:
[1504] Publication of article.
[1505] The device publishes articles on the website after internal checks are complete. Publication is done automatically based on a pre-set schedule and a publication log is recorded, streamlining the article publishing process.
[1506] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1507] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1508] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1509] [Fourth embodiment]
[1510] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1511] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1512] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1513] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1514] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1515] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1516] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1517] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1518] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1519] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1520] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1521] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1522] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1523] The following description is provided as an example of an embodiment of the present invention: The system collects data from advertisers and affiliate sites and uses a generative artificial intelligence model (e.g., ChatGPT) to generate highly accurate affiliate articles.
[1524] System Configuration
[1525] The system mainly consists of the following components:
[1526] 1. Data Collection Module
[1527] 2. Data Preprocessing Module
[1528] 3. Generative AI Model Training Module
[1529] 4. Style and Tone Extraction Module
[1530] 5. Annotation Learning Module
[1531] 6. Article Generation Module
[1532] 7. Internal Check Module
[1533] 8. Article Publishing Module
[1534] Data collection
[1535] First, the server collects the necessary data from the advertiser and affiliate sites. The collected data includes detailed product information, review articles, and past posting data from the affiliate sites. The server obtains the data from each site using web crawling technology.
[1536] Data Preprocessing
[1537] Next, the server preprocesses the collected data. This step involves data cleaning to remove unnecessary information, such as removing HTML tags, deleting unnecessary strings using regular expressions, tokenizing the text, and extracting important key phrases and features.
[1538] Training generative artificial intelligence models
[1539] Using the preprocessed data, the server trains a generative AI model (ChatGPT), which fine-tunes the model based on the initial dataset to learn specific writing styles and expressions.
[1540] Style and Tone Extraction
[1541] The server extracts the style and tone of the advertiser's and affiliate's sites and trains the generative AI model. During this process, the server analyzes each site's writing style, tone, and punctuation usage. This data is then fed back into the generative model to generate articles that match the characteristics of each site.
[1542] Annotation Learning
[1543] The server trains a generative AI model to create appropriate ad descriptions and annotations. This involves collecting standard annotations based on legal requirements and teaching the model necessary annotation information, such as "This product is not a medicine."
[1544] Article Generation
[1545] The server automatically generates affiliate articles based on the collected and learned data. For example, when generating an article about "Healthcare Product A," the generative AI model emphasizes the product's effects and features and generates text that matches the tone and style. Annotations based on legal requirements are also automatically inserted into the article.
[1546] Internal Check
[1547] The generated articles are then checked internally by a user (an internal checker). A checklist is used to check for typographical errors, misrepresentations, and missing notes. During this process, the quality can be checked again using the generative AI model. If necessary, the user can make manual corrections.
[1548] Article Publication
[1549] Finally, articles that have passed the check are published on the website via the terminal, which automatically posts the articles based on the publication schedule and records the publication log.
[1550] Specific examples
[1551] For example, generating an affiliate article for healthcare product A involves the following process:
[1552] 1. The server collects information about the ingredients, efficacy, and reviews of healthcare product A.
[1553] 2. The server preprocesses the collected data, removing HTML tags, deleting unnecessary strings, and tokenizing the text.
[1554] 3. The server trains the ChatGPT model to learn specific writing styles and expressions.
[1555] 4. The server extracts the advertiser's formal tone and the affiliate site's friendly tone and trains the model accordingly.
[1556] 5. The server learns the advertising and annotations and applies them to the generated articles.
[1557] 6. The server generates an article such as, "Healthcare product A contains ingredient B and is expected to be effective in improving immunity."
[1558] 7. The user checks the generated article for errors, misrepresentations, and missing notes, and makes corrections as necessary.
[1559] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[1560] The above system allows high-quality affiliate articles to be generated efficiently and published with reduced legal risks.
[1561] The processing flow will be explained below.
[1562] Step 1:
[1563] The server collects data from advertiser and affiliate sites. Specifically, it uses a crawler to retrieve the HTML of web pages from the specified URLs and stores it in a database. The retrieved data includes titles, body text, meta tags, image URLs, etc.
[1564] Step 2:
[1565] The server preprocesses the collected data, removing HTML tags and extracting only the text data. Regular expressions are also used to remove unnecessary symbols and spaces. Next, the text is tokenized and techniques such as TF-IDF and Word2Vec are used to extract important key phrases and features.
[1566] Step 3:
[1567] The server uses the preprocessed data to train a generative AI model, based on the initial ChatGPT model, which is then fine-tuned to fit specific writing styles and themes. This process uses high-quality samples from the collected data as training data.
[1568] Step 4:
[1569] The server extracts the style and tone of the advertiser and affiliate sites, uses analysis algorithms to identify the unique writing style and tone patterns of each site, and feeds these features back into the generative AI model.
[1570] Step 5:
[1571] The server trains a generative AI model to create appropriate annotations. Standard annotations and warnings based on legal requirements are collected and used as training data for the model. This training gives the model the ability to automatically insert appropriate annotations when generating articles.
[1572] Step 6:
[1573] The server automatically generates affiliate articles based on the data it has collected and learned. For example, it generates articles that emphasize the effectiveness and features of specific products and services using information about those products and services. The generated articles also include necessary annotations.
[1574] Step 7:
[1575] User-generated articles are internally checked using a dedicated checklist to ensure there are no errors, misrepresentations, or missing notes. This process can also be repeated using a generative AI model to check quality.
[1576] Step 8:
[1577] The device publishes the checked articles on the website. Articles are automatically posted according to the publication schedule. The URL and related information after publication are recorded in a database and stored as data for analysis.
[1578] Through the above steps, high-quality affiliate articles can be efficiently generated and published with reduced legal risks.
[1579] Example 1
[1580] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1581] In conventional affiliate article generation systems, a lot of manual work is required to efficiently generate a large amount of content while improving the quality of the articles, which leads to issues such as inconsistency in the articles and inconsistency in the quality of the articles.In addition, inserting accurate annotations based on legal requirements and internal article checking processes require a great deal of effort.
[1582] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1583] In this invention, the server includes means for preprocessing the collected data and training the generative AI model, means for extracting style and tone from advertiser information sources and websites and training the generative AI model, and means for automatically generating text based on the collected and learned data, thereby enabling the efficient generation of high-quality affiliate articles and the provision of content with a consistent style and tone.
[1584] "Collected data" refers to detailed product information, review articles, past posting data, etc. obtained from advertisers and affiliate sites.
[1585] "Preprocessing" refers to a series of processes that remove HTML tags and unnecessary strings from collected data, tokenize the text, and extract important key phrases and features.
[1586] A "generative artificial intelligence model" refers to an artificial intelligence that learns a specific writing style or expression method based on an initial dataset.
[1587] "Training methods" refers to the use of collected data to fine-tune generative AI models to learn specific writing styles and idiomatic expressions.
[1588] "Advertiser Source" refers to a website or database that contains information about an advertiser.
[1589] "Style and tone extraction" refers to the process of analyzing each site's writing style, tone, and punctuation usage and training the model.
[1590] "Means for automatic generation" refers to a method of automatically generating sentences based on collected and learned data using a generative artificial intelligence model.
[1591] "Means for appropriately inserting annotations" refers to a method for automatically inserting standard legal annotations into the generated text.
[1592] "Means for internal checking" refers to a method in which an in-house checker reviews the generated text to check for any typos, misrepresentations, or missing notes.
[1593] "Means of disclosing to the information source" refers to a method in which the generated text is automatically posted on a website after being verified and a disclosure log is recorded.
[1594] "Crawling" refers to the act of collecting data from a website using an automated crawler program.
[1595] "HTML tag stripping" refers to the process of removing unnecessary HTML code from text data obtained from a web page.
[1596] "Extracting text information" refers to the method of extracting the actual required text parts from the preprocessed data.
[1597] "Methods for checking for errors, misrepresentations, and missing notes" refers to methods for checking the content of generated articles and checking for typos, inaccurate expressions, and missing necessary notes.
[1598] This invention relates to a system that collects data from advertisers and affiliate sites and generates highly accurate affiliate articles using a generative artificial intelligence model (e.g., a generative AI model). The system automatically inserts annotations based on legal requirements into the generated articles, and then publishes them on a website after internal checks.
[1599] Hardware and Software
[1600] The system requires the following hardware and software:
[1601] Server: Collects data, preprocesses it, trains the generative AI model, generates articles, and publishes them.
[1602] Terminal: Serves as an interface for publishing checked articles on a website.
[1603] User: Internal checks and corrections of generated articles.
[1604] The main software used includes:
[1605] Web crawler tools: Used to collect data from advertiser and affiliate sites.
[1606] Data preprocessing library: Used to remove HTML tags, tokenize text, and extract key phrases.
[1607] Generative artificial intelligence model (generative AI model): Refers to a generative artificial intelligence platform used to generate articles, such as a generative AI model or ChatGPT.
[1608] Linguistic Style Analysis Library: Used to extract style and tone.
[1609] Text review tools: Used for internal checks and to identify errors, misstatements, and missing notes.
[1610] Specific processing and calculation of data
[1611] 1. Data Collection:
[1612] The server uses a web crawler tool to automatically retrieve product details, reviews, and past posting data from advertiser and affiliate sites, and stores this data in an internal database.
[1613] 2. Data Preprocessing:
[1614] The server performs preprocessing on the collected data, such as removing HTML tags, deleting unnecessary text using regular expressions, tokenizing, and extracting important key phrases.
[1615] 3. Training the generative AI model:
[1616] The server uses the preprocessed data to train a generative artificial intelligence model, a process that customizes the model by teaching it specific writing styles and idioms.
[1617] 4. Extracting style and tone:
[1618] The server uses a language style analysis library to analyze the style and tone of advertiser and affiliate sites, and the analysis data is fed back to a generative AI model.
[1619] 5. Annotation Learning:
[1620] The server collects annotations based on legal requirements and trains a generative artificial intelligence model based on them.
[1621] 6. Article Generation:
[1622] The server uses a generative artificial intelligence model to generate affiliate articles based on the collected and learned data, and the generated articles are automatically supplemented with the necessary legal annotations.
[1623] 7. Internal check:
[1624] The generated articles are then internally checked by users who use text review tools to identify errors, misstatements, and omissions, and make manual corrections as needed.
[1625] 8. Article Publication:
[1626] Finally, articles that have passed the check are published on the website via the terminal, which posts the articles based on the publication schedule and records the publication log.
[1627] Examples of concrete examples and prompts
[1628] For example, let's say you want to generate an affiliate article for "Healthcare Product A":
[1629] 1. The server collects information about the ingredients, efficacy, and reviews of healthcare product A.
[1630] 2. The server preprocesses the collected data and removes HTML tags and unnecessary strings.
[1631] 3. The server trains the generative AI model to learn specific writing styles and expressions.
[1632] 4. The server extracts the advertiser's formal tone and the affiliate site's friendly tone and trains the model.
[1633] 5. The server learns the advertising notations and annotations and applies them to the generated article.
[1634] 6. The server generates an article stating, "Healthcare product A contains ingredient B and is expected to be effective in improving immunity."
[1635] 7. The user checks the generated article for errors and missing notes.
[1636] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[1637] Example prompt sentence:
[1638] "Use a generative AI model to create an affiliate ad article about healthcare product A based on its features, ingredients, and user reviews. Focus on its immune-boosting benefits. Also, be sure to include legal annotations."
[1639] With the above configuration, the present invention makes it possible to efficiently generate high-quality affiliate articles, reduce legal risks, and reliably publish them.
[1640] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1641] Step 1:
[1642] Data collection
[1643] The server collects the necessary data from advertisers and affiliate sites.
[1644] Input: URL list, crawling settings
[1645] Processing: Use a web crawler tool to retrieve detailed product information, reviews, and past posting data from the specified URL.
[1646] Output: Collected data (HTML file, text data, etc.)
[1647] What it does: The server runs a crawler tool that parses the HTML content of each website to extract data and store it in an internal database.
[1648] Step 2:
[1649] Data Preprocessing
[1650] The server preprocesses the collected data.
[1651] Input: Collected data (HTML files, text data, etc.)
[1652] Processing: Strip HTML tags, remove unnecessary strings, tokenize text, and extract important key phrases.
[1653] Output: Preprocessed text data
[1654] What it does: It uses a regular expression library to remove HTML tags and unnecessary text, a morphological analyzer to break down text into words, and a keyphrase extraction algorithm to extract important keywords.
[1655] Step 3:
[1656] Training generative artificial intelligence models
[1657] The server uses the preprocessed data to train a generative artificial intelligence model (ChatGPT).
[1658] Input: Preprocessed text data
[1659] Processing: Fine-tuning the model to learn specific writing styles and phrasing.
[1660] Output: A trained generative artificial intelligence model
[1661] What it does: It processes batches of data, feeds it into the model, and fine-tunes it by setting optimal hyperparameters.
[1662] Step 4:
[1663] Style and Tone Extraction
[1664] The server extracts the style and tone of the advertiser and affiliate sites.
[1665] Input: Text data from advertiser and affiliate sites
[1666] Processing: Uses a linguistic style analysis library to analyze writing style, tone, and punctuation usage.
[1667] Output: Extracted style and tone data
[1668] Specific operations: Analyzes text data from each site, extracts elements that characterize writing style and tone, and feeds this back to a generative AI model.
[1669] Step 5:
[1670] Annotation Learning
[1671] The server trains a generative artificial intelligence model to generate annotations based on legal requirements.
[1672] Input: Legal annotation database
[1673] Processing: Train the model with standard annotations based on legal requirements.
[1674] Output: A generative AI model that can apply annotations
[1675] What it does: Legal annotations are fed into the model as a dataset, and the model is trained to accurately insert annotations under specific conditions.
[1676] Step 6:
[1677] Article Generation
[1678] The server generates affiliate articles using a generative artificial intelligence model.
[1679] Input: Trained generative artificial intelligence model, collected and preprocessed data
[1680] Processing: Using data, we automatically generate affiliate articles that match your tone and style.
[1681] Output: Generated affiliate article
[1682] What it does: The model generates sentences based on the prompts you provide and automatically adds any necessary legal annotations.
[1683] Step 7:
[1684] Internal Check
[1685] Internally check user-generated articles.
[1686] Input: Generated affiliate article
[1687] Action: Check for clerical errors, misrepresentations, and missing notes.
[1688] Output: Affiliate articles checked or corrected
[1689] What you'll do: Proofread articles using text review tools, conduct team reviews, and manually correct as needed.
[1690] Step 8:
[1691] Article published
[1692] The device publishes the checked articles on the website.
[1693] Input: Checked affiliate article
[1694] Processing: Post articles based on the publishing schedule and record publishing logs.
[1695] Output: Published affiliate articles
[1696] What it does: Uses a scheduler tool to automatically post articles at specified times and records the publishing history.
[1697] (Application example 1)
[1698] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1699] Conventional ad article generation systems have problems with inefficiency and time because the data collection and preprocessing process is complicated and style and tone adjustments are often done manually. Other issues include difficulty in previewing and editing generated articles, and difficulty in automatically posting articles to content management systems. Furthermore, there is a lack of a mechanism for quickly generating articles using smartphones.
[1700] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1701] In this invention, the server includes means for preprocessing collected data and training a generative AI model, means for extracting style and tone from the advertiser's website and affiliated sites and training the generative AI model, means for automatically generating advertising articles based on the collected and learned data, means for appropriately inserting advertising notices and annotations into the generated articles, means for internally checking the generated articles, means for publishing the generated articles on the website after the internal check, means for quickly and easily generating articles using a smartphone, means for displaying the generated articles during preview and editing and enabling editing, and means for automatically posting the generated articles via the API of a content management system, thereby enabling efficient and highly accurate generation and publication of advertising articles.
[1702] "Collected data" refers to information such as product information, reviews, and past posting data obtained from advertisers and affiliated sites.
[1703] "Preprocessing" is the process of removing unnecessary information from collected data and shaping the data.
[1704] A "generative artificial intelligence model" is an artificial intelligence algorithm that can generate new text based on input data.
[1705] "Training" is the learning process used to improve the accuracy of a generative artificial intelligence model based on specific data.
[1706] "Style" refers to characteristics such as writing style, tone, and format.
[1707] "Tone" describes the mood or emotional atmosphere of a piece of writing.
[1708] An "advertising piece" is a piece of writing written to promote a particular product or service.
[1709] "Advertising notation" refers to text or notes that clearly indicate that the item is an advertisement.
[1710] An "annotation" is an explanatory text based on supplementary information or legal requirements inserted within an article.
[1711] "Internal checking" is the process of checking the content of generated articles and correcting errors and inappropriate expressions.
[1712] "Publishing on a website" means posting the generated article on the web so that it can be viewed by general internet users.
[1713] "Using a smartphone" means performing various operations and processes using a smartphone.
[1714] "Preview" is a function that allows you to check the state before the final release.
[1715] "Editing" is the process of correcting or changing the content of the generated article.
[1716] A "content management system" is software for organizing, managing, and delivering content.
[1717] An "API" is an interface that allows applications and services to communicate with each other and utilize their functionality.
[1718] "Automatically posting" means that the system automatically uploads content to a website based on a set schedule or conditions.
[1719] As an embodiment of the present invention, a novel advertising article generation system is described, which preprocesses collected data, trains a generative artificial intelligence model, and ultimately automatically generates and publishes highly accurate advertising articles.
[1720] System Configuration
[1721] The system of the present invention includes the following major components:
[1722] 1. Data Collection Module
[1723] 2. Data Preprocessing Module
[1724] 3. Generative AI Model Training Module
[1725] 4. Style and Tone Extraction Module
[1726] 5. Annotation Learning Module
[1727] 6. Article Generation Module
[1728] 7. Internal Check Module
[1729] 8. Article Publishing Module
[1730] 9. Smartphone Application Module
[1731] 10. Preview and Edit Module
[1732] 11. Auto-posting module
[1733] Hardware and Software Used
[1734] The main hardware and software components of this system include:
[1735] Hardware: High-performance servers and smartphones.
[1736] Software: Python, BeautifulSoup, Scrapy, pandas, transformers, React Native, Content Management Systems (CMS) WordPress and Wix, OpenAI's ChatGPT.
[1737] Processing Description
[1738] 1. Data Collection
[1739] The server collects the necessary data from advertisers and partner sites using APIs and web crawling technologies (BeautifulSoup and Scrapy).
[1740] 2. Data Preprocessing
[1741] The collected data is cleaned using a data preprocessing module. Specifically, HTML tags are removed, unnecessary strings are deleted using regular expressions, and important key phrases and features are extracted. The data is formatted and processed using the pandas library.
[1742] 3. Generative AI Model Training
[1743] The server uses the preprocessed data to train a generative AI model (ChatGPT), which learns specific writing styles and expressions.
[1744] 4. Extracting Style and Tone
[1745] The server extracts style and tone from the text of advertisers and partner sites and feeds this back into a generative artificial intelligence model.
[1746] 5. Annotation Learning
[1747] The server trains the model with annotations based on legal requirements and applies them to the generated advertising articles.
[1748] 6. Article Generation
[1749] Using a generative artificial intelligence model, advertising articles are automatically generated based on collected and learned data.
[1750] 7. Internal Checks
[1751] The generated articles are then internally checked by the user to ensure there are no errors, misrepresentations, or missing notes, and any necessary corrections are made. The quality can also be checked again using the generative AI model.
[1752] 8. Public
[1753] Once the article has been checked, it will be published on the website via the terminal.
[1754] 9. Smartphone use
[1755] Users can quickly and easily create, preview, edit and publish articles using their smartphones.
[1756] 10. Preview and Edit
[1757] It displays a preview of the generated article and allows users to easily make any necessary corrections. It is provided with an interface that uses React Native.
[1758] 11. Auto-posting
[1759] Finally, the generated articles are automatically posted to the website via the content management system's API.
[1760] Specific examples
[1761] For example, generating an advertisement for a health device A involves the following process:
[1762] 1. The server collects information about the ingredients, efficacy, and reviews of health device A.
[1763] 2. The server preprocesses the collected data, removing HTML tags, deleting unnecessary strings, and tokenizing the text.
[1764] 3. The server trains the ChatGPT model to learn specific writing styles and expressions.
[1765] 4. The server extracts the advertiser's formal tone and the partner site's friendly tone and trains the model accordingly.
[1766] 5. The server learns the advertising and annotations and applies them to the generated articles.
[1767] 6. The server generates an article such as "Health device A contains ingredient B and is expected to be effective in improving immunity."
[1768] 7. The user checks the generated article for errors, misrepresentations, and missing notes, and makes corrections as necessary.
[1769] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[1770] Prompt Sentence Examples
[1771] Product name: Health equipment A
[1772] Features: Quiet design, foldable, multi-function display
[1773] As described above, the present invention enables efficient and highly accurate generation and publication of advertising articles.
[1774] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1775] Step 1:
[1776] Data collection
[1777] The server uses APIs and web crawling technologies (BeautifulSoup and Scrapy) to retrieve product details, reviews, and past postings from advertisers and partner sites. The input is the URL or API endpoint of each site, and the output is the retrieved raw data. This allows a wealth of information to be stored in the database used.
[1778] Step 2:
[1779] Data Preprocessing
[1780] The server preprocesses the collected data. Specifically, it uses the pandas library to remove HTML tags, delete unnecessary strings, and extract important key phrases and features. The input is the raw data collected in step 1, and the output is cleansed and formatted data. This process processes the data into a form that is easy for the generative artificial intelligence model to process.
[1781] Step 3:
[1782] Generative AI model training
[1783] The server uses the preprocessed data to train a generative AI model (ChatGPT). In the process, it learns specific writing styles and expressions. Specifically, it fine-tunes the model using the transformers library. The input is the preprocessed data, and the output is a trained generative AI model. This allows the model to generate sentences with even greater accuracy.
[1784] Step 4:
[1785] Style and Tone Extraction
[1786] The server performs text analysis to extract style and tone from the text of advertisers' and partner sites. Specifically, it uses machine learning algorithms and natural language processing (NLP) technology to extract text characteristics. The input is text data from advertisers' and partner sites, and the output is the extracted style and tone features. This ensures that the generated articles are tailored to the characteristics of each site.
[1787] Step 5:
[1788] Annotation Learning
[1789] The server collects specific texts (legal annotations) and trains a generative AI model to learn annotations based on legal requirements. Specifically, it identifies patterns in the annotations and incorporates them into the model. The input is a dataset of legal annotations, and the output is a generative AI model that can insert annotations appropriately. This results in articles that reduce legal risk.
[1790] Step 6:
[1791] Article Generation
[1792] The server uses a generative artificial intelligence model to automatically generate advertising articles based on collected and learned data. Specifically, when a prompt is entered, the corresponding trained model generates an article. The input is the prompt and past posting data, and the output is an automatically generated advertising article. This enables fast and highly accurate article generation.
[1793] Example prompt sentence:
[1794] Product name: Health equipment A
[1795] Features: Quiet design, foldable, multi-function display
[1796] Step 7:
[1797] Internal Check
[1798] The generated article undergoes an internal check by the user. Specifically, it checks for typographical errors, misrepresentations, and missing notes, and makes any necessary corrections. The input is the automatically generated advertising article, and the output is the corrected final version of the article. This allows the article to be published with guaranteed quality.
[1799] Step 8:
[1800] Article published
[1801] Once the article has been checked, it is published on the website via the terminal. Specifically, the article is automatically posted via the API of the content management system (CMS). The input is the final, corrected version of the article, and the output is the published article. This allows the article to be published efficiently on the web, saving the user time and effort.
[1802] Step 9:
[1803] Smartphone use
[1804] Users use their smartphones to create, preview, edit, and publish articles. Specifically, articles are created and edited using a dedicated application. The input is the requirements and prompts specified by the user, and the output is the edited article. This allows article creation and management to be done anywhere, anytime.
[1805] Step 10:
[1806] Preview and Edit
[1807] A preview of the generated article is displayed, allowing the user to easily make any necessary corrections. Specifically, the interface uses React Native, and edits are reflected in real time. The input is the automatically generated article, and the output is the article corrected by the user. This further improves the quality of the final article.
[1808] Step 11:
[1809] Auto-post
[1810] Finally, the generated articles are automatically posted to the website via the content management system's API. The server uploads the articles to the website based on the set schedule and conditions. The input is the checked, final version of the article, and the output is the published article and its publication log. This minimizes user effort and efficiently publishes articles through an automated process.
[1811] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1812] The following description is provided as an embodiment of the present invention. This system collects data from advertisers and affiliate sites and generates highly accurate affiliate articles using a generative artificial intelligence model (e.g., ChatGPT). Furthermore, by combining it with an emotion engine that recognizes user emotions, the system can be customized according to the user's emotions.
[1813] System Configuration
[1814] The system mainly consists of the following components:
[1815] 1. Data Collection Module
[1816] 2. Data Preprocessing Module
[1817] 3. Generative AI Model Training Module
[1818] 4. Style and Tone Extraction Module
[1819] 5. Annotation Learning Module
[1820] 6. Emotion Engine Module
[1821] 7. Article Generation Module
[1822] 8. Internal Check Module
[1823] 9. Article Publishing Module
[1824] Data collection
[1825] First, the server collects the necessary data from the advertiser and affiliate sites. The collected data includes detailed product information, review articles, and past posting data from the affiliate sites. The server obtains the data from each site using web crawling technology.
[1826] Data Preprocessing
[1827] Next, the server preprocesses the collected data. This step involves data cleaning to remove unnecessary information, such as removing HTML tags, deleting unnecessary strings using regular expressions, tokenizing the text, and extracting important key phrases and features.
[1828] Training generative artificial intelligence models
[1829] Using the preprocessed data, the server trains a generative AI model (ChatGPT), which fine-tunes the model based on the initial dataset to learn specific writing styles and expressions.
[1830] Style and Tone Extraction
[1831] The server extracts the style and tone of the advertiser's and affiliate's sites and trains the generative AI model. During this process, the server analyzes each site's writing style, tone, and punctuation usage. This data is then fed back into the generative model to generate articles that match the characteristics of each site.
[1832] Annotation Learning
[1833] The server trains a generative AI model to create appropriate ad descriptions and annotations. This involves collecting standard annotations based on legal requirements and teaching the model necessary annotation information, such as "This product is not a medicine."
[1834] Emotion Engine
[1835] The server uses an emotion engine to analyze the user's emotions in real time. The emotion engine recognizes the user's emotional state based on user input and various sensor data (e.g., facial recognition and voice analysis). Based on the results, the style and tone of the generated articles can be adjusted.
[1836] Article Generation
[1837] The server automatically generates affiliate articles based on the collected and learned data. The generative AI model emphasizes the product's benefits and features and generates text that matches the tone and style. For example, if the user is in a positive emotional state, text with a more optimistic tone will be generated. Annotations based on legal requirements are also automatically inserted into the generated articles.
[1838] Internal Check
[1839] The generated articles are then checked internally by a user (an internal checker). A checklist is used to check for typographical errors, misrepresentations, and missing notes. During this process, the quality can be checked again using the generative AI model. If necessary, the user can make manual corrections.
[1840] Article Publication
[1841] Finally, articles that have passed the check are published on the website via the terminal, which automatically posts the articles based on the publication schedule and records the publication log.
[1842] Specific examples
[1843] For example, generating an affiliate article for healthcare product A involves the following process:
[1844] 1. The server collects information about the ingredients, efficacy, and reviews of healthcare product A.
[1845] 2. The server preprocesses the collected data, removing HTML tags, deleting unnecessary strings, and tokenizing the text.
[1846] 3. The server trains the ChatGPT model to learn specific writing styles and expressions.
[1847] 4. The server extracts the advertiser's formal tone and the affiliate site's friendly tone and trains the model accordingly.
[1848] 5. The server learns the advertising and annotations and applies them to the generated articles.
[1849] 6. The server uses an emotion engine to analyze the user's emotions and adjust the tone of the article accordingly. For example, if the user is excited, the article will be generated with a more enthusiastic tone.
[1850] 7. The user reviews the generated article, checking for errors, misrepresentations, and missing notes, and makes corrections as necessary.
[1851] 8. The device will automatically publish the checked articles on the website and record the publishing log.
[1852] The above system makes it possible to efficiently generate high-quality affiliate articles that respond to user sentiment and publish them while reducing legal risks.
[1853] The processing flow will be explained below.
[1854] Step 1:
[1855] The server collects data from advertiser and affiliate sites. Specifically, it uses a crawler to retrieve the HTML of web pages from the specified URLs and stores it in a database. The retrieved data includes titles, body text, meta tags, image URLs, etc.
[1856] Step 2:
[1857] The server preprocesses the collected data, removing HTML tags and extracting only the text data. Regular expressions are also used to remove unnecessary symbols and spaces. Next, the text is tokenized and techniques such as TF-IDF and Word2Vec are used to extract important key phrases and features.
[1858] Step 3:
[1859] The server uses the preprocessed data to train a generative AI model, based on the initial ChatGPT model, which is then fine-tuned to fit specific writing styles and themes. This process uses high-quality samples from the collected data as training data.
[1860] Step 4:
[1861] The server extracts the style and tone of the advertiser and affiliate sites, uses analysis algorithms to identify the unique writing style and tone patterns of each site, and feeds these features back into the generative AI model.
[1862] Step 5:
[1863] The server trains a generative AI model to create appropriate annotations. Standard annotations and warnings based on legal requirements are collected and used as training data for the model. This training gives the model the ability to automatically insert appropriate annotations when generating articles.
[1864] Step 6:
[1865] The server starts an emotion engine that analyzes the user's emotions in real time. The emotion engine detects emotions from the user's input (text, voice, facial recognition data, etc.) and records the results in a database.
[1866] Step 7:
[1867] The server adjusts the style and tone of the generated article based on the user's emotional data analyzed by the emotion engine. For example, if the user is in a positive emotional state, the server generates an optimistic tone of text.
[1868] Step 8:
[1869] Affiliate articles are automatically generated based on the data collected and learned by the server. As a specific example, when generating an article for "Healthcare Product A," the article is generated by incorporating product features and user reviews and adjusting the tone.
[1870] Step 9:
[1871] User-generated articles are internally checked using a dedicated checklist to ensure there are no errors, misrepresentations, or missing notes. Again, during this process, the generative AI model can be used to check the quality.
[1872] Step 10:
[1873] The device publishes the checked articles on the website. Articles are automatically posted according to the publication schedule. The URL and related information after publication are recorded in a database and stored as data for analysis.
[1874] Through the above steps, high-quality affiliate articles that respond to user sentiment can be efficiently generated and published with reduced legal risks.
[1875] Example 2
[1876] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1877] Conventional affiliate article generation systems have had issues with preprocessing collected data, matching the style and tone of generated articles, appropriately inserting advertisements and annotations, and customizing articles to reflect user sentiment. In particular, generating content based on user sentiment is difficult, and generated articles often do not reflect the user's sentiment. There is also a need for more efficient internal checks of generated articles.
[1878] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1879] In this invention, the server includes means for preprocessing collected data and training a generative artificial intelligence model, means for extracting style and tone from the advertiser's site and affiliate site and training the generative artificial intelligence model, and means for the generative artificial intelligence model to analyze user emotions and adjust the style and tone of articles. This enables the generation of more personalized affiliate articles in response to user emotions. Furthermore, using efficiently preprocessed data allows for the efficient generation of high-quality articles, contributing to the efficiency of internal checks.
[1880] "Collected data" refers to information such as detailed product information, review articles, and past posting data obtained from advertiser and affiliate sites using web crawling technology.
[1881] "Preprocessing" refers to a series of processes performed on collected data, such as removing HTML tags, deleting unnecessary strings, tokenizing text, and extracting key phrases and features.
[1882] A "generative artificial intelligence model" is an artificial intelligence algorithm that generates sentences based on training data, and a specific example is ChatGPT.
[1883] "Training" is the process of using preprocessed data to teach a generative artificial intelligence model a particular style or method of expression.
[1884] "Advertiser's Site and Affiliate Site" refers to a website that provides detailed product information and review articles, and includes affiliate marketing content.
[1885] "Style and tone" refers to the writing style, tone, punctuation, and other characteristics of a particular site or document.
[1886] "Learning" is the process by which a generative AI model analyzes data to acquire a particular style or tone, and then applies the results to the model.
[1887] "User emotion" refers to the emotional state analyzed based on the user's input information (text, voice, facial image, etc.), and includes positive, negative, excited, etc.
[1888] "Internal checking" is a process in which the generated articles are checked for errors, misrepresentations, missing notes, etc., and is carried out by in-house checkers.
[1889] "Advertising statements and annotations" refers to standard legally required annotations and advertising statements, such as warnings such as "This product is not a medicine."
[1890] "Publishing on a website" means posting the checked article on a website based on a specific publishing schedule and recording a publishing log.
[1891] To implement this invention, the following system must be constructed. This system uses data collected from advertisers and affiliate sites and a generative artificial intelligence model to automatically generate highly accurate affiliate articles. Furthermore, it is possible to analyze user sentiment and adjust the style and tone of the resulting articles.
[1892] System configuration
[1893] The system consists of the following main components:
[1894] 1. Data Collection Module
[1895] 2. Data Preprocessing Module
[1896] 3. Generative AI Model Training Module
[1897] 4. Style and Tone Extraction Module
[1898] 5. Annotation Learning Module
[1899] 6. Emotion Engine Module
[1900] 7. Article Generation Module
[1901] 8. Internal Check Module
[1902] 9. Article Publishing Module
[1903] Hardware and software used
[1904] The implementation of this system uses the following hardware and software:
[1905] Server: A server with high-performance processing power is used to process collected data, train generative AI models, analyze emotions, and perform other heavy workloads.
[1906] Generative AI models, such as ChatGPT, are used to generate text and adjust style and tone.
[1907] Emotion engine: An engine for analyzing user emotions, using sensors such as facial recognition and voice analysis.
[1908] Program processing
[1909] The server first collects data from advertiser and affiliate sites, including product information, reviews, and past posting data, and the collected data is obtained using web crawling technology.
[1910] The server then preprocesses the collected data, removing HTML tags, eliminating unnecessary strings, tokenizing the text, and extracting important key phrases and features.
[1911] Based on the preprocessed data, the server trains a generative AI model (ChatGPT), fine-tuning the model based on the initial dataset to learn specific writing styles and expressions.
[1912] The server then extracts the style and tone of the advertiser and affiliate sites and trains them into a generative AI model. Each site's writing style, tone, and punctuation usage are analyzed and fed back to the model.
[1913] The server also trains a generative AI model to create appropriate ad labeling and annotations, collecting standard annotations based on legal requirements and training the model to apply them to the generated articles.
[1914] Additionally, the server uses an emotion engine to analyze users' emotions in real time, recognizing their emotional state based on user input and various sensor data (facial recognition and voice analysis), and adjusting the style and tone of the generated articles accordingly.
[1915] As a concrete example, the prompt text to automatically generate an affiliate article for healthcare product A is as follows:
[1916] "Generate an article with an optimistic tone based on the ingredients, efficacy, and reviews of healthcare product A when the user is in a positive emotional state."
[1917] Finally, the generated article is checked internally by the user (an in-house checker). A checklist is used to check for typographical errors, misrepresentations, and missing notes, and manual corrections are made as necessary. The completed article is published on the website via the terminal, and a publication log is recorded.
[1918] summary
[1919] This system allows us to efficiently generate high-quality affiliate articles that reflect user sentiment and publish them while reducing legal risks.
[1920] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1921] Step 1:
[1922] Data collection
[1923] The server collects data from advertiser and affiliate sites. It uses web crawling technology to obtain HTML content based on the URL list specified as input. Specifically, it collects detailed product information, review articles, and past posting data from advertiser and affiliate sites. The output is raw data, which can be in a wide variety of formats.
[1924] Step 2:
[1925] Data Preprocessing
[1926] The data collected by the server is preprocessed. The input is the raw data collected in step 1. Specifically, regular expressions are used to remove HTML tags and extract only the text. Unnecessary strings (e.g., advertising banner text) are also filtered. The text is then tokenized and split into words. Important key phrases are then extracted using algorithms such as TF-IDF and Word2Vec. The output is preprocessed, clean text data.
[1927] Step 3:
[1928] Training generative artificial intelligence models
[1929] The server uses the preprocessed data to train a generative AI model (ChatGPT). The input is preprocessed, clean text data. Specifically, the initial dataset is fed into the generative AI model, and fine-tuned to learn specific writing styles and expressions. The output is a trained generative AI model.
[1930] Step 4:
[1931] Style and Tone Extraction
[1932] The server analyzes the style and tone of advertiser and affiliate sites and trains a generative AI model. The input is the text data of each site. Specific processing involves analyzing writing style (frequency of use of punctuation at the end of sentences, use of honorifics, etc.) and tone (formal or informal, optimistic or pessimistic, etc.). The results of these analyses are then fed back to the generative AI model, enabling it to generate text that matches the characteristics of each site. The output is a model that has learned the characteristics of style and tone.
[1933] Step 5:
[1934] Annotation Learning
[1935] The server trains a generative AI model on appropriate methods for advertising labeling and annotations. The input is standard annotations based on legal requirements (e.g., "This product is not a medicine"). Specific processing involves collecting standard annotations and training the generative AI model so that they can be applied when needed. The output is a model that can insert annotations.
[1936] Step 6:
[1937] Emotion Engine
[1938] The server uses an emotion engine to analyze the user's emotions in real time. The input is the user's input information (text, voice, facial image, etc.). Specifically, the process uses sensors such as facial recognition and voice analysis to recognize the user's emotional state (positive, negative, excited, etc.). Based on the results, the style and tone of the generated article are adjusted. The output is text whose tone and style have been adjusted to reflect the user's emotional state.
[1939] Step 7:
[1940] Article Generation
[1941] The server automatically generates affiliate articles based on collected and learned data. The input is a model that has learned style and tone characteristics that reflect emotional states, as well as requirements such as product information. A generative AI model is used to generate text that highlights the product's features and benefits. Annotations based on legal requirements are also automatically inserted. The output is the completed affiliate article.
[1942] Step 8:
[1943] Internal Check
[1944] The user (an internal reviewer) checks the generated article. The input is the generated affiliate article. Specifically, a checklist is used to check for errors, misrepresentations, and missing notes. Manual corrections may be made as necessary. The output is the final version of the article after corrections have been completed.
[1945] Step 9:
[1946] Article Publication
[1947] The terminal publishes the checked article on the website. The input is the final version of the article. Specific processing involves setting a publishing schedule and automatically posting the article based on the specified date and time. A publishing log is also recorded. The output is the published article and its publishing log.
[1948] (Application example 2)
[1949] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1950] Conventional affiliate article generation systems have difficulty generating advertising content that takes user emotions into account, making it impossible to provide ads that are adapted to user emotions. Furthermore, it is difficult to properly reflect the style and tone of the advertiser or affiliate site, making it impossible to provide an effective advertising experience for users.
[1951] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for preprocessing collected data and training a generative AI model, means for extracting style and tone from the advertiser's website and affiliate site and training the generative AI model, means for automatically generating affiliate articles based on the collected and learned data, means for appropriately inserting advertising notices and annotations into the generated articles, means for internally checking the generated articles, means for publishing the articles generated after the internal check on the website, and means for recognizing user emotions and adjusting the tone and style of the generated advertising content based on the emotions. This makes it possible to provide effective advertising content that corresponds to the user's emotions and generate and publish high-quality affiliate articles while reflecting the style and tone of the advertiser and affiliate site.
[1952] "Collected data" refers to information such as detailed product information, review articles, and past posting data obtained from advertisers and affiliate sites.
[1953] "Preprocessing" refers to the process of cleaning the collected data, removing unnecessary information, and extracting important key phrases and features using regular expressions.
[1954] A "generative artificial intelligence model" is an artificial intelligence model that learns specific writing styles and methods of expression based on given data, and generates new sentences based on the results of that learning.
[1955] "Training" refers to the learning process of using preprocessed data to improve the performance of a generative artificial intelligence model.
[1956] "Style and tone extraction" refers to the process of analyzing characteristics such as writing style and punctuation usage from advertiser websites and affiliate sites, and training a generative AI model to learn these characteristics.
[1957] An "affiliate article" refers to a piece of writing that provides information about a specific product or service and is intended to encourage users to purchase that product or service.
[1958] "Blank and annotation insertion" refers to the process of automatically adding standard legally required annotations to generated articles.
[1959] "Internal Check" refers to the internal review process to ensure that the content of the generated article is accurate and that there are no misrepresentations or omissions.
[1960] "Publication" refers to making the generated article publicly available on a website or other media.
[1961] "Emotion recognition" refers to the process of analyzing a user's emotional state using user input and sensor data (e.g., facial recognition and voice analysis).
[1962] "Tone and style adjustment" refers to the process of changing the tone and presentation of generated text based on the perceived user sentiment.
[1963] The present invention provides an emotion-based advertising content generation system, which includes means for preprocessing collected data and training a generative AI model, means for extracting style and tone from advertiser and affiliate sites and training the generative AI model, means for automatically generating affiliate articles based on the collected and learned data, means for appropriately inserting advertising notices and annotations into the generated articles, means for internally checking the generated articles, means for publishing the generated articles on a website after the internal check, and means for recognizing user emotions and adjusting the tone and style of the generated advertising content based on the emotions.
[1964] Program processing explanation
[1965] The server uses an emotion recognition module to analyze the user's emotions in real time. This module utilizes facial recognition technology using the smartphone camera and TensorFlow. Specifically, it captures the user's face and analyzes their emotions using a trained emotion recognition model (e.g., CNN model).
[1966] The server then uses web crawling techniques to collect the required data from the advertiser's website and affiliate sites, using Python web scraping libraries (e.g., BeautifulSoup and Requests), and preprocesses the collected data by removing HTML tags and extracting text information using regular expressions.
[1967] Based on the preprocessed data, the server trains a generative AI model (e.g., GPT-2). During this process, the collected data is analyzed for its writing style and tone, and the model is trained to learn from it. In particular, the model extracts features of the advertiser's formal writing style and the affiliate site's friendly tone.
[1968] The generated advertisements are automatically inserted with ad captions and annotations to comply with legal requirements. In this process, a generative model is trained on standard annotations collected in advance, and the resulting advertisement captions and annotations are added to the generated articles.
[1969] After generation is complete, the server performs an internal check to ensure there are no errors, misrepresentations, or missing notes. This includes a second quality check using the generative AI model. Once the check is complete, the article is published to the website via the terminal. Here, posts are automatically made based on the publishing schedule, and a publishing log is recorded.
[1970] Examples of concrete examples and prompts
[1971] For example, generating an affiliate article for healthcare product A involves the following process:
[1972] The server uses an emotion recognition module to analyze whether the user is in a positive emotional state and generates an affiliate article with a positive tone. The generated article uses the following prompt:
[1973] "Preprocess this data, normalize it, filter it, tokenize it."
[1974] "Generate ad text based on sentiment and preprocessed data."
[1975] "Please display the generated ad text on your smartphone screen."
[1976] In this way, it becomes possible to provide effective advertising content that matches the user's emotions, and high-quality affiliate articles that reflect the style and tone of the advertiser and affiliate site are generated.
[1977] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1978] Step 1:
[1979] Recognize user emotions.
[1980] The server captures the user's face in real time using the smartphone camera. The captured image is input to an emotion recognition model using TensorFlow to analyze the user's emotional state (e.g., positive, negative, neutral). The analysis result is passed to the next processing step.
[1981] Step 2:
[1982] Data collection.
[1983] The server uses web scraping technology to collect the necessary data from advertiser and affiliate sites. The collected data includes product details, reviews, past posting data, etc. This data is pre-processed in the next step.
[1984] Step 3:
[1985] Data preprocessing.
[1986] The server preprocesses the collected data. Specifically, it uses Python's BeautifulSoup and Requests to remove HTML tags and regular expressions to remove unnecessary characters. It also tokenizes the text and extracts important key phrases. The preprocessed data is used to train generative artificial intelligence models.
[1987] Step 4:
[1988] Training generative artificial intelligence models.
[1989] The server uses the preprocessed data to train a generative artificial intelligence model (e.g., GPT-2). During this training, the model learns the style and tone of the advertiser's website and affiliate site. The trained model is then used in the next step.
[1990] Step 5:
[1991] Generating advertising articles.
[1992] The server uses the trained generative artificial intelligence model to generate advertising articles based on the user's emotional state. For example, if the user is in a positive emotional state, an article with a positive tone is generated. The generated advertising article is then processed in the following steps.
[1993] Step 6:
[1994] Inserting advertising labels and annotations.
[1995] The server automatically inserts appropriate advertisements and annotations into the generated articles by training a generative AI model with standard annotations based on legal requirements, resulting in articles that mitigate legal risks.
[1996] Step 7:
[1997] Internal check.
[1998] The generated articles are then subjected to an internal check by the server, where a quality check is carried out again using the generative AI model to check for errors, misrepresentations, and missing notes. If necessary, an in-house checker manually corrects the errors.
[1999] Step 8:
[2000] Publication of article.
[2001] The device publishes articles on the website after internal checks are complete. Publication is done automatically based on a pre-set schedule and a publication log is recorded, streamlining the article publishing process.
[2002] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2003] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2004] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2005] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2006] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2007] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2008] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2009] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2010] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2011] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2012] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2013] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2014] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Univ...
Claims
1. a means for preprocessing the collected data and training a generative artificial intelligence model; and A means for extracting style and tone from advertiser websites and affiliate sites and training the generative AI model; A means for automatically generating affiliate articles based on collected and learned data; A means of properly inserting advertisements and annotations into the generated articles; a means for performing internal checks on the generated articles; A means of publishing the articles generated after internal checks on the website; A system including:
2. The system of claim 1 , further comprising: crawling the collected data and removing HTML tags to extract text information.
3. 10. The system of claim 1, further comprising means for reviewing the generated affiliate articles to check for typographical errors, misrepresentations, and omissions.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A