Service delivery system and AI model generation method
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ASCON
- Filing Date
- 2025-10-15
- Publication Date
- 2026-04-24
AI Technical Summary
Existing advertising-related service systems require manual work to compare advertisements with manuscript data, leading to high labor costs and expenses.
An advertising-related service system utilizing machine learning to determine consistency between advertisements and manuscript data, minimizing manual work by identifying proofreading portions through machine learning algorithms.
Reduces labor costs and manual effort by automating the proofreading process, ensuring efficient matching of advertisements with manuscript data.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an advertising-related service system that provides services related to advertisements such as flyers, and user-side advertising equipment. More specifically, the present invention relates to an advertising-related service providing system that determines and notifies portions of advertisements that need proofreading, and user-side advertising equipment that can be connected to the Internet on the advertising-related service providing system side that provides services related to advertisement creation. [Background technology]
[0002] One such advertising-related service providing system collates advertisements such as flyers with manuscript data to find any portions that do not conform to the manuscript (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2004-133734 Summary of the Invention [Problem to be solved by the invention]
[0004] The system described in Patent Document 1 has the drawback that manual work is required to compare advertisements such as flyers with their manuscript data, resulting in high labor costs and other expenses.
[0005] The present invention was conceived in view of the above circumstances, and its purpose is to provide an advertising-related service providing system and user-side advertising equipment that can minimize the manual work required to match advertisements with their manuscript data. [Means for solving the problem]
[0006] In one aspect of the present invention, there is provided an advertisement-related service providing system that provides a service related to advertisement creation, comprising: an advertisement data receiving means for receiving advertisement data; manuscript data receiving means for receiving manuscript data used in creating the advertisement; a proofreading portion determining means for determining consistency between the advertisement and the manuscript data based on the advertisement data received by the advertisement data receiving means and the manuscript data received by the manuscript data receiving means, and determining portions of the advertisement that need proofreading; The proofreading portion determining means determines the consistency by reflecting the learning results obtained by the proofreading machine learning means in machine learning for determining the consistency between the advertisement and the manuscript data.
[0007] With this configuration, the learning results obtained by the proofreading machine learning means are reflected to determine the consistency between the advertisement and the manuscript data, and the parts of the advertisement that need to be proofread are identified, thereby minimizing manual work as much as possible.
[0008] Preferably, the advertisement data includes data of a plurality of advertisement products; the manuscript data includes data specifying advertising content for each of a plurality of products to be advertised, The machine learning performed by the calibration machine learning means includes value set classification identification machine learning for classifying and identifying the advertisement data received by the advertisement data receiving means for each single value set, The calibration portion determination means A value group classification identification process that classifies and identifies the advertisement data by single value groups by reflecting the learning results of the value group classification identification machine learning; a manuscript corresponding portion identification process for identifying which portion of the advertisement content in the manuscript data each of the single value sets identified by the value set classification identification process corresponds to; A consistency determination process may be performed in which each single value set identified by the value set classification identification process is matched with the advertising content identified by the manuscript corresponding portion identification process to determine the consistency.
[0009] Preferably, the consistency determination process may determine the consistency between the document data and a character image of each single value set identified by the value set-by-value set classification and identification process.
[0010] Preferably, the machine learning performed by the proofreading machine learning means includes text data conversion machine learning for converting character images of each single value set identified by the value set-by-value set classification identification process into text data, The consistency determination process may be such that the character image is converted into text data by reflecting the learning results of the text data conversion machine learning, and then the consistency between the text data and the manuscript data is determined.
[0011] Preferably, the machine learning performed by the calibration machine learning means further includes product feature learning for learning the features of the product based on data of an original image of an actual product corresponding to the advertised product; The system further includes an association storage means for storing the product features learned by the product feature learning in association with identification information for identifying the product, the manuscript data is data in which the advertisement content is associated with the identification information of each of a plurality of products to be advertised, The manuscript corresponding portion identification process may identify identification information stored in the correspondence storage means in correspondence with the product characteristics used to identify the advertised product, and identify the advertising content associated with the identification information from the manuscript data.
[0012] Preferably, the proofreading machine learning means performs machine learning specialized for each group by performing machine learning based on data grouped by predetermined attributes, The proofreading portion determining means may determine the consistency by reflecting the results of machine learning specialized for each group by the proofreading machine learning means.
[0013] In another aspect of the present invention, there is provided a user-side advertising device connectable to an advertisement-related service providing system that provides advertisement creation services, the user-side advertising device comprising: an advertisement providing means for providing an advertisement to the advertisement-related service providing system; manuscript data providing means for providing manuscript data used in creating the advertisement to the advertisement-related service providing system; a proofreading portion determining means provided on the advertisement-related service providing system side determines the consistency between the advertisement and the manuscript data based on the advertisement data provided by the advertisement providing means and the manuscript data provided by the manuscript data providing means, and determines the portions of the advertisement that need proofreading, and a receiving means receives a service based on the determination result from the advertisement-related service providing system side; The proofreading portion determining means determines the consistency by reflecting the learning results obtained by the proofreading machine learning means in machine learning for determining the consistency between the advertisement and the manuscript data.
[0014] With this configuration, the learning results of the proofreading machine learning means are reflected to determine the consistency between the advertisement and the manuscript data, and the parts of the advertisement that need to be proofread are identified, and services based on the results of this determination can be enjoyed, thereby allowing for services that minimize human work and reduce labor costs. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a diagram illustrating an overall configuration of a service providing system. [Figure 2] FIG. 2 is a diagram illustrating the hardware configuration of a service providing system and an advertisement-related service providing system. [Figure 3] FIG. 2 is a diagram illustrating a hardware configuration of a user terminal. [Figure 4] (A) is a flowchart showing the processing of an advertiser's user terminal, (B) is a flowchart showing the advertising effectiveness measurement processing, and (C) is a flowchart showing the main routine of big data machine learning processing. [Figure 5] (A) is a flowchart showing a subroutine program of machine learning for price fluctuation prediction, and (B) is a flowchart showing a subroutine program of machine learning for generating optimized advertisements. [Figure 6] 1A is a flowchart showing a subroutine program of supervised learning for generating optimized advertisements, and FIG. 1B is a flowchart showing a subroutine program of reinforcement learning for generating optimized advertisements. [Figure 7] 10 is a flowchart of a price fluctuation prediction process. [Figure 8] FIG. 1 is a diagram showing the overall configuration of a calibration machine learning means. [Figure 9] 10 is a flowchart showing a main routine of a flyer proofreading portion determination process. [Figure 10] This is an explanatory diagram explaining the flow of the flyer proofreading part identification process, where (A) is a diagram showing a flyer, (B) is a diagram identifying multiple value sets on the flyer for each single value set, (C) is a diagram of a single value set extracted, (D) is a diagram of a product image and a character image separated, and (E) is a diagram comparing the AI data created by converting the character image into text data using OCR with the manuscript data and outputting the discrepancies. [Figure 11] (A) is a flowchart showing a subroutine program for the reference storage unit selection process, (B) is a flowchart showing a subroutine program for the single value set identification process, (C) is a flowchart showing a subroutine program for the product image identification process, and (D) is a flowchart showing a subroutine program for the target product image extraction process. [Figure 12] (A) is a diagram showing a subroutine program for the corresponding label extraction process, (B) is a flowchart showing a subroutine program for the character position identification process, (C) is a flowchart showing a subroutine program for the AI data creation process using OCR, and (D) is a flowchart showing a subroutine program for the word classification process. [Figure 13](A) is a flowchart showing a subroutine program for processing a comparison between crawler collected data and manuscript data, (B) is a flowchart showing a subroutine program for processing a comparison between AI data and manuscript data, and (C) is a flowchart showing a subroutine program for processing the storage of price data. [Figure 14] 10 is a flowchart showing the processing of a user terminal of an advertisement production company. [Figure 15] (A) is a flowchart showing a subroutine program for the learning process using training data in the practical stage, and (B) is a diagram showing the learned models grouped and stored in the learned storage section using training data in the practical stage. DETAILED DESCRIPTION OF THE INVENTION
[0016] The service providing system, advertisement-related service providing system, user-side equipment, and user-side advertising equipment in this embodiment will be described in detail with reference to the drawings. Hereinafter, database will be referred to as DB, and artificial intelligence will be referred to as AI. The service providing system in this embodiment utilizes machine learning by AI, and the overall system configuration will be described with reference to FIG.
[0017] This service provision system has the function of having a machine learning means 34 perform machine learning using training data that utilizes big data, inputting big data into the trained model 50 generated thereby, predicting price fluctuations and conducting futures trading based on the price fluctuation predictions, and generating optimized advertisements and providing (transmitting) them to advertisers 39 and advertisement creation companies 31.
[0018] The collected big data is input and stored in the big data DB 36 in three ways: by inputting flyers 27 provided by an advertising production company 31 and the manuscript data 24 used to create them; by inputting data collected by a crawler 37 crawling the Internet (e.g., weather, natural disasters, stock prices, exchange rates, interest rates, economic trends, income fluctuations, geopolitical risks, etc.); and by inputting data for measuring advertising effectiveness provided by advertisers 30 (e.g., POS data, number of flyer coupons used, number of flyer inquiries, etc.).
[0019] Upon receiving a request from advertiser 30 to create an advertisement, advertising creation company 31 creates manuscript data 24 for creating flyer 27 as an example of an advertisement, and creates flyer 27 based on the manuscript data 24. When this manuscript data 24 and the data of flyer 27 are input into advertising-related service providing means 35, advertising-related service providing means 35 uses AI to match the data of manuscript data 24 and flyer 27, finds parts of flyer 27 that differ from manuscript data 24, and provides them to advertising creation company 31 as advertisement proofreading points.
[0020] By being provided with the advertisement proofreading parts, advertisement creation company 31 will have an incentive to provide flyer 27 and its manuscript data 24 to advertisement-related service providing means 35, and will take the initiative to provide flyer 27 and its manuscript data 24 to advertisement-related service providing means 35. In this way, the data related to the advertisement provided by advertisement creation company 31, particularly the product data including the price of each product (product name, volume, price, sales limit, etc.), is accumulated in big data DB 36 as price data 41.
[0021] This service providing system is provided with an advertising effectiveness measurement means 38. The advertising effectiveness measurement means 38 measures advertising effectiveness based on effectiveness data provided by the advertiser 30 (for example, POS data, the number of times a flyer coupon is used, the number of inquiries about the flyer, etc.), and provides the results of this effectiveness measurement to the advertiser 30. As a result, the advertiser 30 receives an incentive to provide effectiveness data, and is encouraged to provide it on their own initiative. The provided effectiveness data 42 is stored in the big data DB 36.
[0022] The crawler 37 periodically patrols the Internet to collect various data necessary for predicting price fluctuations and creating optimized advertisements. The collected data is stored in the big data DB 36 as crawler collected data 40.
[0023] When the machine learning means 34 performs supervised learning, for example, a large amount of data set consisting of input information (vector x) and correct answer information y extracted from various data stored in the big data DB 36 is input to the machine learning means 34 as training data. The machine learning means 34 performs supervised learning based on the training data to generate a trained model 50. In the case of supervised learning, the function ci(x) (ci:x→y) that maps the input x to the correct answer y is learned, and therefore the trained model includes the function ci. Note that the machine learning performed by the machine learning means 34 is not limited to supervised learning, and may be any type of unsupervised learning such as model estimation or pattern mining (data mining), or semi-supervised learning, reinforcement learning, deep learning, or the like, which is an intermediate method between supervised learning and unsupervised learning.
[0024] By inputting input data extracted from the big data DB 36 into the generated trained model 50, it becomes possible to predict, for example, price fluctuations as a supervised learning regression problem. The prediction results are provided to users (e.g., advertisers 30, advertisement creation companies 31, etc.) as a cloud service. Furthermore, various services (e.g., futures trading) that utilize the price fluctuation prediction results may be provided.
[0025] Next, the hardware configuration of the service providing system and the advertisement-related service providing system will be described with reference to Fig. 2. Services provided by the service providing system and the advertisement-related service providing system are provided to users (advertisers 30, advertisement creation companies 31, etc.) from a cloud 43 via the Internet. For this purpose, a data center 44 in which a large number of servers 14 are installed is provided, and terminals (personal computers, smartphones, etc.) installed at the advertisers 30, advertisement creation companies 31, etc. are connected to the data center 44 via the Internet so that information can be communicated.
[0026] The server 14 is composed of a CPU (Central Processing Unit) 10 as the control center, a RAM (Random Access Memory) 9 that functions as the work area of the CPU 10, a ROM (Read Only Memory) 11 that stores data and programs, storage means such as an HDD (hard disk drive) 12, an input operation unit 7 such as a display and keyboard, a communication unit 5, a display unit 6, an interface 8, a bus 13, and various other hardware.
[0027] Server 14 uses a typical von Neumann-type computer, but it can also use a neural net processor (NNP). The NNP chip is equipped with a large number of "artificial neurons" modeled after real neurons, and each neuron connects with each other via a network. It can also use a quantum computer that employs the "quantum annealing method." This can significantly reduce the time required for optimization calculations in machine learning.
[0028] 3 shows the hardware configuration of the user terminal 74. In a plurality of stores 73 (only one is shown in the drawing) operated by the advertiser 30, products handled by the advertiser 30 are displayed and sold. Each store 73 is equipped with a POS (Point Of Sale) system, which collects POS data such as the price of each product, the number of units sold, the date and time of sale, as well as the weather at the time of sale, the customer's age group, gender, purchased products, point management, etc., and stores the data in the user terminal 74 of each store 73.
[0029] Each user terminal 74 of the advertiser 30 is connected by an in-house LAN (Local Area Network) 84. The user terminal 74 of each store 73 is connected to the in-house LAN 84 via the Internet or the like. POS data accumulated in the user terminal 74 of each store 73 is transmitted to the user terminal 74 of the advertiser 30 via the Internet and the in-house LAN 84. The user terminal 74 of the advertiser 30 tally up the transmitted POS data and periodically transmits it to the data center 44 as effectiveness data for measuring advertising effectiveness.
[0030] Each user terminal 74 of the advertisement creation company 31 is also connected by an in-house LAN 84. The in-house LAN 84 is connected to the data center 44 via the Internet or the like. The user terminal 74 of the advertisement creation company 31 transmits manuscript data for creating flyers to the data center 44, and also transmits data of flyers created based on the manuscript data to the data center 44.
[0031] 3 shows the hardware configuration of the user terminal 74 of each store 73, the user terminal 74 of the advertiser 30, and the user terminal 74 of the advertisement creation company 31. This hardware configuration is the same as the hardware configuration of the server 14 shown in FIG. 2, and a repeated explanation will be omitted here.
[0032] Next, referring to FIG. 4(A), the processing of the user terminal 74 of the advertiser 30 will be described. In step S (hereinafter simply referred to as "S") 150, it is determined whether POS data has been received from the user terminal 74 of each store 73. If not, control proceeds to S152. If received, the received POS data is classified and stored as effectiveness data in S151. Next, in S152, it is determined whether the time has come to transmit the effectiveness data to the data center 44. If the time has not yet come, control proceeds to S155. If the time has come, the stored effectiveness data is transmitted to the server 14 of the data center 44 (S153a). In the server 14, the transmitted effectiveness data is received and stored as effectiveness data 42 in the big data DB 36 in S153b of the advertising effectiveness measurement process.
[0033] Next, the user terminal 74 of the advertiser 30 erases the stored effectiveness data in S154. Next, in S155, it is determined whether an optimized advertisement or a price fluctuation forecast has been received, and if not, control returns to S150. On the other hand, if an optimized advertisement is transmitted from the server 14 in S48 (described later), a YES determination is made in S155, the received optimized advertisement is stored in S156, and control returns to S150. Also, if a price fluctuation forecast is transmitted from the server 14 in S58 (described later), a YES determination is made in S155, the received price fluctuation forecast is stored in S156, and control returns to S150.
[0034] Next, the details of the advertising effectiveness measurement process by the advertising effectiveness measurement means 38 will be described with reference to Figure 4(B). In S1, a process is performed in which each value set classified by the single value set identification process is designated as gj. Here, j is the jth example in the training data and takes the value 1, 2, 3, ..., J. J is the total number of value sets. A value set is composed of "a product image + a text image such as the price and product name" on a flyer (see value sets 99a, 99b, 99c, 99d in Figure 10(B)). The "single value set identification process" is a process in which a plurality of value sets listed on a flyer are classified into single value sets and identified, and will be described later with reference to Figure 11(B).
[0035] Next, in S2, an initial value of "1" is set for j. Next, in S3, effect data 42 related to gj is read from the big data DB 36. At this stage, j=1, so effect data related to the value set of g1 is read. For example, if g1 is the value set of advertiser A's product chicken breast (see value set 99d in Figures 10(B) and 10(C)), the effect data for chicken breast sent from advertiser A (POS data, number of flyer coupon uses, number of flyer inquiries, etc.) is read.
[0036] Next, in S4, the number of units sold, sales, number of flyer coupon usages, and number of flyer inquiries sent by the advertiser are divided by their respective standard values to calculate the values. For example, in the case of g1 (chicken breast), these standard values are the current average values for the number of units sold, sales, number of flyer coupon usages, and number of flyer inquiries for g1 (chicken breast) for that advertiser A.
[0037] Next, in S5, a process is performed to calculate an advertising effectiveness score tj for gj by adding up the calculated values. In S6, the advertising effectiveness (tj) for gj is output from the advertising effectiveness measurement means 38 and provided to the corresponding advertiser 30.
[0038] In S7, 1 is added to the current j, and in S8, it is determined whether j has exceeded J. If it has not yet exceeded J, control returns to S3. Each time the loop of S3 → S4 → S5 → S6 → S7 → S8 → S3 is circulated, j is incremented by "1" in S7, and j is incremented to 1, 2, 3, ... The value set gj is also updated to g1, g2, g3, ... Then, when j exceeds the total number of value sets J, S8 determines YES, and the advertising effectiveness measurement process returns.
[0039] Next, the big data machine learning process by the machine learning means 34 will be described with reference to Fig. 4(C). Machine learning for price fluctuation prediction is performed in S9, and machine learning for optimized advertisement generation is performed in S10.
[0040] The flowchart of the subroutine program for machine learning to predict price fluctuations in S9 is explained based on Figure 5(A). In S15, each product type is defined as kj, where j is the jth product type among all product types and takes the value 1, 2, 3, ..., J. J is the total number of product types.
[0041] Next, in S16, j is set to an initial value of 1. In S17, training data {(xji, yji)} is generated for product type kj, consisting of a dataset in which past crawler collected data is vector xji and the price at that time is correct answer information yji. Here, i is the i-th example in the training data and takes the value 1, 2, 3, . . . , I. I is the total number of training data.
[0042] Next, in S18, supervised learning is performed using the training data {(xji, yji)}, and a process is performed to learn a function cj(x) (cj:x→y) that maps an input x to a correct answer y as a regression problem.
[0043] In S19, 1 is added to the current j, and in S20, it is determined whether j has exceeded J. If it has not yet exceeded J, control returns to S17. Each time the loop of S17 → S18 → S19 → S20 → S17 is executed, j is incremented by 1 in S19, and j is incremented to 1, 2, 3, ... The product type kj is also updated to k1, k2, k3, ... Then, when j exceeds the total number of product types J, S20 determines YES, and the machine learning for price fluctuation prediction returns.
[0044] Next, a flowchart of the subroutine program for machine learning for generating optimized advertisements in S10 will be described with reference to FIG. 5(B). In S25, each product type to be learned is defined as kj. Here, j is the jth product type among all product types and takes the value 1, 2, 3, . . . , J. J is the total number of product types. Next, in S26, j is set to the initial value "1." In S27, it is determined whether supervised learning has been completed for product type kj. At this stage, since j = 1, it is determined whether supervised learning has been completed for product type k1. If supervised learning has not yet been completed, supervised learning for generating optimized advertisements is performed in S28.
[0045] On the other hand, if supervised learning has already been completed, reinforcement learning for optimized ad generation is performed in S29. In other words, in machine learning for optimized ad generation, supervised learning is performed first, and reinforcement learning is performed for product types for which supervised learning has been completed. This enables efficient machine learning for optimized ad generation. Reinforcement learning is a mechanism in which an agent placed in a certain environmental state acquires a strategy that maximizes the cumulative reward from the initial state to the goal based on the reward given when selecting an action. In reinforcement learning, learning progresses through interaction between a software agent (hereinafter referred to as "agent"), a type of AI, and the environment. An agent is a type of AI software that behaves autonomously and persistently with a certain degree of judgment while communicating with users and software. When the agent performs a certain action a on the environment, the state s of the environment changes and a certain goal state is reached, and a reward r is given to the agent. The agent learns a function that takes the state s as input and outputs action a with the goal of maximizing this reward r.
[0046] Reinforcement learning progresses over time by repeating the following simple steps: 1 The agent receives an observation o from the environment (or directly the state s of the environment) and returns an action a to the environment based on a policy π. 2 Based on the action a received from the agent and the current state s, the environment changes to the next state s', and based on that transition, it returns to the agent the next observation o' and a single number (scalar quantity) called the reward r, which indicates the quality of the previous action. 3 Progression of time: t←t+1 Here, ← represents an assignment operation.
[0047] After execution of step S28 or S29, control proceeds to S30, where 1 is added to the current j, and in S31, it is determined whether j has exceeded J. If j has not yet exceeded J, control returns to S27. Each time the loop of S27 → S28 or S29 → S30 → S31 → S27 is executed, j is incremented by 1 in S30, and j is incremented to 1, 2, 3, ... The product type kj is also updated to k1, k2, k3, ... Then, when j exceeds the total number of product types J, S31 determines YES, and the machine learning for generating optimized advertisements returns.
[0048] Next, a flowchart of the supervised learning subroutine program for generating optimized advertisements in S28 will be explained based on Figure 6(A). In S36, a process is performed to extract all prices Pji listed on flyers 27 with high scores among the past advertising effectiveness scores tj calculated by the advertising effectiveness measurement process (see Figure 4(B)) for product type kj, and the predicted prices pji at that time. Here, i is the ith extracted data among all extracted data, and takes the value of 1, 2, 3, ..., I. I is the total number of extracted data. Specifically, a "high-scoring flyer" is a flyer with an advertising effectiveness score in, for example, the top 10% of the past advertising effectiveness scores tj.
[0049] Next, in S37, training data {(xji, yji)} is generated for kj, consisting of a data set in which pji is input information (vector xji) and Pji is correct answer information yji. Pji is the price listed in a flyer with a high score for product type kj (see S36), and is data suitable for correct answer information.
[0050] Next, in S38, supervised learning is performed using {(xji, yji)}, and a function cj(x) (cj: x → y) that maps input x to correct answer y as a regression problem is learned. When supervised learning for generating an optimized advertisement for product type kj is completed, for example, a completion flag is associated with kj and stored. Whether or not this completion flag is stored is determined in the above-mentioned S27, thereby determining whether or not kj has completed supervised learning.
[0051] Next, a flowchart of the subroutine program of reinforcement learning for generating optimized advertisements in S29 will be explained based on Fig. 6(B). In S46, a process is performed to calculate a reward r for a product type kj listed in the flyer 27 of the optimized advertisement provided to the advertiser 30 based on the advertising effectiveness score tj calculated by the advertising effectiveness measurement process (see Fig. 4(B)). The reward r is controlled so that the higher the advertising effectiveness score tj, the larger the reward r. Next, in S47, the optimal policy π * The process of calculating the price Pi as an action a according to the above is performed. TD learning stands for Temporal Difference Learning, which estimates the Q value using a model-free method. The state at time t is s t In general, the optimal policy is π * (s t ):Choose a t * ifQ * (s t ,a t * ) Next, in S48, an optimized advertisement with the product price set to Pj is generated and sent (provided) to the advertiser 30 or the advertisement creation company 31.
[0052] Next, the contents of the price fluctuation prediction process will be explained with reference to FIG. 7. In S55, each product type is defined as kj. Here, j is the jth product type among all product types and takes the value 1, 2, 3, ..., J. J is the total number of product types. In S56, j is set to the initial value "1". In S57, each current effect data 42 is read from the big data DB 36 and substituted into the function cj(x), thereby predicting the price Pj. Here, the function cj(x) is the price fluctuation prediction function learned in S18 described above.
[0053] Next, the predicted price pj is transmitted in S58. Specifically, this transmission involves transmitting the price pj to the advertiser or ad creation company that handles the product type kj. Note that instead of transmitting the prices pj one by one in S58, S58 may aggregate the prices pj for each product type kj and transmit the aggregated results to the advertiser or ad creation company. In S59, 1 is added to the current j, and in S60, it is determined whether j has exceeded J. If j has not yet exceeded J, control returns to S57. Each time the loop of S57 → S58 → S59 → S60 → S57 is executed, j is incremented by 1 in S59, and j is incremented to 1, 2, 3, . . . The product type kj is also updated to k1, k2, k3, . . . When j exceeds the total number of product types J, S60 determines YES, and the price fluctuation prediction process returns.
[0054] Next, machine learning that the proofreading machine learning means 33 performs in advance so that the advertising-related service providing means 35 can determine the proofreading portions of the flyer 27 will be described with reference to Figure 8. Product feature learning data 87, which associates a large number of original product images (images of actual products) 25 with labels such as JAN codes of the products, is input to the proofreading machine learning means 33, and product feature learning 22 is performed based on the product feature learning data 87. In the case of packaged products, JAN codes are assigned and affixed, and the JAN codes are used as labels to identify the products. However, in the case of non-packaged products such as salmon fillets, no JAN codes are assigned, and so a house code assigned independently by the entity (vendor) providing the advertising-related service is used as a product identification label.
[0055] A large amount of flyer 27 data is input to the proofreading machine learning means 33, and based on the flyer 27 data, flyer feature learning 23, advertising value set boundary learning 3, product image identification learning 4, character position learning 28, OCR learning 16, and word classification learning 39 are performed. Original product images 25 are image data of the actual products featured in the flyer 27, and also include image data of similar products. The more similar products there are, the more features can be extracted.
[0056] Product feature learning 22 learns the features of actual products because the same product image on a flyer may be viewed from different angles or in different colors, and learns the features of each product using a neural network (NN) or a deep neural network (DNN). The learning results, learned product features 20, are stored in the learning stage learned storage unit 21. The learned product features 20 are stored by associating the features of each product with the product's label (JAN code, house code, etc.).
[0057] Leaflet feature learning 23 learns the features of a leaflet from the image data of the input leaflet 27, and uses NN or DNN for learning. The learning result, learned leaflet feature 1, is stored in learning stage learned storage unit 21.
[0058] The advertising price boundary learning 3, product image identification learning 4, character position learning 28, OCR learning 16, and word classification learning 39 use supervised learning similar to that shown in Figures 5(A) and 6(A), but the correct answer is label classification. Furthermore, since machine learning is performed using large amounts of image data, we use convolutional neural networks (CNNs), recurrent neural networks (RNNs), and long short-term memories (LSTMs) to minimize the human burden of extracting features and inputting them as input vectors x. CNNs are convolutional neural networks, a type of forward propagation artificial deep neural network. They are widely used in image and video recognition. RNNs are recurrent neural networks, an algorithm that has achieved excellent results in the field of natural language processing. LSTMs, which emerged in 1995 as an extension of RNNs, are a type of model or architecture for sequential data that enables long-term utilization of short-term memory.
[0059] The advertising value set boundary learning 3 is a machine learning method for identifying the boundaries between multiple value set images published on a flyer and extracting a single value set image from the multiple value set images, and uses CNN, RNN, and LSTM. The advertising value set boundary learned model 2, which is the learning result of the advertising value set boundary learning 3, is stored in the learning stage learned storage unit 21.
[0060] The commodity image classification training 4 is machine learning for identifying commodity images from extracted single-value pair images, and uses CNN or the like. A single-value pair image includes a commodity image and a text image (character image) such as the price and name of the commodity, and the commodity image classification training 4 is machine learning for identifying commodity images from such single-value pair images. A commodity image classification trained model 18, which is the learning result of this commodity image classification training 4, is stored in the learning stage trained storage unit 21.
[0061] Character position learning 28 is machine learning to identify the position of a text image (character image) relative to a product image based on the extracted single-valued image, and uses CNN. A character position trained model 19, which is the learning result of this character position learning 28, is stored in the learning stage trained storage unit 21.
[0062] OCR learning 16 is machine learning that uses LSTM or CNN to learn the font of characters used in flyers 27, designs with different character sizes and thicknesses, character strings (one line) of different lengths, character spacing, line spacing, background color (with noise), etc. The OCR trained model 17, which is the learning result of this OCR learning 16, is stored in the learning stage trained storage unit 21.
[0063] Word classification learning 39 is machine learning to identify the category (price, capacity, product name, manufacturer name, etc.) to which the content of each character string identified by OCR belongs, and a word classification trained model 32, which is the learning result of this word classification learning 39, is stored in the learning stage trained storage unit 21. AI data is created using the OCR trained model 17 and the word classification trained model 32, as described below (see Figure 12(C)). The manuscript data 24 is Excel data in which items are separated by category (see "label, product name, meter, price," etc. in Figure 10(E)), and by using this word classification trained model 32, it is possible to compare the manuscript data 24 and the AI data for each item.
[0064] Next, the flyer proofreading portion determination process performed by the advertising-related service providing means 35 will be described with reference to FIG. 9. In S65, it is determined whether a flyer image has been received. If the flyer image data is sent to the data center 44 and input to the advertising-related service providing means 35 in S183, which will be described later, then S65 determines YES. Note that instead of sending the flyer image data, the advertising creation company 31 may send the flyer itself to the entity (vendor) providing the advertising-related service. In that case, the entity (vendor) providing the advertising-related service scans the image of the flyer page or converts the data typeset on a personal computer into a PDF, and then transmits the image data to the server 14 for input to the advertising-related service providing means 35.
[0065] If the image data of the flyer 27 provided by the advertisement creation company 31 is input to the advertisement-related service providing means 35, a YES determination is made in S65 and a determination is made in S67 as to whether or not the manuscript data 24 has been accepted. If the manuscript data 24 transmitted (provided) by the advertisement creation company 31 is input to the advertisement-related service providing means 35 in S181, which will be described later, a YES determination is made in S67 and control proceeds to S68.
[0066] On the other hand, if S65 returns NO or S67 returns NO, control proceeds to S66, where crawler processing is performed. This crawler processing involves a crawler periodically patrolling the Internet to collect information such as the products sold and business hours from the homepage of advertiser 30. The collected information is compared with input manuscript data 24, and discrepancies are extracted. For example, if advertiser 30 is drugstore A, information on the products sold (e.g., Category 2 and Category 3 drugs) is collected from the homepage of drugstore A and compared with the manuscript data 24 and image data of flyer 27 of drugstore A. As a result, if a Category 1 drug is included in the image data of manuscript data 24 or flyer 27, it is determined that the manuscript data 24 or flyer 27 is incorrect (see FIG. 13(A)).
[0067] In S68, a reference storage unit selection process is performed. In this embodiment, two types of storage units for trained models are provided: the aforementioned learning stage trained model storage unit 21 and the trained model storage unit 29 using training data in the practical stage. In S68, it is determined which of the two storage units the trained model stored in should be used.
[0068] In S69, a single value pair identification process is performed, in S70 a product image identification process is performed, in S71 a target product image extraction process is performed, in S72 a corresponding label search process is performed, in S73 a character position identification process is performed, in S74 an AI data creation process using OCR is performed, and in S75 a word classification identification process is performed. Each of these processes will be described in detail later.
[0069] Next, in S76, M is decremented by 1. This M is the number of divisions defined in S89 in the single value pair identification process of Fig. 11(B), and is the number of divisions into which the multiple value pairs listed in the flyer 27 received in S65 are divided into single value pairs.
[0070] Next, S77 determines whether M=0. If it is not yet 0, control returns to S71, and the loop of S71 → S72 → S73 → S74 → S75 → S76 → S77 → S71 is repeated. Each time this loop is repeated, M is decremented by 1 in S76, and when M=0, control proceeds to S78. In other words, the processing of S71 to S75 is repeated for the number of value pairs listed in the flyer accepted in S65, and when the processing of S71 to S75 has been completed for all value pairs, control proceeds to S78.
[0071] In S78, a comparison process is performed between the crawler collected data and the manuscript data, in S79 a comparison process is performed between the AI data and the manuscript data, and in S80 the proofreading part is output and provided to the client, the advertisement creation company 31. Next, in S81, a storage process is performed for the price data, after which control returns to S65.
[0072] The main part of this flyer proofreading part determination process will be explained with reference to Fig. 10. Referring to Fig. 10(A), a plurality of value sets are listed on flyer 27 provided by advertisement production company 31. Fig. 10 shows four types of value sets: a value set 99a for tomatoes, a value set 99b for green peppers, a value set 99c for domestic beef loin, and a value set 99d for chicken breast.
[0073] From the image of the flyer 27 containing these multiple value sets, the advertising value set boundaries are determined and each single value set is identified (see FIG. 10(B)). Next, a single value set (product frame) 99d is cut out (see FIG. 10(C)). The image of this cut-out single value set 99d is separated into a product image portion 97 and a character image portion 98 (see FIG. 10(D)).
[0074] Next, the separated character image portion 98 is converted into text data using OCR to generate AI data, and the AI data is compared (matched) with the manuscript data 24 to identify and output inconsistent areas (proofreading areas) (see Figure 10(E)).
[0075] Next, a flowchart of the subroutine program for the reference storage unit selection process in S68 will be described with reference to FIG. 11A. In S83, the business type, region, and season are identified from the received flyer image. Specifically, the insertion date listed on the flyer is read using OCR to recognize the date and identify the season, the store address listed on the flyer is read using OCR to identify the region, and the label (e.g., JAN code or house code) extracted by the corresponding label extraction process in FIG. 12A is used to identify the products and determine the product number composition, thereby identifying the business type. Specifically, the types of multiple products listed on the flyer are identified from the labels such as JAN codes or house codes, and the number of products by type (product number composition) is determined to determine which type of product the flyer is mainly about, thereby identifying the business type. Specific examples of business types include grocery stores, general merchandise stores (GMS), drug stores, home centers (including general discount stores), liquor discount stores, real estate, etc. Real estate includes new single-family homes, used single-family homes, new condominiums, and used condominiums.
[0076] Next, in S84, it is determined whether a trained model corresponding to the received business type, region, and season is stored in the training data storage unit 29 in the practical stage. In the training data storage unit 29 in the practical stage, the data is grouped into a three-dimensional table consisting of business type, region, and season (spring, summer, fall, winter), as shown in Figure 15(B), and a trained model is stored for each group. The generation process of the trained models stored for each group will be explained first with reference to Figure 15(A).
[0077] FIG. 15(A) is a flowchart of the learning process using training data in the practical phase. Various trained models that have completed the learning phase are utilized in the practical phase (FIGS. 9 to 13) where advertising-related services are actually provided. However, errors occasionally occur even in the practical phase. This is because in the practical phase, flyers 27 and their manuscript data 24 are provided from special business types or special regions, and special character fonts or special expressions that were not handled in the learning phase may be encountered, resulting in errors. In addition, special expressions may be used depending on the season. Therefore, in order to respond to the particularities of business types, regions, and seasons, in the practical phase, groups are created by business type, region, and season, and specialized machine learning is performed for each group.
[0078] In S129, it is determined whether the number of errors discovered during the production phase has reached a predetermined number R, and the process waits until the predetermined number is reached. When the advertising-related service providing means 35 notifies the advertising production company 31 of the advertisement proofreading points, and the advertising production company 31 attempts to proofread the flyer 27 based on the notification, if there are no errors, a notice (complaint) stating that the notified proofreading points were not actually errors is sent to the advertising-related service provider, making it possible to discover errors during the production phase. However, if an error occurs in which a mismatch between the flyer 27 and the manuscript data 24 is overlooked, the advertising production company 31 is not notified of the advertisement proofreading points, and therefore no notice (complaint) from the advertising production company 31 is sent to the advertising-related service provider, making it difficult to discover errors during the production phase. One way to solve this problem is to pay an appropriate compensation to the advertiser 30, etc. who discovers an error in the flyer 27, thereby creating an incentive for notifying the advertising-related service provider of the discovery of the error in the flyer 27.
[0079] When the number of errors reaches a predetermined number R, control proceeds to S130, where the value set in which the error occurred is grouped into a three-dimensional table consisting of business type, region, and season. This process is performed by specifying the business type, region, and season using the method described above in S83.
[0080] Next, in S131, correct answer information yi is created for each group. Here, i is the ith correct answer information among all the correct answer information and takes the value of 1, 2, . . . I. I is the total number of correct answer information. More specifically, this correct answer information creation process determines at which processing stage the error occurred, i.e., in which of the single value pair identification process, product image identification process, target product image extraction process, corresponding label search process, character position identification process, word classification identification process, or OCR AI data creation process, the error occurred, and creates correct answer information yi for machine learning corresponding to the process in which the error occurred. For example, if an error occurs in the character position identification process, correct answer information yi for character position learning 28 is created.
[0081] Next, in S132, the input data of the error part for each group is designated as xi. Here, i is the ith input data among all input data and takes the value of 1, 2, . . . I. I is the total number of input data. Training data {(xi, yi)} for supervised learning is generated from a data set of this input data xi and correct answer information yi. For example, if this training data {(xi, yi)} is generated when an error occurs in the character position identification process, it is used as training data {(xi, yi)} for retraining the character position learned model 19.
[0082] In S133, a single example (the i-th example), {(xi, yi)}, is reused multiple times (e.g., 20 times) as different examples, and this reuse is repeated for all i to create additional training data. In S134, supervised learning is performed using the additional training data added to the existing training data used in the learning stage as training data for the group, and a function c(x) (c: x → y) that maps input x to the correct answer y is learned for each group. This allows the learned model of a group in which an error occurred to be retrained using additional training data in which the correct answer information yi for the error and the input data xi for the erroneous part are reused multiple times (e.g., 20 times), thereby providing the advantage of machine learning that increases the influence of training data generated based on errors that occurred in the practical stage.
[0083] Another method for performing machine learning that increases the influence of training data generated based on errors made in the practical stage is to thin out the existing training data by the number I of pieces of training data {(xi, yi)} generated in S132, and then perform supervised learning using the data obtained by adding the training data {(xi, yi)} generated in S132 to the thinned-out existing training data as the training data for the group. This method also has the advantage of being able to perform machine learning that increases the influence of training data generated based on errors made in the practical stage.
[0084] Next, in S135, the trained model for each group is updated to the function c(x), and the trained models having the updated function c(x) are classified by group and stored in the trained storage unit 29 using training data from the practical stage (see FIG. 15(B)). More specifically, the machine learning specialized for each group retrains the trained model used for the process in which the error occurred, among the single value pair identification process, product image identification process, target product image extraction process, corresponding label search process, character position identification process, word classification identification process, or AI data creation process using OCR. Therefore, the trained models classified by group and stored include the retrained flyer feature trained model 1, advertising value pair boundary trained model 2, product image identification trained model 18, character position trained model 19, OCR trained model 17, and word classification trained model 32.
[0085] 11(A), if S84 returns YES, control proceeds to S86, where the learned storage unit 29 using the training data from the practical stage is selected. As a result, advanced processing becomes possible using a learned model that has been retrained using the training data from the practical stage and can handle the particularities of business type, region, and season. On the other hand, if S87 returns NO, control proceeds to S85, where the learning stage learned storage unit 21 is selected.
[0086] Next, a flowchart of the subroutine program for the single value set identification process of S69 will be described with reference to Figure 11 (B). A process of reading out the advertising value set boundary trained model 2 from the learned storage unit selected in S85 or S86 is performed in S87. Next, based on the read advertising value set boundary trained model 2, single value set classification identification is performed in S88 to classify the multiple value sets published in the flyer 27 received in S65 into single value sets. Next, a process of setting the number of classifications identified to M is performed in S89, and this single value set identification process returns.
[0087] Next, a flowchart of the subroutine program for the commodity image identification process of S70 will be described with reference to Figure 11(C). A process of reading the data of the trained flyer feature 1 and the trained model for commodity image identification 18 from the trained storage unit selected in S85 or S86 is performed in S95. Next, a process of identifying a commodity image from the image published in the flyer 27 received in S65 is performed in S96 based on the read-out data of the trained flyer feature 1 and the trained model for commodity image identification 18.
[0088] Next, a flowchart of the subroutine program for the target product image extraction process of S71 will be described with reference to Figure 11(D). A process of reading the trained model 18 for commodity image identification from the trained storage unit selected in S85 or S86 is performed in S98. Next, a process of sequentially extracting target product images from the single-value set images sorted in S88 based on the read trained model for commodity image identification is performed in S99.
[0089] Next, a flowchart of the subroutine program of the corresponding label extraction process of S72 will be described with reference to Figure 12(A). A process of comparing the data of the learned product feature 20 stored in the learned storage unit selected in S85 or S86 with the target product image extracted in S99 and selecting a product with a high degree of similarity is performed in S104. Next, a process of reading out the label (see Figure 8) stored in correspondence with the feature of the selected product is performed in S105.
[0090] Next, a flowchart of the subroutine program for the character position identification process in S73 will be explained with reference to Fig. 12(B). Based on the character position trained model 19 stored in the trained storage unit selected in S85 or S86, a process is performed in S108 to discriminate the single value set image identified in S88 into a product image and a character image and identify the character image portion.
[0091] Next, a flowchart of the subroutine program for the AI data creation process by OCR in S74 will be described with reference to Figure 12(C). A process of reading the OCR trained model 17 from the trained storage unit selected in S85 or S86 is performed in S111. Next, a process of converting the character image portion identified in S108 into text data by OCR using the OCR trained model 17 is performed in S112. Next, a word classification identification process is performed in S113. Next, a process of creating a table (AI data) that associates the label read in S105 with the text data word-classified in S113 is performed in S114.
[0092] Next, a flowchart of the subroutine program for the word classification identification process of S75 will be described with reference to Figure 12(D). A process of reading out the word classification trained model 32 from the trained storage unit selected in S85 or S86 is performed in S116. A process of classifying and identifying the words in the text data using the word classification trained model 32 is performed in S117.
[0093] Next, a flowchart of the subroutine program for processing to compare the crawler-collected data with the manuscript data in S78 will be described with reference to Fig. 13(A). In S125, the crawler-collected data is compared with the manuscript data 24. As a result of the comparison, a determination is made in S126 as to whether or not the two match. If they match, the processing to compare the crawler-collected data with the manuscript data returns. If they do not match, control proceeds to S127, where all data of the mismatched parts (data and labels of the mismatched parts) is output as proofreading parts and provided to the client, advertising production company 31. Furthermore, in addition to or instead of the processing to compare the crawler-collected data with the manuscript data, processing to compare the crawler-collected data with the data of the flyer 27 may be performed.
[0094] Next, a flowchart of the subroutine program for the process of comparing the AI data and the manuscript data 24 in S79 will be described with reference to Figure 13(B). In S120, a process is performed to compare data corresponding to the same label between the AI data and the manuscript data 24. As a result of the comparison, a determination is made in S121 as to whether or not the two match. If they match, control proceeds to S123, but if they do not match, control proceeds to S122, where the data and label of the mismatched part are stored as a proofread part, and then control proceeds to S123.
[0095] In S123, it is determined whether or not comparison of all AI data has been completed, and if not, control returns to S120. By circulating this loop of S120 → S121 → S122 or S123 → S120, if comparison of all AI data has been completed, a YES determination is made in S123, control moves to S124, and processing is performed in which all proofreading points (data and labels of mismatched points) stored in S122 are output and provided (transmitted) to the client, advertisement creation company 31.
[0096] Next, a flowchart of the subroutine program for the price data storage process shown in S81 will be described with reference to Figure 13(C). In S128, product data (product name, volume, price, sales limit, etc.) including the price of each product other than the proofreading part is stored together with the date in the big data DB 36. This makes it possible to read and use the price data 41 stored in the big data DB 36 in, for example, S17 and S36.
[0097] Next, the control processing by the user terminal 74 of the advertisement production company 31 will be described with reference to Figure 14. In S180, it is determined whether or not manuscript data for creating a flyer will be sent, and if it has not yet been sent, control proceeds to S182. If manuscript data has been created and will be sent, the determination in S180 is YES, and the manuscript data is sent to the data center 44 in S181.
[0098] Next, in S182, it is determined whether or not to send a flyer, and if not yet sent, control proceeds to S184. If sent, image data of the flyer is sent to the data center 44 in S183. Note that the manuscript data sent in S181 and the flyer image data sent in S183 may be sent simultaneously rather than separately. Next, in S184, it is determined whether or not the proofread portions of the flyer have been received, and if not, control returns to S180. On the other hand, if the proofread portions are sent from the server 14 of the data center 44 in S124, YES is determined in S184, and the proofread portions are stored in S185.
[0099] Next, in S186, it is determined whether other data has been received. "Other data" is, for example, the optimized advertisement transmitted in S48 or the price fluctuation forecast data transmitted in S58.
[0100] Modifications of the present embodiment described above are listed below.
[0101] (Variant 1) In the above-described embodiment, the advertising-related service providing means 35 finds the corrections in the flyer 27 and provides them to the advertising production company 31, but this is not limited to this. It may also be a service that creates an electronic flyer based on an advertising request from the advertiser 30 and makes the electronic flyer searchable and viewable on the web.
[0102] (Variant 2) In the above-described embodiment, a flyer 27 was shown as an example of an advertisement, but it may also be a paper medium advertisement other than a flyer, such as a pamphlet, or even an advertisement other than a paper medium, such as a web advertisement.
[0103] (Variant 3) The data stored in the big data DB 36 may be grouped by predetermined attributes (for example, a three-dimensional table consisting of business type, region, and season (spring, summer, fall, winter)), similar to the learned storage unit 29 using training data in the practical stage. Then, the machine learning means 34 may perform machine learning specialized for each group, and each learned model that is the learning result may be grouped and stored by attribute (for example, the above-mentioned three-dimensional table). The learned models 50 specialized for each group may be used to perform price fluctuation predictions, futures trading, and the provision of optimized advertisements specialized for each group.
[0104] (Variant 4) While supervised learning has been shown as machine learning for price fluctuation prediction (see Figure 5(A)), reinforcement learning may be used instead or in addition, in which agents are rewarded based on how accurate their price fluctuation predictions are, and price fluctuation predictions may also be made using data mining.
[0105] (Variation 5) Deep Q-Network may be adopted as reinforcement learning for generating optimized advertisements (S29). Deep Q-Network is a technique that applies deep learning to function approximation in reinforcement learning. The greatest feature of Q-Learning in reinforcement learning is that if samples (s, a, r, s') can be obtained infinitely from all pairs of (s, a), then the optimal value function Q will always be obtained regardless of the order in which they are given. * The key point is that (s, a) can be obtained. Creating a table function Q(s, a) for all states and actions would result in a huge amount of data to be processed, so function approximation is used for Q(s, a). Deep reinforcement learning is the application of deep learning techniques to this function approximation. Therefore, deep reinforcement learning is also a type of reinforcement learning, and the term "reinforcement learning" is a broad concept that includes "deep reinforcement learning."
[0106] (Variant 6) The advertising-related service providing means 35 described above discovers parts of the flyer 27 that differ from the manuscript data 24 and need to be proofread, and notifies the advertising production company 31 of this. However, it is also possible to provide the advertising production company 31 with data on a constructed flyer in which the parts of the flyer 27 that differ from the manuscript data 24 have been proofread to correct the content.
[0107] (Variant 7) In the above-described embodiment, the optimized advertisement was an advertisement for a product price Pi that maximizes the advertising effect, but in addition to or instead of that, it may also be an advertisement created by predicting best-selling products.
[0108] The subject matter disclosed in the above-described embodiments will now be described.
[0109] [Theme 1] One service provision system that accumulates price data as big data and provides services using that big data involves having information providers provide market information on fluctuations in various trading prices, and then predicting price fluctuations based on that information (e.g., JP 2001-22849 A).
[0110] The technique described in Patent Document 1 has the drawback that it is difficult to collect information from information providers, and in order to give information providers an incentive to provide information, the method of having them provide information for a fee has to be adopted.
[0111] The present subject matter 1 was conceived in view of the above circumstances, and its purpose is to provide a service providing system that can give an incentive to an information provider for providing information.
[0112] This subject 1 can be expressed, for example, as the following items:
[0113] (Item 1) A service providing system that accumulates at least price data (e.g., product data including the price of each product (product name, volume, price, sales limit, etc.)) as big data (e.g., S128) and provides a service using the big data, data receiving means (e.g., S65, S67) for receiving at least price data (e.g., product data including the price of each product (product name, volume, price, sales limit, etc.)); big data storage means (e.g., big data DB36, S128) that stores the price data (e.g., price data 41) received by the data receiving means as big data; A service providing means (e.g., S48, S55) that provides a service (e.g., price fluctuation prediction, creation and provision of optimized advertisements, futures trading) that reflects the learning results obtained by machine learning means (e.g., machine learning means 34, S9, S10) that performed machine learning based on the big data accumulated by the big data accumulation means; An advertisement-related service providing means (for example, advertisement-related service providing means 35, S65 to S80) for providing a service related to advertisement creation, The advertising-related service providing means provides a service (e.g., S65 to S80) related to advertisement creation that reflects the price data received by the data receiving means to a provider of the price data (e.g., advertisement creation company 31), a service providing system.
[0114] According to this service providing system, information providers are provided with a service relating to the creation of advertisements that reflect price data, and thus receive an incentive to provide information, which encourages them to take the initiative in providing information.
[0115] (Item 2) The big data storage means also stores fluctuation factor data that causes price fluctuations (for example, weather, natural disasters, stock prices, exchange rates, interest rates, economic trends, income fluctuations, and geopolitical risks as crawler collected data 40) as big data, Further comprising a prediction means (e.g., S55 to S60) for predicting price fluctuations based on the price data and the fluctuation factor data accumulated by the big data accumulation means, The prediction means predicts the price fluctuations by reflecting the learning results obtained by the machine learning means performing machine learning (e.g., S15 to S20) based on the price data and the fluctuation factor data (e.g., S57); The service providing system according to item 1, wherein the service providing means provides a service that reflects the price fluctuation prediction by the prediction means (e.g., price fluctuation prediction, creation and provision of optimized advertisements).
[0116] According to this service providing system, information providers can receive services that reflect forecasts of price fluctuations.
[0117] (Item 3) The machine learning means performs machine learning (e.g., S25 to S31, S36 to S38, S46, S47) to create an optimized advertisement based on the big data stored by the big data storage means, and the machine learning means creates an advertisement by reflecting the learning results (e.g., S48). 3. The service providing system according to item 1 or 2, wherein the service providing means provides the advertisement created by the advertisement creating means (for example, S48).
[0118] According to this service providing system, it is possible to obtain advertisements that are created by reflecting the results of machine learning for creating optimized advertisements.
[0119] (Item 4) Further provided is an advertisement effectiveness measurement means (for example, advertisement effectiveness measurement means 38, S1 to S8) for measuring the effectiveness of the advertisement created by the advertisement creation means, The machine learning for creating the optimized advertisement is reinforcement learning in which an agent is given a reward according to the measurement results by the advertising effectiveness measurement means, and the agent learns a strategy to maximize the accumulation of the reward (e.g., S46, S47).
[0120] According to this service providing system, it is possible to provide advertisements with high advertising effectiveness as a result of reinforcement learning.
[0121] (Item 5) Further provided is an advertisement effectiveness measurement means (for example, advertisement effectiveness measurement means 38, S1 to S8) for measuring the effectiveness of the advertisement created by the advertisement creation means, The machine learning for creating the optimized advertisement is a service provision system described in item 3, in which training data (e.g., {(xji, yji)}) is generated (e.g., S37) using input information as the price predicted by a prediction means based on the price data accumulated by the big data accumulation means and fluctuation factor data that causes price fluctuations, and advertising content whose effectiveness is measured to be at least a predetermined level by the advertising effectiveness measurement means as correct answer information, and supervised learning is performed using the training data to learn a function (e.g., cj(x)(cj:x→y)) that maps the input to a correct answer as a regression problem (e.g., S36 to S38).
[0122] According to this service providing system, it is possible to provide advertisements with high advertising effectiveness as a result of supervised learning.
[0123] (Item 6) A user-side facility (e.g., a user terminal 74 connected to an internal LAN 84 of an advertiser 30, a user terminal 74 connected to an internal LAN 84 of an advertisement creation company 31) that can be connected to the Internet to a service providing system (e.g., a data center 44) that accumulates at least price data (e.g., product data including the price of each product (product name, volume, price, sales limit, etc.)) as big data (e.g., S128) and provides a service using the big data, data providing means (e.g., S181, S183) for providing at least price data (e.g., product data including the price of each product (product name, volume, price, sales limit, etc.)) to the service providing system; a service receiving means (e.g., S155, S187) that receives from the service providing system a service (e.g., optimized advertisement, price fluctuation prediction) that reflects the learning results of machine learning performed by machine learning means based on the big data stored by a big data storing means provided on the service providing system side; an advertisement-related service receiving means (e.g., S184) for receiving a service related to advertisement creation from an advertisement-related service providing means provided on the service providing system side; The advertisement-related service receiving means receives services (for example, S65 to S80) relating to the creation of advertisements that reflect the price data provided by the data providing means.
[0124] According to this service provision system, price data providers can receive services related to creating advertisements that reflect the price data provided by the data provision means, which gives them an incentive to provide price data and encourages them to take the initiative in providing price data.
[0125] [Theme 2] One advertising-related service providing system that identifies and notifies parts of an advertisement that need proofreading is one that compares advertisements such as flyers with their manuscript data and finds parts that do not conform to the manuscript (for example, JP 2004-133734 A).
[0126] However, the technique described in this patent document has the drawback that manual work is required to match the advertisement with its manuscript data.
[0127] The present subject matter 2 was devised in view of the above situation, and its purpose is to provide an advertising-related service providing system that can minimize the manual work required to match advertisements with their manuscript data.
[0128] This theme 2 can be expressed, for example, as the following items:
[0129] (Item 1) An advertising-related service providing system that provides services related to advertisement creation, advertising data receiving means (e.g., S65) for receiving data of advertisements (e.g., flyers 27, pamphlets, web advertisements, etc.); manuscript data receiving means (e.g., S67) for receiving manuscript data (e.g., manuscript data 2) used in creating the advertisement; and proofreading portion determining means (e.g., S68 to S79) for determining consistency between the advertisement and the manuscript data based on the advertisement data received by the advertisement data receiving means and the manuscript data received by the manuscript data receiving means, and determining portions of the advertisement that need proofreading; The proofreading part determination means determines the consistency by reflecting the learning results (e.g., trained flyer feature 1, advertising value set boundary trained model 2, trained product feature 20, product image recognition trained model 18, character position trained model 19, OCR trained model 17, word classification trained model 32) performed by a proofreading machine learning means (e.g., proofreading machine learning means 33) in machine learning (e.g., flyer feature learning 23, advertising value set boundary learning 3, product feature learning 22, product image recognition training 4, character position trained model 28, OCR trained model 16, word classification training 39) for determining the consistency between the advertisement and the manuscript data (e.g., S87, S88, S95, S96, S98, S99, S104, S105, S108, S116, S117, S111, S112, S116, S117), advertising-related service providing system.
[0130] According to this advertising-related service provision system, the consistency between an advertisement and manuscript data is determined by reflecting the learning results of the proofreading machine learning means that performs machine learning to determine the consistency between the advertisement and manuscript data, thereby minimizing the manual work required to match the advertisement with the manuscript data.
[0131] (Item 2) The advertisement data includes data of a plurality of advertisement items (for example, value sets 99a, 99b, 99c, and 99d shown in FIG. 10(B)), The manuscript data includes data specifying the advertisement content for each of the multiple products to be advertised (for example, the product name, manufacturer, price, etc. shown in FIG. 10(E)), The machine learning performed by the calibration machine learning means includes value set classification identification machine learning (e.g., flyer feature learning 23, advertising value set boundary learning 3) for classifying and identifying the advertising data received by the advertising data receiving means for each single value set, The calibration portion determination means A value group classification identification process (e.g., S87 to S89) that reflects the learning results of the value group classification identification machine learning and classifies and identifies the advertisement data for each single value group; a manuscript corresponding portion identification process (e.g., S95, S96, S98, S99, S104, S105, S108) for identifying which portion of the advertisement content in the manuscript data each of the single value pairs identified by the value pair classification identification process corresponds to; An advertising-related service providing system as described in item 1, which performs a consistency determination process (e.g., S111 to S114, S116, S117, S120 to S123) that matches each single value set identified by the value set-by-value set classification identification process with the advertising content identified by the manuscript corresponding portion identification process to determine the consistency.
[0132] According to this advertising-related service provision system, it is determined which part of the advertising content in the manuscript data each single-value pair corresponds to, and the determined advertising content is matched with each single-value pair to determine consistency, thereby enabling accurate consistency determination.
[0133] (Item 3) The advertising-related service providing system described in item 2, wherein the consistency determination process determines the consistency between the character image of each single value set identified by the value set classification identification process and the manuscript data (e.g., S111 to S114, S116, S117, S120 to S123).
[0134] According to this advertising-related service providing system, it is possible to determine whether or not a character image of a single value set matches manuscript data.
[0135] (Item 4) The machine learning performed by the proofreading machine learning means includes text data conversion machine learning (e.g., character position learning 28, OCR learning 16) for converting character images of each single value set identified by the value set classification identification process into text data, The advertising-related service providing system described in item 3, wherein the consistency determination process converts the character image into text data by reflecting the learning results of the text data conversion machine learning, and then determines the consistency between the text data and the manuscript data (e.g., S111 to S114, S116, S117, S120 to S123).
[0136] According to this advertising-related service providing system, character images are converted into text data, and then the consistency between the text data and manuscript data is determined, making it possible to determine good consistency between the two data.
[0137] (Item 5) The machine learning performed by the calibration machine learning means further includes product feature learning (e.g., product feature learning 22) for learning the features of the product based on data of an original image of an actual product corresponding to the advertised product, The system further includes a correspondence storage unit (e.g., a learned commodity feature quantity 20 in a learning stage learned storage unit 21) that stores the commodity feature learned by the commodity feature learning in association with identification information for identifying the commodity (e.g., a label such as a JAN code or a house code), The manuscript data is data in which the advertising content (e.g., product name, manufacturer, price, etc.) of each of the multiple products to be advertised is associated with the identification information (e.g., label) of the product (e.g., FIG. 10(E)), The manuscript corresponding portion identification process identifies identification information (e.g., label) stored in the correspondence storage means in correspondence with the product characteristics used to identify the advertised product, and identifies the advertising content (e.g., product name, manufacturer, price, etc.) associated with the identification information from the manuscript data (e.g., manuscript data 24) (e.g., S120), in an advertising-related service providing system described in any of items 2 to 4.
[0138] According to this advertising-related service providing system, the advertisement content associated with the identification information for identifying the product is identified from the manuscript data to determine consistency, so that the manuscript data can be accurately identified.
[0139] (Item 6) the calibration machine learning means performs machine learning specialized for each group (e.g., S129 to S135) by performing machine learning based on data grouped by predetermined attributes (e.g., business type, region, season (spring, summer, fall, winter)); An advertising-related service providing system described in any one of items 1 to 5, wherein the proofreading part determination means determines the consistency by reflecting the results of machine learning specialized for each group by the proofreading machine learning means (e.g., S84, S86).
[0140] According to this advertising-related service provision system, consistency determination processing is performed by reflecting the results of machine learning specialized for each group grouped by specified attributes, making it possible to utilize the results of machine learning that match the attributes of the group in the consistency determination processing.
[0141] (Item 7) A user-side advertising facility (e.g., a user terminal 74 connected to an internal LAN 84 of an advertisement creation company 31) that can be connected to an advertisement-related service providing system (e.g., a data center 44) that provides services related to advertisement creation, an advertisement providing means (e.g., S183) for providing an advertisement to the advertisement-related service providing system; manuscript data providing means (e.g., S181) for providing manuscript data used in creating the advertisement to the advertisement-related service providing system; a proofreading portion determining means provided on the advertisement-related service providing system side determines consistency between the advertisement and the manuscript data based on the advertisement data provided by the advertisement providing means and the manuscript data provided by the manuscript data providing means, and determines portions of the advertisement that need proofreading, and a receiving means (e.g., S184) receives a service based on the determination result from the advertisement-related service providing system side; The proofreading part determination means determines the consistency by reflecting the learning results (e.g., learned flyer feature 1, advertising value set boundary learned model 2, learned product feature 20, product image recognition learned model 18, character position learned model 19, OCR learned model 17, word classification learned model 32) obtained by the proofreading machine learning means to determine the consistency between the advertisement and the manuscript data (e.g., S87, S88, S95, S96, S98, S99, S104, S105, S108, S116, S117, S111, S112, S116, S117).
[0142] This user-side advertising equipment reflects the learning results of the proofreading machine learning means to determine the consistency between the advertisement and the manuscript data, and identifies the parts of the advertisement that need to be proofread, allowing users to enjoy services based on the results of this determination, thereby allowing them to receive services that reduce labor costs by minimizing human work as much as possible.
[0143] The embodiments disclosed herein should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined by the claims, not by the above description, and is intended to include all modifications within the meaning and scope of the claims. [Explanation of symbols]
[0144] 1 Flyer feature trained model, 2 Advertising value set boundary trained model, 3 Advertising value set boundary learning, 4 Training for product image identification, 14 Server, 16 OCR training, 17 OCR trained model, 18 Trained model for product image identification, 19 Character position trained model, 20 Trained product features, 22 Product feature learning, 23 Flyer feature learning, 24 Manuscript data, 27 Flyer, 28 Character position learning, 29 Trained storage unit using training data in the practical phase, 32 Word classification trained model, 33 Machine learning for advertising-related services, 34 Machine learning means, 35 Advertising-related service provision means, 50 Trained model.
Claims
1. A storage means for storing price data and fluctuation factor data that are factors influencing price fluctuations, The system includes a service provision means that predicts fluctuating prices based on the price data and fluctuation factor data accumulated by the accumulation means, and provides a service that reflects the prediction, The service provision means includes an advertising creation means for creating web advertisements with prices that reflect the forecast, and the service provision system provides the web advertisements created by the advertising creation means.
2. The service provision system according to claim 1, wherein the service provision means predicts fluctuating prices using an AI model that has been machine-learned by the machine learning means.
3. The service provision system according to claim 2, further comprising an advertising effectiveness measuring means for measuring the effectiveness of an advertisement created by the advertising creation means.
4. The service provision system according to claim 3, wherein the advertisement creation means generates training data in which the predicted fluctuating price is input information and the advertisement content for which an effect of a predetermined level or higher is measured by the advertisement effectiveness measurement means is used as correct information, and creates advertisements using an AI model that has undergone supervised learning, which learns a function that maps the input to the correct answer as a regression problem.
5. The service provision system according to claim 3, wherein the advertisement creation means creates advertisements using an AI model that has undergone reinforcement learning, in which the agent is rewarded according to the measurement results by the advertisement effectiveness measurement means, and the agent learns a strategy to maximize the accumulation of the rewards.
6. A method for generating an AI model for a web advertisement in which the AI model predicts price fluctuations of an advertised target and displays a price that reflects the prediction, The process involves accumulating price data and data on factors that influence price fluctuations, An AI model generation method comprising the step of machine learning the AI model using machine learning means based on the price data and fluctuation factor data accumulated in the aforementioned accumulation step.
7. Further comprising an advertising effectiveness measurement step for measuring the effectiveness of the web advertisement, The aforementioned machine learning step is, A step of generating training data using the predicted fluctuating price as input information and the advertising content for which an effect of a predetermined level or higher has been measured in the advertising effectiveness measurement step as the correct information, The method for generating an AI model according to claim 6, comprising the step of performing supervised learning using the aforementioned training data, and learning a function that maps the input to the correct answer as a regression problem.
8. Further comprising an advertising effectiveness measurement step for measuring the effectiveness of the web advertisement, The method for generating an AI model according to claim 6, wherein the machine learning step includes a step in which the agent is rewarded according to the measurement results of the advertising effectiveness measurement step, and the agent performs reinforcement learning in which it learns a strategy to maximize the accumulation of the reward.
9. The AI model generation method according to claim 7 or 8, wherein the advertising effectiveness measurement step measures the advertising effectiveness using data including sales of the advertised target.