Information processing device, information processing method, and program

WO2026190948A1PCT designated stage Publication Date: 2026-09-17NTT DOCOMO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/009087
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2026-09-17

Smart Images

  • Figure JP2025009087_17092026_PF_FP_ABST
    Figure JP2025009087_17092026_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device according to the present invention comprises: an advertisement selection unit that acquires text information by inputting, into a VLM generative AI model, image data acquired by imaging a person present around the information processing device and a VLM prompt for acquiring pedestrian information in the image data, and selects an advertisement target on the basis of the text information and the current location of the information processing device; and an advertisement generation unit that determines information to be used in an advertisement on the basis of the advertisement target and the text information, and generates a prompt for generating advertisement copy on the basis of the text information and the information to be used in the advertisement.
Need to check novelty before this filing date? Find Prior Art

Description

Information Processing Apparatus, Information Processing Method, and Program

[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a program.

[0002] In recent years, a marketing method in which people wear advertisements and walk around the city to promote products or stores has attracted attention, and an increasing number of companies are launching mobile digital advertisements such as walking advertisements and car advertisements. For example, there is a new advertising service in which staff carry large monitors (digital signage) on their backs and walk around a city in a specified area and time slot to display advertisements. Since this service moves at the same eye level and speed as passersby, it has high visibility and can effectively appeal to a large number of people. In addition, because it is a unique form of advertisement, diffusion power on SNS can also be expected.

[0003] When a company or the like runs an advertisement, it is required to advertise products and services that match the age and gender of the users who view the advertisement. However, in the above-mentioned "walking advertisement", since the situation around the advertisement is constantly changing, it cannot respond to changes in situations such as the attributes of people present in the area and the attributes of surrounding advertisement viewers, so the advertising effect is insufficient, and there is room for improvement.

[0004] One aspect of the present disclosure provides an information processing apparatus, an information processing method, and a program that can generate an optimal advertisement according to the attributes of people present in an area and the attributes of surrounding advertisement viewers.

[0005] The information processing apparatus according to one aspect of the present disclosure includes: an advertisement selection unit that acquires text information by inputting image data obtained by photographing people existing around the information processing apparatus and a VLM (Vision Language Model) prompt for acquiring pedestrian information in the image data into a VLM generation AI model, and selects an advertisement target based on a current location of the information processing apparatus and the text information; and an advertisement generation unit that determines information to be used for an advertisement based on the advertisement target and the text information, and generates a prompt for generating an advertisement text based on the text information and the information to be used for the advertisement.

[0006] This figure shows an example of an advertising display device according to an embodiment. This figure shows an example of the case when the advertising display device according to an embodiment is installed. This figure shows an example of an advertising display device according to an embodiment. This figure shows an example of the case when the advertising display device according to an embodiment is installed. This figure shows an example of the configuration of the advertising system according to the first embodiment. This figure shows an example of the advertising DB according to an embodiment. This figure shows an example of the processing flow according to the first embodiment. This figure shows an example of location information according to an embodiment. This figure shows an example of statistical information using location data according to an embodiment. This figure shows an example of information collected from surrounding viewers according to an embodiment. This figure shows an example of the age and gender estimation results by image recognition according to the first embodiment. This figure shows an example of the age, gender, and other attribute estimation results by image recognition according to the first embodiment. This figure shows an example of the age and gender estimation results by speech recognition according to an embodiment. This figure shows an example of estimated attribute data acquired by the advertising display device according to the first embodiment. This figure shows an example of the calculation result of the accuracy rate between advertising information and each item of estimated attribute data according to an embodiment. This figure shows an example of the positional relationship between a store and an advertising display device according to the first embodiment. This figure shows an example of a prompt for generating an advertisement according to the first embodiment. This figure shows an example of an advertisement text generated by the generation AI according to the first embodiment. This figure shows an example of an advertisement according to the first embodiment. This figure shows an example of the advertising DB according to an embodiment. This figure shows an example of the configuration of the advertising system according to the second embodiment. This figure shows an example of a knowledge DB according to an embodiment. This figure shows an example of the processing flow according to the second embodiment. This figure shows an example of a VLM prompt according to the second embodiment. This figure shows an example of the execution result of the VLM generation AI according to the second embodiment. This figure shows an example of estimated attribute data acquired by the advertisement display device according to the second embodiment. This figure shows an example of a prompt for generating an advertisement according to the second embodiment. This figure shows an example of an advertisement text generated by the generation AI according to the second embodiment. This figure shows an example of an advertisement according to the second embodiment. This figure shows an example of the hardware configuration of the advertisement management device, statistical information management device, generation device and VLM processing device 50 according to the embodiment.

[0007] Hereinafter, an embodiment relating to one aspect of this disclosure will be described with reference to the drawings. Note that the embodiment described below is merely an example, and the embodiments to which this disclosure applies are not limited to the embodiments described below.

[0008] <Appearance of the advertising display device> Figure 1 shows an example of the appearance of advertising display devices 10a and 10b used in a mobile advertisement according to an embodiment of the present disclosure. The advertising display devices 10a and 10b may be, for example, digital signage that displays advertising information using a large display. The upper part of the display unit 131 is equipped with a camera 121, a microphone 122, and a speaker 132. The advertising display devices 10a and 10b are equipped with a shoulder strap 18 on the back so that a person carrying the advertising display devices 10a and 10b can carry them on their back.

[0009] Figure 2 shows an example of a person carrying the advertising display devices 10a and 10b shown in Figure 1, with the devices on their back. The person carrying the advertising display devices 10a and 10b places the shoulder straps 18 of the advertising display devices 10a and 10b over their shoulders and carries the devices on their back. The advertising display devices 10a and 10b are equipped with a large display unit 131 that dynamically displays digital content such as store advertisements, product introductions, and service information. As a result, people walking behind the advertising display devices 10a and 10b can view advertising information, such as videos, displayed on the display unit 131. The advertising display devices 10a and 10b also output audio of the digital content of the advertisements from a speaker 132. In addition, the advertising display devices 10a and 10b acquire information about the surrounding viewers using a camera 121 and a microphone 122.

[0010] Figure 3 shows an example of the appearance of advertising display devices 10a and 10b used in a mobile advertisement according to an embodiment of the present disclosure. The advertising display devices 10a and 10b are equipped with a medium-sized display for displaying advertising information on a backpack-shaped body. The configuration and functions of the devices are the same as those described in Figure 1.

[0011] Figure 4 shows an example of a person carrying the advertising display devices 10a and 10b shown in Figure 3, with the devices on their back. The operation and functions of the devices are the same as those described in Figure 2. The devices in Figures 3 and 4 are smaller than the devices in Figures 1 and 2, making them easier to carry and allowing for a wider range of movement.

[0012] <Configuration of the advertising system in the first embodiment> Figure 5 is a block diagram showing an example of the configuration of the advertising system 1a according to the first embodiment of the present disclosure. The advertising system 1a comprises a plurality of advertising display devices 10a, an advertising management device 20a, a statistical information management device 30, and a generation device 40. The devices 10a to 40 are interconnected by a network 60 and can communicate with each other.

[0013] The advertising display device 10a displays digital content such as store advertisements, product introductions, and service information used in mobile advertising. The advertising display device 10a can communicate with the advertising management device 20a and acquires digital content such as the advertising database 171, advertising video data, and advertising data from the advertising management device 20a. The advertising display device 10a selects the advertising information and advertising data to display from the acquired digital content such as the advertising database 171, advertising video data, and advertising data.

[0014] The advertising display device 10a acquires statistical information from the statistical information management device 30 when selecting advertising information and advertising data to display. The advertising display device 10a grasps the attributes of people present in the surrounding area from the location information acquired from the location acquisition unit 14, the age and gender of surrounding viewers estimated by the viewer estimation unit 113a, the statistical information acquired from the statistical information management device 30, and the collected surrounding viewer information, and selects advertising information and advertising data that are optimal for the attributes of people present in the surrounding area.

[0015] Furthermore, the advertising display device 10a generates prompts for generating advertising text based on the current location information, the location information of the advertised store, and information regarding the relative positional relationship. A prompt is a request, such as an instruction or question, that the user inputs to the generation AI model. The generation AI model is composed of text generation models such as LLM (Large-Scale Language Model) and outputs an execution result (response) based on the input data request, such as the given prompt. An LLM is a model that exhibits natural language capabilities by learning from a large amount of data through a large number of parameters, based on a transformer architecture.

[0016] The ad display device 10a sends a prompt to the generation device 40 and receives the ad text, which is the result of processing the prompt returned from the generation device 40. Based on the selected ad information and the generated ad text, the ad display device 10a generates and displays an ad.

[0017] The advertising management device 20a manages the advertising DB 171, which is a database of advertising information distributed to the advertising display device 10a, as well as digital content such as advertising videos and advertising data. The advertising display device 10a can communicate with the advertising management device 20a and acquires the advertising DB 171 and digital content such as advertising videos and advertising data from the advertising management device 20a.

[0018] The statistical information management device 30 estimates real-time demographic statistics (including population distribution, population composition, population trends, population inflow and outflow, etc., collectively referred to as "statistical information") for a region (e.g., a 500m mesh) every hour, based on operational data of mobile phone base stations, which are operational data of the mobile phone network, and attribute data, which are data related to the attributes of mobile phone users. The population composition of this information is based on men and women aged 15 to 79. The statistical information management device 30 aggregates the population composition values ​​every hour. For example, if a person stays in a certain area for 15 minutes, the statistical value is obtained as 1 / 4 of the population. This information is subject to statistical disclosure control (de-identification and concealment processing).

[0019] The generation device 40 includes a Generative AI (Generative AI) model composed of text generation models such as an LLM (Large-Scale Language Model). The Generative AI model of the generation device 40 receives requests to the Generative AI model that are input, executes processing by the Generative AI model, and outputs an execution result (response) to the prompt.

[0020] Furthermore, the advertising management device 20a, the statistical information management device 30, the generation device 40, and other devices connected via the network 60 may be computer devices consisting of server devices or clouds, and provide services and resources to other computer devices via the network 60.

[0021] [Configuration of the advertising display device of the first embodiment] As shown in Figure 5, the advertising display device 10a according to the first embodiment comprises a control unit 11a, an input unit 12, an output unit 13, a position acquisition unit 14, a communication unit 15, a main memory unit 16, and an auxiliary memory unit 17a. The control unit 11a, the main memory unit 16, and the auxiliary memory unit 17a may be composed of a small computer such as an NUC (Next Unit of Computing) or an embedded AI platform. The advertising display devices 10a and 10b may also be referred to as information processing devices.

[0022] The control unit 11a is composed of a central processing unit (CPU) or the like. The control unit 11a controls the operation of the entire advertising display device 10a. The control unit 11a may also be referred to as a processing unit, processor, or controller. The control unit 11a may be composed of hardware such as a microprocessor, digital signal processor (DSP), application specific integrated circuit (ASIC), programmable logic device (PLD), or field programmable gate array (FPGA), and some or all of each functional block may be realized by such hardware. The control unit 11a further includes various functional units such as an image recognition unit 111a, a voice recognition unit 112, a viewer estimation unit 113a, an advertising selection unit 114a, and an advertising generation unit 115a.

[0023] The image recognition unit 111a detects faces from image data acquired by the camera 121 and extracts feature points (such as the positions of the eyes, nose, mouth, and chin, and the contour). If the image data contains the faces of multiple people, the image recognition unit 111a processes each of the faces individually. OpenCV, MediaPipe, YOLO, etc., may be used as the face detection library for the image recognition unit 111a.

[0024] The speech recognition unit 112 removes noise and excludes silent portions from the audio data acquired by the microphone 122 to detect speech segments (VAD: Voice Activity Detection). The speech recognition unit 112 performs speaker separation (diarization) to identify the voices of multiple people. Then, the speech recognition unit 112 extracts acoustic features such as Mel-Frequency Cepstrum Coefficients (MFCC), fundamental frequency, formant frequency, volume, and speech rate from each of the separated voices. Libraries such as Praat, Librosa, and OpenSMILE may be used for acoustic feature extraction by the speech recognition unit 112.

[0025] The speech recognition unit 112 uses a speech recognition engine to convert the viewer's speech into text. Whisper, Speech-to-Text, or the like may be used as the speech recognition engine of the speech recognition unit 112. The speech recognition unit 112 detects specific words and phrases from the transcribed speech and filters the speakers. For example, the speech recognition unit 112 removes speakers (sound sources) of words and phrases unrelated to the viewer's speech (such as automated voice announcements in public transport or facilities, "The doors are closing," "Watch your step," etc.). The speech recognition unit 112 may also process the speech to select speakers who uttered specific words and phrases highly relevant to the act of viewing the advertisement (for example, reactions such as "This ad is interesting!" or words related to the content of the advertisement). The speech recognition unit 112 may also extract characteristic speech content and frequently occurring speech content (sentence endings, interjections, words) from the viewer's speech.

[0026] The viewer estimation unit 113a inputs individual image data of multiple people recognized by the image recognition unit 111a into a pre-trained machine learning model to estimate age, gender, or outward appearance characteristics such as clothing. This pre-trained machine learning model that performs image recognition mainly operates using deep learning technology. The pre-trained machine learning model that performs image recognition extracts features from the image data through convolutional layers and pooling layers, and uses the extracted image features to estimate age, gender, or outward appearance characteristics such as clothing. ResNet, VGG16, etc., may be used as the pre-trained machine learning model that performs age estimation, gender classification, or classification of outward appearance characteristics such as clothing from image data in the viewer estimation unit 113a. Similarly, the viewer estimation unit 113a inputs speaker-specific speech data recognized by the speech recognition unit 112 into a pre-trained machine learning model to estimate age and gender. This pre-trained machine learning model uses the acoustic features of the speech data to perform age estimation and gender classification. The viewer estimation unit 113a may use a pre-trained machine learning model such as TensorFlow, PyTorch, or DeepVoice3 to perform age estimation and gender classification from the audio data.

[0027] The ad selection unit 114a calculates a relevance rate for each ad in the ad database 171 based on location information obtained from the location acquisition unit 14, the age and gender of surrounding viewers estimated by the viewer estimation unit 113a, or other attributes, statistical information obtained from the statistical information management device 30, and the collected surrounding viewer information, and selects the ad with the highest relevance rate as the "advertising target (target store)".

[0028] The ad generation unit 115a collects and calculates information regarding the positional relationship with the advertised target (target store), and generates a prompt for generating ad text based on the information regarding the positional relationship with the advertised target (target store). The ad generation unit 115a transmits the generated prompt for generating ad text to the generation device 40 and obtains the execution result of the prompt. The ad generation unit 115a generates ad information from the obtained ad text and video file and displays it on the display unit 131.

[0029] The input unit 12 inputs information used by the advertising display device 10a to the control unit 11a. The input unit 12 includes a camera 121, a microphone 122, and an input device 123.

[0030] Camera 121 captures images of surrounding viewers who are viewing the digital content (advertising data) on the display unit 131 and outputs image data to the control unit 11a. Camera 121 can be any imaging device capable of shooting video, and is installed above or around the display unit 131.

[0031] The microphone 122 collects the voices of viewers in the vicinity who are viewing the digital content (advertising data) on the display unit 131 and outputs the voice data to the control unit 11a. The microphone 122 can be any device capable of acquiring sound and is installed on or around the display unit 131. The microphone 122 may be a directional microphone, and the use of a directional microphone reduces noise and ambient noise, allowing for effective collection of the voices of viewers viewing the display unit 131.

[0032] The input device 123 inputs information to the control unit 11a. The input device 123 can be any device capable of inputting information to the control unit 11a, and may consist of, for example, a small keyboard, a touch panel, or a mouse.

[0033] The output unit 13 converts the information output from the control unit 11a into images and sound and outputs them. The output unit 13 includes a display unit 131 and a speaker 132.

[0034] The display unit 131 dynamically displays digital content (advertisements) and text information, such as store advertisements, product introductions, and service information, output from the control unit 11a. The display unit 131 may be a liquid crystal display (LCD), LED display, OLED display, etc., and is composed of a large or medium-sized display.

[0035] The speaker 132 can be any device capable of outputting sound, and it outputs the sound of digital content related to advertising.

[0036] The position acquisition unit 14 is a functional unit for acquiring position, incorporating a GPS receiver (Global Positioning System Receiver) or a GNSS (Global Navigation Satellite System) receiver. For example, using the GPS receiver or GNSS receiver functions of the position acquisition unit 14, it receives signals from multiple GPS satellites or GNSS satellites and acquires the latitude and longitude of the location of the advertising display device 10a. If the position acquisition unit 14 cannot acquire signals from GPS satellites or GNSS satellites, it may measure the signal strength from multiple base stations or APs (access points) to estimate the distance and acquire the location of the advertising display device 10a. Alternatively, the position acquisition unit 14 may acquire the location of the advertising display device 10a by using a location information provision service using LCS (Location Service) provided by the mobile phone network. The position acquisition unit 14 may also be equipped with other position detection functions for acquiring position.

[0037] The communication unit 15 performs communication between computers via a wireless or wired network. The communication unit 15 also performs communication between the advertising display device 10a, the advertising management device 20a, the statistical information management device 30, and the generation device 40.

[0038] The main memory unit 16 is the main memory of a computer device configured as random access memory (RAM) or other memory (ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM)), and is an area for temporarily storing programs and their data that are currently running. Programs and data are read from the auxiliary storage unit 17a to the main memory unit 16, and the main memory unit 16 stores program instructions and temporary data during calculation processing that are necessary for the computer device to operate.

[0039] The auxiliary storage unit 17a holds the OS (operating system), applications and programs for realizing each functional unit, user data, etc. The auxiliary storage unit 17a can be any device that can retain its contents even when power is not supplied, and may consist of, for example, a hard disk drive (HDD), a solid state drive (SSD), an optical disc such as a CD-ROM (Compact Disc ROM), a flexible disc, a magneto-optical disc (e.g., compact disc, digital multipurpose disc, Blu-ray® disc), a smart card, flash memory (e.g., card, stick, key drive), a floppy® disc, a magnetic strip, etc. The auxiliary storage unit 17a stores the advertising DB 171 acquired from the advertising management device 20a, as well as digital content such as advertising videos and advertising data.

[0040] The advertising database 171 stores advertising information, such as that shown in Figure 6. Each record (row) in the advertising database 171 constitutes information about one advertisement (advertising information). The advertising information in the advertising database 171 includes, for example, advertising ID, store name, location, latitude and longitude, business type, target attributes, video advertisement (file name), store / series code, and trade area. The trade area indicates the geographical range in which the store can attract customers and varies depending on the surrounding competitors and the price range of other stores. Figure 6 shows stores in the food and beverage service industry, but it is not limited to the food and beverage service industry; it can also include advertising information for stores in other industries. Other industries can include manufacturing, electricity and gas, information and communication, transportation, retail, finance, insurance, real estate, professional and technical services, accommodation, entertainment, education, medical care, and welfare, and is not limited to any particular industry.

[0041] <Flowchart of Operation of the Advertising Display Device in the First Embodiment> Figure 7 is a flowchart of an example of the operation of the advertising display device 10a. The steps of this flowchart will be described below.

[0042] In step S1, the camera 121 of the advertising display device 10a takes pictures of the area around the advertising display device 10a and acquires the captured image data. The camera 121 transfers the captured image data to the control unit 11a.

[0043] In S2, the image recognition unit 111a performs image recognition processing on the image data acquired by the camera 121 to detect faces and extract feature points (eyes, nose, mouth, and contour). If the image data contains the faces of multiple people, the image recognition unit 111a processes each face individually.

[0044] In step S3, the image recognition unit 111a detects, as a "viewer", a person recognized within a certain angular range and a certain distance range relative to the display unit 131 among people walking behind the advertisement display device 10a recognized in step S2. As another method, the image recognition unit 111a may estimate the orientation of the face in a three-dimensional space by a PnP (Perspective n Point) algorithm from a plurality of detected feature points of the face (such as eyes, nose, mouth, jaw position, and contour), and detect a "viewer" who is estimated to be viewing the display unit 131.

[0045] In step S4, the microphone 122 of the advertisement display device 10a collects sound around the advertisement display device 10a and obtains audio data. The microphone 122 outputs the collected audio data to the control unit 11a. If the microphone 122 is a directional microphone, it may be configured to collect the voice of a person within a certain angular range and a certain distance range relative to the display unit 131.

[0046] In step S5, the voice recognition unit 112 performs voice recognition processing on the audio data acquired by the microphone 122 to extract acoustic features (such as MFCC, formant frequency, volume, and speech rate). When the audio data contains voices of a plurality of people, the voice recognition unit 112 performs speaker separation and performs recognition processing on the voices of the plurality of people respectively.

[0047] In step S6, the voice recognition unit 112 converts the utterance of the viewer into text using a voice recognition engine. The voice recognition unit 112 detects specific words and phrases from the text-converted utterance and performs filtering processing on speakers. For example, the voice recognition unit 112 may remove speakers (sound sources) of irrelevant words and phrases (such as automatic voice guidance), and perform processing to select a speaker who has uttered specific words and phrases highly relevant to the advertisement viewing behavior. The voice recognition unit 112 may extract characteristic utterance content and frequently-occurring utterance content (sentence endings, interjections, words) from the utterance of the viewer.

[0048] In step S7, the position acquisition unit 14 of the advertisement display device 10a acquires the latitude and longitude at which the advertisement display device 10a is located as position information. The position acquisition unit 14 transfers the acquired position information to the control unit 11a. The control unit 11a performs reverse geocoding on the acquired latitude and longitude information to convert it into an address notation in Japan. For example, latitude, longitude, and the current position are acquired as shown in FIG. 8.

[0049] In step S8, the control unit 11a of the advertisement display device 10a acquires statistical information using area presence data around the current position from the statistical information management device 30 based on the position information acquired in step S7. For example, the control unit 11a acquires statistical information as shown in FIG. 9. The statistical information shown in FIG. 9 is a population composition ratio estimated to exist within a 500 m square (mesh) around the current position at a certain time (in 1-hour units).

[0050] In step S9, the control unit 11a of the advertisement display device 10a collects information from smartphones of surrounding viewers using wireless communication. For example, as shown in FIG. 10, the collected information includes member information registered on an application, action history information (purchase history and other history information), and the like. These demographic pieces of information are acquired based on a contract agreement (opt-in) for the use of personal information.

[0051] In S10, the viewer estimation unit 113a inputs the individual image data recognized by the image recognition unit 111a for each "viewer" to a trained machine learning model to estimate their age and gender. For example, in the example shown in Figure 11, the viewer estimation unit 113a estimates that among the 12 viewers of the advertisement detected within a predetermined time period in the past, the largest group was male in his 20s. As shown in Figure 12, the image recognition unit 111a may also acquire results for each attribute, including characteristics such as clothing, in addition to age and gender, and use them as estimated attribute data. Furthermore, the viewer estimation unit 113a inputs the voice data for each speaker recognized by the voice recognition unit 112 to a trained machine learning model to estimate their age and gender. For example, in the example shown in Figure 13, the viewer estimation unit 113a estimates that among the 5 viewers of the advertisement detected within a predetermined time period in the past, the largest group was male in his 20s.

[0052] In S11, the ad selection unit 114a acquires estimated attribute data as shown in Figure 14 (the information in Figure 14 is collectively referred to as "estimated attribute data"). Based on the location information (latitude and longitude, current location in Japanese addresses) acquired from the location acquisition unit 14, the age and gender of surrounding viewers estimated by the viewer estimation unit 113a (image recognition and voice recognition), statistical information using location data acquired from the statistical information management device 30, and the collected surrounding viewer information, the ad selection unit 114a calculates the precision rate for each ad in the ad database 171. The ad selection unit 114a calculates the precision rate by quantifying how well multiple items in the advertising information match the attributes of people estimated to be present in the surrounding area, such as the degree of agreement between the "target attributes" in the advertising information of the advertising DB 171 and the "age and gender of surrounding viewers" in the estimated attribute data, the degree of agreement between the "market area" in the advertising information of the advertising DB 171 and the "current location" in the estimated attribute data, and the degree of agreement between the "location" in the advertising information of the advertising DB 171 and the "current location" in the estimated attribute data. The ad selection unit 114a may also calculate the precision rate by averaging the degree of agreement for each item between the advertising information of the advertising DB 171 and the estimated attribute data. For example, the ad selection unit 114a calculates the precision rate for each item between the advertising information of the advertising DB 171 and the estimated attribute data (see Figure 15). The ad selection unit 114a may also obtain the precision rate as the average value calculated by weighting each item of the estimated attribute data.

[0053] In S12, the ad selection unit 114a selects the ad information with the highest relevance to each piece of acquired information (location information, age and gender of surrounding viewers, statistical information, or surrounding viewer information) as the "advertised target (target store)". In other words, the ad selection unit 114a grasps the attributes of people present in the surrounding area and selects the ad information that best matches the attributes of people present in the surrounding area. In the example shown in Figure 15, the ad selection unit 114a selects the ad information with ad ID A001 as the "advertised target". Note that when the ad selection unit 114a selects the ad information with the highest relevance as the "advertised target", it may use the items with the highest priority in the estimated attribute data to make the decision. For example, the priority order for which items of the estimated attribute data to prioritize may be determined as follows.

[0054] Priority 1: Current location and estimated age and gender of nearby viewers using image recognition. Priority 2: Current location and estimated age and gender of nearby viewers using voice recognition. Priority 3: Current location and information collected from nearby viewers using wireless communication. Priority 4: Statistical information within a 500m mesh using location data.

[0055] In S13, the ad generation unit 115a collects and calculates information regarding the positional relationship with the advertised target (target store). For example, let's assume that the ad display device 10a and the advertised target (target store) are in the positional relationship shown in Figure 16. In this case, the ad generation unit 115a detects the direction of travel of the ad display device 10a from the time-series information of the position information of the ad display device 10a acquired from the position acquisition unit 14. The ad generation unit 115a also calculates the distance between the ad display device 10a and the advertised target (target store), and the direction in which the advertised target (target store) is located relative to the ad display device 10a, from the current position information of the ad display device 10a acquired from the position acquisition unit 14 and the position information of the advertised target selected by the ad selection unit 114a.

[0056] In S14, the ad generation unit 115a generates prompts for generating ad text based on the target of the advertisement (target store), the distance between the ad display device 10a and the target of the advertisement (target store), the orientation of the target of the advertisement (target store) relative to the ad display device 10a, and the direction of travel of the ad display device 10a (see Figure 17). The ad generation unit 115a may also generate prompts to generate ad text according to conditions. For example, if the distance from the current location to the target of the advertisement (target store) is large, or if the target of the advertisement (target store) is in the opposite direction to the direction of travel, ad text offering coupons or service benefits may be added.

[0057] In S15, the ad generation unit 115a sends a prompt to the generation device 40 for generating the generated ad text. The generation AI model of the generation device 40 receives the prompt for generating the generated ad text, generates the ad text, and outputs it. The ad generation unit 115a obtains the execution result of the prompt (the generated ad text) from the generation device 40 (see Figure 18).

[0058] In S16, the ad generation unit 115a generates an advertisement by combining the generated ad text with digital content such as an ad video and ad data. The ad generation unit 115a then transfers the generated advertisement to the display unit 131 and displays the advertisement (see Figure 19).

[0059] <Modification of the First Embodiment> In steps S11-12 of Figure 7, if the "viewer" detected by the image recognition unit 111a is a single person, and the attributes of the person estimated by the viewer estimation unit 113a are estimated to be "one person", the advertising selection unit 114a may select an "advertising target" considering the age and gender of the target viewer, as well as the fact that the target viewer is "one person". The advertising DB 171 may store advertising information such as that shown in Figure 20. If the current location of the advertising display device 10a is Shinjuku, Shinjuku-ku, the advertising selection unit 114a may refer to the target attribute category in the advertising DB 171 that matches the estimated person's attributes ("in their 30s", "female", and "one person"), and select a store for "one person" as the "advertising target".

[0060] In S14 of Figure 7, the ad generation unit 115a may generate a prompt to change the expression of the ad text according to the age and gender of the surrounding audience. For example, the ad generation unit 115a may customize the ending of the ad text to a specific form according to the age and gender of the surrounding audience, such as "nominalization," "polite style" ("desu," "masu"), "plain style" ("dearu," "da"), or "conversational style" ("dayo," "dane"). For example, if there are many older people or people dressed formally (see Figure 12), the ad generation unit 115a may use polite language; if there are many younger people or people dressed casually, it may use conversational style; and if there is a mix of different generations, it may use nominalization. In this way, expressions that are well-received by the majority of the audience can be selectively adopted.

[0061] In S14 of Figure 7, the ad generation unit 115a may extract specific endings and expressive features from the utterance transcribed by the speech recognition unit 112 and generate prompts to generate ad text in an expression that matches the speaker's characteristics. For example, depending on the location or the speaker's characteristics recognized in S6, the ad generation unit 115a may generate prompts to generate ad text using dialects ("dabe", "yaro", "yanaa", "dosu", "yaken"), unique endings ("ja", "degozaru", "nya", "dayoon"), etc. By changing the endings, the nuance and feeling of the ad text can be greatly altered, making the ad text more relatable.

[0062] (Effects) As described above, the advertising display device of the first embodiment estimates the attributes of viewers walking around or behind the advertising display device, selects appropriate advertising targets (target stores) based on location and attributes, and generates appropriate advertising text that matches the location and attributes in real time, thereby enabling the display of impactful and effective advertisements that accurately convey the message.

[0063] <Configuration of the advertising system of the second embodiment> Next, the advertising system 1b according to the second embodiment of this disclosure will be described. Figure 21 is a block diagram showing an example of the configuration of the advertising system 1b according to the second embodiment of this disclosure. In describing the second embodiment, the explanation of matters similar to those of the first embodiment (configurations with the same figure numbers) will be omitted, and the explanation will focus on the parts that have been changed from the first embodiment.

[0064] In the second embodiment, the "advertising system 1b" is a replacement of the "advertising display device 10a" with the "advertising display device 10b", the "advertising management device 20a" with the "advertising management device 20b", and a "VLM processing device 50" is newly added. The "advertising display device 10b" has some functions changed or added compared to the "advertising display device 10a". Note that the "statistical information management device 30" and "generation device 40" of the "advertising system 1b" are the same as those of the "advertising system 1a", so their description is omitted.

[0065] The advertising management device 20b manages the advertising information database 171 and knowledge database 172, which are databases of advertising information to be distributed to the advertising display device 10b, as well as digital content such as advertising videos and advertising data. The advertising display device 10b can communicate with the advertising management device 20b and obtains the advertising database 171 and knowledge database 172 from the advertising management device 20b.

[0066] The advertising management device 20b differs from the aforementioned advertising management device 20a in that it manages the knowledge database 172.

[0067] The VLM processing unit 50 is equipped with a Generative AI (Generative AI) model composed of a generative model such as a VLM (Vision Language Model). The VLM model is a generative AI model that can understand both images (Vision) and text (Language) by embedding them in a common vector space and learning the relationship between the two, thereby enabling it to process them in relation to each other. The VLM generative AI model of the VLM processing unit 50 receives requests to the VLM generative AI model that are input, executes processing by the VLM generative AI model, and outputs an execution result (response) to the prompt. The VLM processing unit 50 may be a computer device composed of server devices or a cloud, and provides services and resources to other computer devices via a network.

[0068] [Configuration of the Advertisement Display Device of the Second Embodiment] Figure 21 shows a block diagram illustrating an example of the configuration of the "Advertisement Display Device 10b" according to the second embodiment of this disclosure. Compared to the "Advertisement Display Device 10a" according to the first embodiment, the "Advertisement Display Device 10b" according to the second embodiment is configured such that the "control unit 11a" is replaced by the "control unit 11b", the "auxiliary storage unit 17a" is replaced by the "auxiliary storage unit 17b", the "image recognition unit 111a" is replaced by the "VLM instruction unit 111b", the "viewer estimation unit 113a" is replaced by the "viewer estimation unit 113b", the "advertisement selection unit 114a" is replaced by the "advertisement selection unit 114b", and the "advertisement generation unit 115a" is replaced by the "advertisement generation unit 115b", and a "knowledge DB 172" is added to the "auxiliary storage unit 17b". Note that the "input unit 12", "output unit 13", "position acquisition unit 14", "communication unit 15", "main memory unit 16", "voice recognition unit 112", "camera 121", "microphone 122", "input device 123", "display unit 131", and "speaker 132" of the "advertising display device 10b" are the same as those of the "advertising display device 10a", so their explanation will be omitted.

[0069] The following describes the various functions and storage units of the advertising display device 10b.

[0070] The VLM instruction unit 111b generates prompts for questions to detect the situation, attributes, and characteristics of multiple people in the surrounding area in the image data acquired by the camera 121. The VLM instruction unit 111b estimates that people walking behind the advertising display device 1 captured in the image data acquired by the camera 121 who are walking in the same direction as the advertising display device 10b, walking at approximately the same speed, and within a certain distance range from the display unit 131 are "viewers" of the advertising display device 10b. The VLM instruction unit 111b generates prompts to acquire information on pedestrians who are estimated to be walking behind the advertising display device 10b, whose direction of travel is approximately the same as the advertising display device 10b, and who are within a certain distance range from the display unit 131. The VLM instruction unit 111b may also acquire information on people whose walking speed is approximately the same as the advertising display device 10b by generating prompts for questions to detect the situation, attributes, and characteristics of multiple people using multiple image data acquired at multiple different timings. The VLM instruction unit 111b transmits the image data acquired by the camera 121 and the generated prompt to the VLM processing unit 50 and obtains the execution result of the prompt. From the obtained execution result, the VLM instruction unit 111b obtains text information such as the estimated age, gender, situation, actions, or behaviors of the viewer (or person presumed to be a viewer) in the image data acquired by the camera 121.

[0071] The viewer estimation unit 113b inputs the speaker-specific voice data recognized by the speech recognition unit 112 into a trained machine learning model to estimate the age and gender of the viewer.

[0072] The viewer estimation unit 113b, unlike the viewer estimation unit 113a described above, has only the function of estimating age and gender from audio data, excluding the function of estimating age, gender classification, or classification of external characteristics such as clothing from image data.

[0073] The ad selection unit 114b calculates the relevance rate for each ad in the ad database 171 based on location information obtained from the location acquisition unit 14, text information such as the estimated age, gender, situation, actions, or behaviors of surrounding viewers obtained from the VLM instruction unit 111b, the age, gender, or other attributes of surrounding viewers estimated by the viewer estimation unit 113b, statistical information obtained from the statistical information management device 30, and the collected surrounding viewer information, and selects the ad with the highest relevance rate as the "ad target (target store)".

[0074] The ad selection unit 114b differs from the ad selection unit 114a described above in that, instead of using the age and gender of surrounding viewers estimated from image data, it uses text information such as the estimated age, gender, situation, actions, or behaviors of surrounding viewers obtained from the VLM instruction unit 111b.

[0075] The ad generation unit 115b refers to the "registered services" targeted by the advertised target (target store) from the acquired knowledge DB 172. The ad generation unit 115b determines the "registered services" that are highly relevant to the text information such as the estimated age, gender, situation, actions, or behaviors of the surrounding audience acquired from the VLM instruction unit 111b as "information to be used in the advertisement."

[0076] The ad generation unit 115b generates prompts for generating ad copy based on the determined registration service. This function may be implemented using RAG (Retrieval-Augmented Generation), a type of prompt augmentation technology. A vector search engine may be constructed using libraries such as FAISS, Pinecone, and ScaNN, and RAG may be constructed by utilizing the functions of LangChain and Haystack in the library. RAG is used to provide necessary information and knowledge to complement information and knowledge that cannot be covered by a single generation AI. Specifically, when a user makes a request to the generation AI, it is known that a search system searches for similar documents in external data in advance and requests those documents together with the generation AI. This allows the generation AI to generate a response based on a specific document.

[0077] The ad generation unit 115b sends a prompt to the generation device 40 to generate the generated ad text and obtains the result of executing the prompt. The ad generation unit 115b generates ad information from the obtained ad text and video file and displays it on the display unit 131.

[0078] The ad generation unit 115b differs from the ad generation unit 115a described above in that, instead of generating prompts for generating ad text based on information regarding the positional relationship with the advertised target (target store), it generates prompts for generating ad text based on the determined "registered service".

[0079] The auxiliary storage unit 17b stores the advertising DB 171, the knowledge DB 172, and digital content such as advertising videos and advertising data acquired from the advertising management device 20b.

[0080] Knowledge DB 172 stores advertising information, such as that shown in Figure 22. Each record (row) in Knowledge DB 172 represents information about one registered service. The advertising information in Knowledge DB 172 includes, for example, a service ID, store / series code, business type, target attributes (age and gender, and category), time of day, registered service, and keywords. In Knowledge DB 172 in Figure 22, services in the food and beverage service industry are stored, but it is not limited to the food and beverage service industry; it may also store information such as product names, item names, campaign names, and service names for other industries. Other industries may include manufacturing, electricity and gas, information and communication, transportation, retail, finance, insurance, real estate, professional and technical services, accommodation, entertainment, education, medical care, and welfare, and is not limited to any particular industry.

[0081] <Flowchart of Operation of the Advertising Display Device in the Second Embodiment> Figure 23 is a flowchart of an example of the operation of the advertising display device 10b. The steps of this flowchart will be explained below.

[0082] The flowchart of the operation of the "advertising display device 10b" in Figure 23 according to the second embodiment differs from the flowchart of the operation of the "advertising display device 10a" in Figure 7 according to the first embodiment in that the processes from S22 to S23, S27, and S31 to S36 are different.

[0083] The steps related to S21, S24-S26, and S28-S30 in Figure 23 are the same as those in S1, S4-S9 in Figure 7 described above, so their explanation will be omitted.

[0084] In S22, the VLM instruction unit 111b generates a VLM prompt to acquire information on pedestrians who are estimated to be walking behind the advertising display device 10b, whose direction of travel is approximately the same as the advertising display device 10b, and whose distance from the display unit 131 is within a certain range (see Figure 24).

[0085] In S23, the VLM instruction unit 111b transmits the image data acquired by the camera 121 and the generated VLM prompt to the VLM processing unit 50. The VLM generation AI model of the VLM processing unit 50 receives the VLM prompt request to the input VLM generation AI model, executes processing by the VLM generation AI model, and outputs an execution result (response) for the VLM prompt. The VLM instruction unit 111b acquires the execution result of the VLM prompt. From the obtained execution result, the VLM instruction unit 111b acquires text information that describes the estimated age, gender, situation, actions or behaviors of surrounding viewers in the image data acquired by the camera 121 (see Figure 25).

[0086] In S27, the viewer estimation unit 113b inputs the speaker-specific voice data recognized by the speech recognition unit 112 into a trained machine learning model to estimate the age and gender.

[0087] In S31, the ad selection unit 114b acquires estimated attribute data as shown in Figure 26 (the information in Figure 26 is collectively referred to as "estimated attribute data"). Based on the location information (latitude and longitude, current location in Japanese address notation) acquired from the location acquisition unit 14, text information such as the estimated age, gender, situation, actions or behaviors of surrounding viewers acquired by the VLM instruction unit 111b (hereinafter referred to as "information acquired by VLM"), the age and gender of surrounding viewers estimated by the viewer estimation unit 113b (voice recognition), statistical information using location data acquired from the statistical information management device 30, and the collected surrounding viewer information, the ad selection unit 114b calculates the precision rate for each ad information in the ad database 171. The ad selection unit 114b may calculate the precision rate by averaging the degree of match for each item (location, target attributes). Alternatively, the ad selection unit 114b may calculate the precision rate by adding a weighted value to the degree of match between the information obtained by VLM and the information in the ad database, based on the degree of match for each other item of the estimated attribute data (estimated age and gender of surrounding viewers (voice recognition), statistical information within a 500m mesh using location data, and collected surrounding viewer information).

[0088] In S32, the ad selection unit 114b selects the ad information with the highest relevance rate from the ad DB 171 (see Figure 6) to the information obtained by VLM (and other items of the estimated attribute data) as the "ad target". From the ad information, the ad selection unit 114b selects the ad information for "Shibuya XX Store" with ad ID A002 as the "ad target".

[0089] In S33, the ad generation unit 115b determines from the knowledge database 172 (see Figure 22) that the registered services that are highly related to the information obtained by VLM will be used as the information to be used for the advertisement. First, the ad generation unit 115b refers to the registered services that match the identifier (store / series code: R-10) of the target of the advertisement (target store: Shibuya XX store) from the acquired knowledge database 172. Then, the ad generation unit 115b refers to the registered services that are highly related to the information obtained by VLM for the matching registered services (K001, K007, and K009). For example, the information obtained by VLM includes information about "high school students (approximately 15-18 years old)," and the service ID in the knowledge database 172 matches the target attribute "teenagers" of K001. Furthermore, the "information acquired by VLM" includes information about "going out with friends" and "a cheerful atmosphere," which matches the content of the keywords "share with friends" and "fun" in the knowledge database 172, where the service ID is K001. Therefore, the ad generation unit 115b decides to use the "discount campaign for U20s" of the registered service in the knowledge database 172, where the service ID is K001, as "information to be used for advertising."

[0090] In S34, the ad generation unit 115b generates a prompt for generating ad copy based on the determined "information to be used in the ad". Using the RAG function, the ad generation unit 115b references the registered service name "Discount campaign for U20" and the keywords "Share with friends, great value, fun, energetic" from the knowledge database to generate a prompt for generating ad copy (see Figure 27).

[0091] In S35, the ad generation unit 115b sends a prompt to the generation device 40 to generate the generated ad text. The generation AI model of the generation device 40 generates and outputs the ad text. The ad generation unit 115b obtains the execution result of the prompt (the generated ad text) (see Figure 28).

[0092] In S36, the ad generation unit 115b generates an advertisement by combining the generated ad text with digital content such as an ad video and ad data. The ad generation unit 115b then transfers the generated advertisement to the display unit 131 and displays the advertisement (see Figure 29).

[0093] <Modification of the Second Embodiment> In S32 of Figure 23, the ad selection unit 114b may select an "advertised target" from the ad DB 171 by considering the age and gender of the target audience, as well as whether the target audience is a "group" or "individual," from the "information acquired by VLM." The ad DB 171 may store advertising information such as that shown in Figure 20. If the current location of the ad display device 10a is "Shibuya, Shibuya-ku" and the "information acquired by VLM" is "in their 20s," "male," and "individual," the ad selection unit 114b may refer to the location and target attribute category of the ad DB 171 that matches the "information acquired by VLM" and select a store in "Shibuya, Shibuya-ku" that is "in their 20s," "male," and "individual" as the "advertised target."

[0094] In step S33 of Figure 23, the ad generation unit 115b may determine from the knowledge database 172 a "registered service" that is highly relevant to the time of day as "information to be used in the advertisement." For example, if a family restaurant is determined to be the target store for advertising, the ad generation unit 115b may refer to the time of day and determine the registered service as "information to be used in the advertisement" based on the matching time of day. For example, the ad generation unit 115b may refer to the identifier of the target store (store / series code: G-21) and the time of day and select "morning menu" for the morning hours, "half-price takeout" for the afternoon hours, and "house wine service" for the evening hours.

[0095] Furthermore, in steps S33 to S34 of Figure 23, the ad generation unit 115b may refer to customer attributes from the knowledge database 172, determine registered services as "information to be used in advertising" according to generation, and generate prompts to include keywords in the ad copy that reflect the consumption behavior and values ​​of each generation. For example, if a family restaurant is determined to be the target of advertising, the ad generation unit 115b may select the "Popular Menu Voting!" registered service for university students and young working adults (in their 20s) and generate prompts to include keywords such as "SNS share" and "trend" in the ad copy; select the "Family Set" registered service for families with children and generate prompts to include keywords such as "family-friendly" and "kids menu" in the ad copy; select the "Seasonal Organic Dinner" registered service for middle-aged people (in their 40s and 50s) and generate prompts to include keywords such as "organic" and "relaxing atmosphere" in the ad copy; and select the "Senior Day Discount" registered service for seniors (in their 60s) and generate prompts to include keywords such as "healthy" and "traditional taste" in the ad copy. The ad generation unit 115b may select services that take into account the consumer behavior and values ​​of each generation within the same store, and generate prompts that include keywords that resonate with each generation in the ad copy.

[0096] In steps S33 to S34 of Figure 23, the ad generation unit 115b may generate a prompt that selects a relevant registered service from the utterance transcribed by the speech recognition unit 112 based on the content of the speaker's conversation, and generates an ad copy that proposes a specific optimal menu from the knowledge DB acquired by the RAG function. For example, if a family restaurant is selected as the target store for advertising, the ad generation unit 115b may generate a prompt that generates an ad copy that reflects the content of the surrounding conversation in real time to the registered service. For example, if the utterance transcribed by the speech recognition unit 112 contains topics such as "breakfast" or "morning," it may present the "morning menu" of the registered service; if the utterance contains a conversation such as "I'm hungry," it may present the "5% off coupon" of the registered service; if the utterance contains a conversation about sweets, it may present the "seasonal parfait and cake set" of the registered service; and if there is a fun conversation about family, it may present the "family set" menu of the registered service.

[0097] In steps S33 to S34 of Figure 23, the ad generation unit 115b may generate a prompt to generate ad copy that suggests a matching registered service from the knowledge DB acquired by the RAG function based on the hobbies and interests of the collected surrounding viewer information. For example, if a fast food restaurant is determined to be the target of the advertisement, the ad generation unit 115b may select a registered service for a "Discount Campaign for U20s" from the information of "active in the soccer club" in the collected surrounding viewer information and the information of "group of students returning from club activities" from the execution result of the VLM prompt, and generate a prompt to generate ad copy for a group of male students returning from soccer club activities. In S35, the ad generation unit 115b obtains ad copy such as "Hungry men, gather at Shibuya XX store! Strengthen team bonds! U20 limited, all-you-can-eat set."

[0098] Similar to the modified version of the first embodiment, in S34 of Figure 23, the advertisement generation unit 115b may customize the ending of the advertisement text to a specific form such as "nominalization," "polite style" ("desu," "masu"), "plain style" ("dearu," "da"), or "conversational style" ("dayo," "dane").

[0099] Furthermore, similar to the modified version of the first embodiment, in S34 of Figure 23, the advertisement generation unit 115b may extract specific endings and expressive features from the utterance transcribed into text by the speech recognition unit 112, and generate a prompt to generate an advertisement text with expressive forms such as "dialects" ("dabe", "yaro", "yanaa", "dosu", "yaken") and "unique ending forms" ("ja", "degozaru", "nya", "dayoon") that match the characteristics of the speaker.

[0100] (Effects) As described above, the advertising display device of the second embodiment can display effective advertisements by estimating the attributes and circumstances of viewers walking around or behind the advertising display device using a VLM generation AI model, selecting appropriate advertising targets and services based on the attributes and circumstances, and generating appropriate ad copy that matches the attributes and circumstances in real time.

[0101] <Modifications of the First and Second Embodiments> In the advertising systems 1a and 1b described above, the generation AI model of the generation device 40 or the VLM generation AI model of the VLM processing device 50 is implemented on different devices connected to the advertising display devices 10a and 10b via a network, but it may also be implemented on the advertising display devices 10a and 10b. Furthermore, the RAG function implemented on the advertising display devices 10a and 10b may be implemented on the advertising management devices 20a and 20b, or it may be implemented on the generation device 40 together with the generation AI model.

[0102] Furthermore, each functional unit of the advertising display devices 10a, 10b, advertising management devices 20a, 20b, statistical information management device 30, generation device 40, and VLM processing device 50 may be implemented in a distributed manner on the cloud. Also, each functional unit may be implemented on multiple devices. Furthermore, the same functional unit may be realized by multiple devices or multiple advertising display devices 10a, 10b. The implementation of each functional unit on each device can be freely modified.

[0103] The advertising management devices 20a, 20b, statistical information management device 30, generation device 40, and VLM processing device 50 in embodiments of this disclosure may function as a computer that performs the processing of this disclosure. Figure 30 shows an example of the hardware configuration of the advertising management devices 20a, 20b, statistical information management device 30, generation device 40, and VLM processing device 50 according to embodiments of this disclosure. The advertising management devices 20a, 20b, statistical information management device 30, generation device 40, and VLM processing device 50 described above may be physically configured as a computer device including a control unit 1001, main memory unit 1002, auxiliary memory unit 1003, communication unit 1004, input device 1005, output device 1006, bus 1007, etc.

[0104] In the following explanation, the term "device" can be replaced with "circuit," "device," "unit," etc. The hardware configuration of the advertising management devices 20a, 20b, the statistical information management device 30, the generation device 40, and the VLM processing device 50 may be configured to include one or more of the devices shown in the figure, or it may be configured to omit some of the devices.

[0105] The functions of the advertising management devices 20a and 20b, the statistical information management device 30, the generation device 40, and the VLM processing device 50 are realized by loading predetermined software (programs) onto hardware such as the control unit 1001 and the main memory unit 1002, which allows the control unit 1001 to perform calculations, control communication by the communication unit 1004, and control at least one of data reading and writing in the main memory unit 1002 and the auxiliary memory unit 1003.

[0106] The control unit 1001 controls the entire computer, for example, by running the operating system. The control unit 1001 may be composed of a central processing unit (CPU) that includes interfaces with peripheral devices, control devices, arithmetic units, registers, etc.

[0107] Furthermore, the control unit 1001 reads programs (program code), software modules, data, etc., from at least one of the auxiliary storage unit 1003 and the communication unit 1004 into the main memory unit 1002, and executes various processes accordingly. The program used is one that causes a computer to execute at least a part of the operations described in the above embodiment. For example, each part of the advertising display device 10a, 10b, the generation device 40, and the user terminal 30 may be stored in the main memory unit 1002 and implemented by a control program that operates in the control unit 1001, and other functional blocks may be implemented similarly. The above-mentioned various processes have been described as being executed by one control unit 1001, but they may be executed simultaneously or sequentially by two or more control units 1001. The control unit 1001 may be implemented by one or more chips. The program may also be transmitted from a network via a telecommunications line.

[0108] The main memory unit 1002 is a computer-readable recording medium and may consist of at least one of the following: ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. The main memory unit 1002 may also be called a register, cache, main memory, etc. The main memory unit 1002 can store executable programs (program code), software modules, etc., for implementing the wireless communication method according to the embodiments of this disclosure.

[0109] The auxiliary storage unit 1003 is a computer-readable recording medium and may consist of at least one of the following: an optical disc such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., compact disk, digital multipurpose disk, Blu-ray® disk), an SSD (Solid State Drive), a smart card, flash memory (e.g., card, stick, key drive), a floppy® disk, a magnetic strip, etc. The auxiliary storage unit 1003 may also be called an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, server, or other suitable medium including at least one of the main storage unit 1002 and the auxiliary storage unit 1003.

[0110] The communication unit 1004 is hardware (transceiver / receiver device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as a network device, network controller, network card, communication module, etc.

[0111] The input device 1005 is an input device that accepts input from an external source (e.g., a keyboard, mouse, microphone, switch, button, sensor, etc.). The output device 1006 is an output device that outputs to an external source (e.g., a display, speaker, LED lamp, etc.). The input device 1005 and the output device 1006 may be configured as an integrated unit (e.g., a touch panel).

[0112] Furthermore, each device, such as the control unit 1001 and the main memory unit 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or different buses may be configured for each device.

[0113] Furthermore, the advertising management devices 20a, 20b, the statistical information management device 30, the generation device 40, and the VLM processing device 50 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), and an FPGA (Field Programmable Gate Array), and some or all of each functional block may be realized by such hardware. For example, the control unit 1001 may be implemented using at least one of these hardware components.

[0114] One aspect of this disclosure is useful for an information processing device, an information processing method, and a program that can generate optimal advertisements based on the attributes of people present in an area and the attributes of surrounding advertisement viewers.

[0115] 1a Information processing system (first embodiment) 1b Information processing system (second embodiment) 10a Advertising display device (first embodiment) 10b Advertising display device (second embodiment) 11a Control unit (first embodiment) 11b Control unit (second embodiment) 12 Input unit 13 Output unit 14 GPS receiver 15 Communication unit 16 Main memory unit 17a Auxiliary memory unit (first embodiment) 17b Auxiliary memory unit (second embodiment) 18 Shoulder strap 20a Advertising management device (first embodiment) 20b Advertising management device (second embodiment) 30 Statistical information management device 40 Generation device 50 VLM processing device 60 Network 111a Image recognition unit (first embodiment) 111b VLM instruction unit (second embodiment) 112 Voice recognition unit 113a Viewer estimation unit (first embodiment) 113b Viewer estimation unit (second embodiment) 114a Ad selection unit (first embodiment) 114b Ad selection unit (second embodiment) 115a Ad generation unit (first embodiment) 115b Ad generation unit (second embodiment) 121 Camera 122 Microphone 123 Input device 131 Display unit 132 Speaker 171 Ad DB 172 Knowledge DB

Claims

1. An information processing device comprising: an advertising selection unit that acquires text information by inputting image data obtained by photographing people present around the information processing device and a VLM (Vision Language Model) prompt for acquiring pedestrian information in the image data into a VLM generation AI model, and selects an advertising target based on the current location of the information processing device and the text information; and an advertising generation unit that determines information to be used for advertising based on the advertising target and the text information, and generates a prompt for generating advertising text based on the text information and the information to be used for advertising.

2. The information processing apparatus according to claim 1, wherein the VLM prompt includes an instruction for obtaining the age, gender, situation, movement, or behavior of the pedestrian in the image data, and the text information includes the estimated age, gender, situation, movement, or behavior of a person present around the information processing apparatus.

3. The information processing apparatus according to claim 1, wherein the advertisement selection unit selects as the advertisement target the advertisement in which the current location and the location of the store in the advertisement information are within a certain range and the content contained in the text information is most suitable for the target attributes of the advertisement information, and the advertisement generation unit determines as the information to be used for the advertisement the service in which the information to be used for the advertisement is a service provided at the store in the advertisement target and the content contained in the text information is most suitable for the target attributes of the service provided at the store in the advertisement target.

4. The information processing apparatus according to claim 3, wherein the advertisement generation unit generates a prompt for generating an advertisement using the service and information related to the service, and obtains the advertisement by inputting the generated prompt into a generation AI model.

5. An information processing method comprising: an information processing device acquiring text information by inputting image data obtained by photographing people present around the information processing device and a VLM (Vision Language Model) prompt for acquiring pedestrian information in the image data into a VLM generation AI model; selecting an advertising target based on the current location of the information processing device and the text information; determining information to be used in the advertisement based on the advertising target and the text information; and generating a prompt to generate an advertisement text based on the text information and the information to be used in the advertisement.

6. A program to cause an information processing device to perform the following processes: acquire text information by inputting image data obtained by photographing people present around the information processing device and a VLM (Vision Language Model) prompt for acquiring pedestrian information in the image data into a VLM generation AI model; select an advertising target based on the current location of the information processing device and the text information; determine information to be used in the advertisement based on the advertising target and the text information; and generate a prompt to generate an advertisement text based on the text information and the information to be used in the advertisement.