system
The taxi-based system uses facial and voice analysis to deliver personalized advertisements with QR codes, addressing the lack of individual targeting in conventional systems and enhancing advertising effectiveness.
Patent Information
- Application Number
- JP2024128428
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-02
- Publication Date
- 2026-02-16
AI Technical Summary
Conventional taxi advertising systems fail to target advertisements to individual passengers based on their attributes, lacking real-time analysis and interactive elements, resulting in reduced advertising effectiveness.
A system installed in taxis that captures passenger facial images and voices, analyzes gender, age, and ride purpose, generates tailored advertisements with QR codes for coupons, and displays them on a display, enhancing interaction.
Provides personalized advertisements that improve passenger experience and maximize advertising effectiveness by targeting individual needs.
Smart Images

Figure 2026025619000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional taxi advertising systems display the same advertisements to all passengers, making it difficult to target advertisements to meet individual needs. Furthermore, technology for analyzing passenger attributes in real time and generating and displaying optimal advertisements based on that information was underdeveloped. Furthermore, interactive elements to enhance the effectiveness of advertisements were lacking, resulting in reduced advertising effectiveness. The present invention aims to solve these problems and provide a system that delivers optimal advertisements to each passenger. [Means for solving the problem]
[0005] The present invention is a system using a terminal installed in a taxi. Specifically, it includes an imaging means for capturing facial images of passengers and an audio collection means for capturing passengers' voices. It also includes an analysis means for analyzing data obtained from these means and determining the passenger's gender, age, and purpose of the ride. It also includes an advertisement generation means for generating optimal advertisements based on the passenger information determined by the analysis means. The generated advertisements are displayed via a display means, and interactive elements are added by a QR code generation means for providing coupons and links related to the advertisements. This makes it possible to provide advertisements that are tailored to each passenger and maximize the effectiveness of the advertisements.
[0006] "Imaging means" refers to a camera or other imaging device for capturing facial images of passengers.
[0007] "Audio collection means" means a microphone or other audio collection device for recording passenger speech.
[0008] The "analysis means" is a system or algorithm for processing data obtained from the imaging means and audio collection means to determine the gender, age, and purpose of the passenger.
[0009] The "advertising generation means" is a system or software for creating optimal advertisements based on passenger information determined by the analysis means.
[0010] The "display means" refers to a display or other display device for displaying the advertisements generated by the advertisement generating means to passengers.
[0011] A "QR code generator" is a system or software for creating a QR code containing a coupon or link related to an advertisement.
[0012] A "face image" is image data of a passenger's face acquired by an imaging means.
[0013] "Voice data" refers to passenger voice information recorded by a voice collection means.
[0014] "Gender" is gender information such as male or female of the passenger determined by the analysis means.
[0015] "Age" refers to the generation or specific age of the passenger as determined by the analytical means.
[0016] The "purpose of the ride" is the purpose of the passenger's taxi use determined by the analysis means, and is a specific purpose such as a business meeting or sightseeing.
[0017] "Advertising materials" are content such as images, videos, and text that are used by the advertisement generation means to create advertisements.
[0018] A "coupon" is special offer information provided by a QR code generating means, and is provided to passengers in the form of discounts on products and services. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] System Configuration
[0041] The system of the present invention includes the following major components:
[0042] 1. Terminal installed in the taxi: Equipped with imaging means, audio collection means, and display means.
[0043] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means.
[0044] 3. Network: Data communication between the terminal and the server.
[0045] System Operation
[0046] Collecting passenger information
[0047] When a passenger gets into a taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image. At the same time, the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time as soon as the passenger gets into the taxi.
[0048] Data analysis
[0049] The server runs a facial recognition algorithm on the received image data to estimate the passenger's gender and age, and applies a voice recognition system to the audio data to analyze the passenger's purpose. The results of these analyses are stored in a database.
[0050] Ad Generation
[0051] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses ad generation AI to create ads for passengers. The ads contain text information and graphics tailored to the passenger's interests and goals. QR codes containing special coupons and related links are also generated and embedded in the advertising content.
[0052] Advertisement display and coupon offer
[0053] The advertisement data sent from the server is returned to the terminal and provided to passengers through a display. The advertisement contains a QR code, which passengers can scan to receive special coupons and link information.
[0054] Specific program processing
[0055] Program processing explanation
[0056] Scenario: A man in his 30s gets into a taxi for a business meeting.
[0057] 1. Information collection during the ride:
[0058] The device "uses a camera to capture passengers' faces and a microphone to record the purpose of the ride."
[0059] Device: "Sends the captured image and audio data to the server"
[0060] 2. Data Analysis:
[0061] Server: "From the received image data, a facial recognition algorithm estimates the gender as male and the age as being in their 30s."
[0062] Server: "The voice recognition system converts the voice data into text and extracts the keyword 'business negotiations'."
[0063] 3. Ad generation:
[0064] Server: "Based on the analysis results, select business-related advertisements from the advertising materials."
[0065] Server: Generates ads that combine business-related videos and text.
[0066] Server: Embed a QR code containing a special coupon or link in the ad.
[0067] 4. Advertising and Coupon Offers:
[0068] Server: "Sends the generated advertising data to the device"
[0069] Device: "The advertisement will be displayed on the tablet screen, and the QR code will be placed in an easily visible position."
[0070] User: "Scan the QR code with your smartphone and receive a special coupon."
[0071] This will improve passengers' in-cab experience by providing tailored advertising, while also enabling advertisers to effectively target their ads and maximize their advertising effectiveness.
[0072] The processing flow will be explained below.
[0073] Program processing steps
[0074] Specific scenario: A man in his 30s gets into a taxi for a business meeting.
[0075] Step 1:
[0076] The device: "The moment a passenger gets into the taxi, the camera activates and captures the passenger's face."
[0077] The device photographs the passenger's face from multiple angles to obtain clear image data.
[0078] Step 2:
[0079] "At the same time, the microphone collects passenger conversations and records voice data in real time."
[0080] The device records the conversation audio in high quality and performs pre-processing to remove noise.
[0081] Step 3:
[0082] Terminal: "Sends collected image data and audio data to the server."
[0083] The terminal encrypts the image and audio data and transmits it to the server using a secure communication protocol.
[0084] Step 4:
[0085] Server: Based on the received image data, it runs a facial recognition algorithm to estimate the passenger's gender and age.
[0086] The server uses a deep learning model to classify the gender as male and the age as in their 30s from the image.
[0087] Step 5:
[0088] Server: "Apply a voice recognition system to the voice data and analyze the purpose of the ride."
[0089] The server uses a speech recognition API to convert the speech into text and extract the keyword "business negotiation."
[0090] Step 6:
[0091] Server: "Save the analysis results in the database"
[0092] The server stores data such as passenger ID, estimated gender, age, and purpose of ride in a database.
[0093] Step 7:
[0094] Server: "Based on the analysis results, we will create optimal ads using ad generation AI."
[0095] The server selects business-related videos and text from advertising materials aimed at businessmen in their 30s.
[0096] Step 8:
[0097] Server: Generate a QR code containing a special coupon or link for the ad and embed it in the ad content.
[0098] The server uses a QR code generation algorithm and integrates it into the advertisement.
[0099] Step 9:
[0100] Server: "Sends the generated advertising data to the taxi terminal."
[0101] The server re-encrypts the advertisement data and transmits it to the terminal.
[0102] Step 10:
[0103] Terminal: "Decodes the received advertising data and displays it on the tablet display."
[0104] The terminal places advertisements in a location that is easy for passengers to see.
[0105] Step 11:
[0106] The terminal "displays the QR code included in the advertisement in front of the passenger."
[0107] Place your device in the center of the screen so that the QR code is easy to scan.
[0108] Step 12:
[0109] User: "Scan the QR code with your smartphone to get a special coupon."
[0110] Users can scan the QR code using their smartphone's camera app to receive the coupon information.
[0111] The above are the specific processing steps of the program of the present invention. In this way, the processing at each step works together to provide passengers with optimal targeted advertising and interactive experiences.
[0112] Example 1
[0113] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0114] Conventional taxi ride systems lacked a method for providing appropriate and effective advertisements to passengers. Furthermore, technology for generating and displaying advertisements tailored to passengers' interests in real time was underdeveloped. As a result, the effectiveness of advertisements was low, and passenger convenience was difficult to improve. The present invention aims to solve these problems and provide more effective and personalized advertisements to passengers.
[0115] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0116] In this invention, the server includes an imaging means for capturing facial images of passengers, an audio collection means for capturing passenger voices, an analysis means for analyzing data obtained from the imaging means and the audio collection means to determine the passenger's gender, age, and purpose of the ride, an advertisement generation means for generating an optimal advertisement based on the passenger information determined by the analysis means, a display means for displaying the advertisement generated by the advertisement generation means, a QR code generation means for providing benefits and links related to the advertisement, a communication means for encrypting data acquired by the imaging means and the audio collection means and transmitting the encrypted data to the server, an image recognition means for executing the facial recognition algorithm to estimate the passenger's gender and age, a speech conversion means for converting speech data into text using the speech recognition system and extracting important keywords, an advertising material selection means for creating optimal advertising content from advertising materials using a generative AI model, and a QR code generation means for generating a QR code containing a benefit coupon or link destination information. This allows passengers to receive advertisements tailored to their needs, improving their in-taxi experience.
[0117] "Imaging means" is a hardware or software component for capturing facial images of passengers.
[0118] "Voice collection means" is a hardware or software component for collecting passenger voices.
[0119] "Analysis means" is a general term for hardware or software for analyzing data obtained from the imaging means and audio collection means and determining the gender, age, and purpose of the passenger's ride.
[0120] The "advertising generation means" is a hardware or software component for generating optimal advertisements based on passenger information determined by the analysis means.
[0121] The "display means" refers to a display device or software for visually presenting the advertisements generated by the advertisement generating means to passengers.
[0122] A "QR code generator" is a hardware or software component that generates a QR code to provide a special offer or link related to an advertisement.
[0123] The "communication means" is a hardware or software function for encrypting data acquired by the imaging means and the audio collecting means and transmitting the data to the server.
[0124] "Image Recognition Means" means a hardware or software component that runs a facial recognition algorithm and estimates the gender and age of a passenger.
[0125] "Speech conversion means" refers to hardware or software functionality that converts voice data into text using a voice recognition system and extracts important keywords.
[0126] "Advertising material selection means" means a hardware or software function that uses a generative AI model to select and create optimal advertising content from advertising materials.
[0127] The system of the present invention mainly comprises a terminal installed in a taxi, a server that performs data analysis and advertisement generation, and a network that communicates between these components. The operation of the system proceeds as follows.
[0128] System Configuration
[0129] 1. Terminal
[0130] Imaging means: Includes a high-resolution camera for capturing facial images of passengers. The camera is automatically activated when a passenger gets into the taxi and captures facial images. As a specific example, a Sony IMX586 camera is used.
[0131] Audio collection means: Includes a microphone for collecting passengers' voices. The microphone is activated when passengers board the vehicle and records their conversations. As a specific example, a Rode NT-USB microphone is used.
[0132] Display means: includes a high-resolution touchscreen display for displaying advertising content. As a specific example, a 16-inch high-resolution touchscreen display is used.
[0133] Communication means: Includes a communication module for encrypting data acquired by the imaging means and audio collection means and transmitting it to a server. As a specific example, a 4G LTE module is used.
[0134] 2. Server
[0135] Analysis means: includes a software component for analyzing data obtained from the imaging means and audio collection means and determining the gender, age, and purpose of the passenger.
[0136] Image Recognition: Implements a facial recognition algorithm to estimate gender and age. For example, the OpenCV library is used.
[0137] Speech conversion method: A speech recognition system is used to convert the audio data into text and extract important keywords. Specific examples include DeepSpeech and Amazon Transcribe.
[0138] Ad generation methods:
[0139] Advertising material selection means: Based on passenger information determined by the analysis means, the optimal advertising material is selected and advertising content is created using a generative AI model. As a specific example, GPT-4 is used.
[0140] QR code generation method: Generates a QR code containing special coupons and link information. As a concrete example, the Libqrencode library is used.
[0141] 3. Network
[0142] This includes a network for fast and secure data communication between the terminals installed in the taxi and the server. For example, 4G / 5G networks are used.
[0143] Specific examples
[0144] The following is a specific example of the system's operation.
[0145] Scenario: A man in his 30s gets into a taxi for a business meeting.
[0146] 1. Terminal: When a passenger gets into a taxi, the terminal automatically activates the camera to capture the passenger's face image, and simultaneously activates the microphone to record the passenger's conversation.
[0147] 2. Terminal: The acquired facial image and audio data are encrypted and sent to the server.
[0148] 3. Server: Analyzes the received facial image and estimates the passenger's gender as "male" and age as "30s." The voice data is converted into text using a voice recognition system, and the keyword "business negotiation" is extracted.
[0149] 4. Server: Based on the analysis results, the generative AI model generates optimal business-related advertising content. For example, the prompt sentence should include keywords such as "men in their 30s" and "business negotiations" to generate business-related advertising content.
[0150] 5. Server: Embeds a QR code containing a special coupon or link information into the generated advertising content and sends it to the device.
[0151] 6. Terminal: Display the advertisement on the display inside the taxi and place the QR code in a location that is easy to see.
[0152] 7. User: Scans the QR code displayed on the screen with a smartphone to receive special coupons and link information.
[0153] This system will enable passengers to receive advertisements tailored to their needs, improving their in-cab experience, while also enabling advertisers to target their ads more effectively, maximizing their advertising impact.
[0154] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0155] Program processing
[0156] Step 1:
[0157] Device: When a passenger gets into the taxi, the device automatically activates its camera to capture the passenger's facial image, and simultaneously activates its microphone to begin recording the passenger's conversation.
[0158] Input: Passenger riding behavior
[0159] Data processing: Images are taken with a camera and image data (JPEG format) is generated. Audio is collected with a microphone and audio data (WAV format) is generated.
[0160] Output: Facial image data and audio data
[0161] Step 2:
[0162] Terminal: The acquired facial image data and voice data are end-to-end encrypted using SSL and sent to the server.
[0163] Input: Facial image data and audio data
[0164] Data processing: SSL encryption is applied and data communication is converted into secure data.
[0165] Output: Encrypted facial image data and audio data
[0166] Step 3:
[0167] Server: Based on the received facial image data, a facial recognition algorithm (OpenCV library) is executed to estimate the passenger's gender and age. The results are stored in a database.
[0168] Input: Encrypted facial image data
[0169] Data processing: The data is decrypted and facial recognition algorithms are used to generate gender and age data.
[0170] Output: Gender and age information
[0171] Step 4:
[0172] Server: The voice data is converted into text using a speech recognition system (DeepSpeech or Amazon Transcribe), and the purpose of the ride is analyzed using a natural language processing (NLP) algorithm. The analysis results are stored in a database.
[0173] Input: Encrypted audio data
[0174] Data processing: The data is decoded and generated into text data using a speech recognition system, after which keywords are extracted using NLP algorithms.
[0175] Output: Keywords for trip purpose
[0176] Step 5:
[0177] Server: Based on the analysis results (gender, age, purpose of ride), a generative AI model (GPT-4) is used to generate the optimal advertisement. Examples of prompt sentences include keywords such as "male in his 30s" and "business negotiation."
[0178] Input: Gender, age, purpose of ride keywords
[0179] Data processing: Input prompt text into the generative AI model to generate advertising content.
[0180] Output: Ad content data
[0181] Step 6:
[0182] Server: Generates a QR code containing a special coupon or link information in the generated advertising content and embeds it in the advertising content.
[0183] Input: Ad content data
[0184] Data processing: Use a QR code generation library to generate a QR code containing special coupon information and link information. This can then be incorporated into advertising content.
[0185] Output: QR code embedded advertising content data
[0186] Step 7:
[0187] Server: Sends advertising content data with an embedded QR code to the terminal.
[0188] Input: QR code embedded advertising content data
[0189] Data processing: The advertising content data is encrypted and prepared for transmission.
[0190] Output: Encrypted advertising content data
[0191] Step 8:
[0192] Terminal: Display the received advertising content data on the display and place the QR code in an easily visible location.
[0193] Input: Encrypted ad content data
[0194] Data processing: Decrypting the data and optimizing it for display.
[0195] Output: Advertising content and QR code displayed on the display
[0196] Step 9:
[0197] User: Scans the QR code displayed on the screen with a smartphone to receive special coupons and link information.
[0198] Input: QR code displayed on the screen
[0199] Data processing: Scan the QR code with your smartphone to obtain special coupons and link information.
[0200] Output: Bonus coupons and link information saved on your smartphone
[0201] This ensures smooth operation of the entire system, improves passengers' in-cab experience by allowing them to receive advertisements tailored to their needs, and enables advertisers to effectively target their advertisements, maximizing their advertising effectiveness.
[0202] (Application example 1)
[0203] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0204] Conventional in-taxi advertising systems have the challenge of making it difficult to provide advertisements tailored to passengers' interests and needs in real time. Furthermore, because advertisements are displayed on fixed monitors, there is no guarantee that passengers will see them. In such situations, advertising effectiveness is not fully realized, resulting in an unsatisfactory experience for both advertisers and passengers.
[0205] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0206] In this invention, the server includes an imaging means for capturing facial images of passengers, an audio collecting means for collecting passenger voices, an analysis means for analyzing data obtained from the imaging means and the audio collecting means and determining the gender, age, and purpose of the passenger's ride, an advertisement generating means for generating an optimal advertisement based on the passenger information determined by the analysis means, a display means for displaying the advertisement generated by the advertisement generating means, a QR code generating means for providing coupons and links related to the advertisement, and a display means for identifying that the display means is a smart glasses display. This allows advertisements tailored to individual passenger needs to be provided in real time and the advertisements to be displayed in a location that is easily visible through the smart glasses.
[0207] "Imaging means" refers to a function or device used to capture facial images of passengers.
[0208] "Voice collection means" refers to the functionality or device used to collect passenger voices.
[0209] The "analysis means" is a function or device that analyzes the data obtained from the imaging means and the audio collection means and determines the gender, age, and purpose of the passenger.
[0210] The "advertisement generation means" is a function or device that generates optimal advertisements based on passenger information determined by the analysis means.
[0211] The "display means" refers to a function or device that displays the advertisement generated by the advertisement generation means, and includes the display of the smart glasses.
[0212] The term "QR code generating means" refers to a function or device that generates a QR code for providing a coupon or link related to an advertisement.
[0213] System Configuration
[0214] An embodiment of the present invention includes the following major components:
[0215] 1. Terminal: Equipped with an imaging means (camera), an audio collection means (microphone), and a display means (display) for displaying advertisements, all of which are installed in the smart glasses.
[0216] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means for providing coupons and links.
[0217] 3. Network: Data communication between the terminal and the server.
[0218] System Operation
[0219] Collecting passenger information
[0220] When a user puts on the smart glasses, the camera automatically captures facial images and the microphone starts recording the user's conversation, and these data are sent to the server in real time.
[0221] Data analysis
[0222] The server analyzes the received image and voice data, using a facial recognition algorithm to estimate the user's gender and age, and a voice recognition system to analyze the user's intent based on the voice data.
[0223] Ad Generation
[0224] Based on the analysis results, the server generates the most suitable advertisement, which includes text information and graphics tailored to the user's interests and goals, as well as a QR code containing special coupons and related links.
[0225] Advertisement display and coupon offer
[0226] The generated advertising data is displayed on the smart glasses display, and users can receive special coupons and link information by scanning the QR code.
[0227] Specific program processing explanation
[0228] The server uses a facial recognition library (e.g., OpenCV) and a voice recognition library (e.g., SpeechRecognition) to analyze the user's facial image and voice data. This allows the server to extract the user's gender, age, and purpose, and then uses ad generation AI to create optimal ads. It also generates QR codes containing coupons and link information, and incorporates them into the ad content.
[0229] Specific examples
[0230] Scenario: A woman in her twenties puts on smart glasses while out shopping.
[0231] 1. The camera in the smart glasses captures the woman's face, and the microphone collects the audio of her saying "shopping."
[0232] 2. The server uses facial recognition to determine that the user is a woman in her 20s, and uses voice analysis to detect the keyword "shopping."
[0233] 3. Generate ads for the latest fashion and cosmetics for women and embed QR codes containing special coupons and shop links into the ads.
[0234] 4. Advertisements will be displayed on the smart glasses display, and users can scan the QR code to receive reward coupons.
[0235] Prompt Sentence Examples
[0236] Give your generative AI model the following inputs:
[0237] User information: Female in her 20s
[0238] Keywords: Shopping
[0239] Generated advertising content: Latest fashion, cosmetics advertising
[0240] QR code content: Bonus coupon, shop link QR code
[0241] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0242] Step 1:
[0243] The device detects that the user has put on the smart glasses. At this time, the camera is activated and an image of the user's face is captured. The captured image is saved in a data format (e.g., JPEG format) and sent to the server.
[0244] Input: User's face image data
[0245] Output: JPEG format face image file
[0246] Step 2:
[0247] The device collects the user's voice using the microphone in the smart glasses. The collected voice data is saved as an audio file (e.g., WAV format) and sent to the server.
[0248] Input: User's voice data
[0249] Output: WAV format audio file
[0250] Step 3:
[0251] The server receives the facial image data and uses a facial recognition algorithm (e.g., OpenCV) to estimate the user's gender and age. During this process, image processing operations are performed to obtain the estimated gender and age.
[0252] Input: JPEG format face image file
[0253] Output: User's gender and age (e.g., female in her 20s)
[0254] Step 4:
[0255] The server receives the voice data and converts the voice data into text data using a voice recognition system (e.g., SpeechRecognition). Keywords (e.g., "shopping") are extracted from the converted text data.
[0256] Input: WAV format audio file
[0257] Output: Text data and extracted keywords
[0258] Step 5:
[0259] The server uses ad generation AI to generate optimal ads based on user information (gender, age, keywords) obtained through facial and voice recognition. During this process, ad materials are selected and combined.
[0260] Input: User information (gender, age, keywords)
[0261] Output: Generated advertising data (text information, graphics, QR code)
[0262] Step 6:
[0263] The server transmits the generated advertisement data to the terminal, and the terminal displays the received advertisement data on the display of the smart glasses.
[0264] Input: Generated ad data
[0265] Output: Advertisement displayed on the smart glasses display
[0266] Step 7:
[0267] Users scan the QR code displayed on the smart glasses display with their smartphone to receive reward coupons and shop links.
[0268] Input: QR code displayed on the screen
[0269] Output: Special coupons and shop links acquired by the user's smartphone
[0270] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0271] System Configuration
[0272] The present invention includes the following major components:
[0273] 1. Terminal installed in the taxi: equipped with imaging means, audio collection means, display means, and emotion engine.
[0274] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means.
[0275] 3. Network: Data communication between the terminal and the server.
[0276] System Operation
[0277] Collecting passenger information
[0278] When a passenger gets into a taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image. At the same time, the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time as soon as the passenger gets into the taxi.
[0279] Data analysis
[0280] The server runs a facial recognition algorithm based on the received image data to estimate the passenger's gender and age. It also applies a voice recognition system to the audio data to analyze the passenger's purpose. It also uses an emotion engine to determine the passenger's emotional state based on their facial expressions and voice. The results of these analyses are stored in a database.
[0281] Ad Generation
[0282] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses ad generation AI to create ads for passengers. The ads include text information and graphics tailored to the passenger's gender, age, purpose of the ride, and emotional state. QR codes containing special coupons and related links are also generated and embedded in the advertising content.
[0283] Advertisement display and coupon offer
[0284] The advertising data sent from the server is returned to the terminal and provided to passengers via a display. The advertisement contains a QR code, which passengers can scan to receive special coupons or link information. In addition, real-time emotional feedback from passengers regarding the displayed advertisement is collected again, and the content of the advertisement is dynamically adjusted to provide a more effective advertising experience.
[0285] Specific program processing
[0286] Program processing explanation
[0287] Scenario: A man in his 30s gets into a taxi for a business meeting and feels a little nervous.
[0288] 1. Information collection during the ride:
[0289] The device "uses a camera to capture passengers' faces and a microphone to record the purpose of the ride."
[0290] Device: "Sends the captured image and audio data to the server"
[0291] 2. Data Analysis:
[0292] Server: "From the received image data, a facial recognition algorithm determines the gender as male and the age as in their 30s."
[0293] Server: "The voice recognition system converts the voice data into text and extracts the keyword 'business negotiations'."
[0294] Server: "Using an emotion engine to detect tension levels from passengers' facial expressions and voices."
[0295] 3. Ad generation:
[0296] Server: "Based on the analysis results, select the most suitable business-related advertising materials."
[0297] The server "generates ads containing text and images that soften products and services and incorporates content that reduces tension."
[0298] Server: Generates QR codes containing special coupons and links and embeds them in advertising content.
[0299] 4. Advertising and Coupon Offers:
[0300] Server: "Sends the generated advertising data to the device"
[0301] Device: "The advertisement will be displayed on the tablet screen, and the QR code will be placed in an easily visible position."
[0302] User: "Scan the QR code with your smartphone to get a special coupon."
[0303] The device "recaptures the passenger's facial expression in response to the displayed advertisement and sends emotional feedback to the server."
[0304] This allows passengers to see ads that reflect their real-time emotional state, providing them with an optimal advertising experience, while enabling advertisers to achieve more effective targeting and interactive advertising methods.
[0305] The processing flow will be explained below.
[0306] Program processing steps
[0307] Specific scenario: A man in his 30s is getting into a taxi for a business meeting and is feeling a bit nervous.
[0308] Step 1:
[0309] The moment a passenger gets into the taxi, the camera activates and captures the passenger's facial image.
[0310] The device photographs the passenger's face from multiple angles to obtain clear image data.
[0311] Step 2:
[0312] "At the same time, the microphone collects passenger conversations and records voice data in real time."
[0313] The device records passengers' conversations in high quality and performs pre-processing to remove noise.
[0314] Step 3:
[0315] The device "encodes the collected image and audio data and transmits it to a server using a secure communication protocol."
[0316] The terminal encrypts the data to ensure secure communication.
[0317] Step 4:
[0318] Server: Based on the received image data, it runs a facial recognition algorithm and determines the passenger's gender as male and age as in their 30s.
[0319] The server uses a deep learning model to estimate the gender and age of passengers from facial images.
[0320] Step 5:
[0321] Server: "Apply a voice recognition system to the voice data and analyze the purpose of the ride."
[0322] The server uses a speech recognition API to convert the speech into text data and extract the keyword "business negotiation."
[0323] Step 6:
[0324] Server: "Uses an emotion engine to determine the passenger's emotional state from their facial expressions and voice."
[0325] The server performs facial expression analysis and voice tone analysis to detect the passenger's state of tension.
[0326] Step 7:
[0327] Server: "Save the analysis results in the database"
[0328] The server stores passengers' gender, age, purpose of the ride, and emotional state data in a database.
[0329] Step 8:
[0330] Server: "Based on the analysis results, we will use ad generation AI to create ads for passengers."
[0331] The server selects advertising materials that are business-related and de-escalating, and generates customized advertisements.
[0332] Step 9:
[0333] Server: Generates QR codes containing special coupons and links and embeds them in advertising content.
[0334] The server uses a QR code generation algorithm and binds it to the advertisement.
[0335] Step 10:
[0336] Server: "Sends the generated advertising data to the taxi terminal."
[0337] The server encodes the advertisement data and transmits it to the terminal.
[0338] Step 11:
[0339] Terminal: "Decodes the received advertising data and displays it on the tablet display."
[0340] The terminal displays advertising content in real time and places the advertisements in a location that is easy for passengers to see.
[0341] Step 12:
[0342] The terminal "displays the QR code included in the advertisement in front of the passenger."
[0343] Place your device in the center of the screen so that the QR code is easy to scan.
[0344] Step 13:
[0345] User: "Scan the QR code with your smartphone to get a special coupon."
[0346] Users scan the QR code with their smartphone's camera app to obtain coupon and link information.
[0347] Step 14:
[0348] The device "recaptures the passenger's facial expressions in response to the displayed advertisement and obtains emotional feedback."
[0349] The device then takes another photograph of the passenger's face and analyzes their emotional state.
[0350] Step 15:
[0351] The device "sends emotional feedback to the server and readjusts the advertising content."
[0352] The terminal transmits the obtained emotional feedback to the server and dynamically updates the advertisement content.
[0353] This allows passengers to see ads that are tailored to their real-time emotional state, providing a more satisfying advertising experience for passengers and enabling advertisers to deliver more targeted and effective ads.
[0354] Example 2
[0355] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0356] Conventional in-taxi advertising systems do not optimize advertisements based on the passenger's gender, age, and purpose of the ride, nor do they optimize advertisements based on the passenger's current emotional state. This makes it difficult to effectively deliver advertisements that are most appropriate for each passenger. It is also difficult to grasp the effectiveness of advertisements in real time and dynamically adjust them. To solve these problems, the present invention aims to provide optimal advertisements based on detailed passenger information, including the passenger's emotional state, and to provide feedback on the effectiveness of advertisements in real time.
[0357] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0358] In this invention, the server includes an analysis means for analyzing the received data and determining the gender, age, purpose of riding, and emotional state of the passenger, an advertisement generation means for generating an optimal advertisement based on the passenger information and emotional state determined by the analysis means, and a feedback collection means for acquiring passenger emotional feedback on the displayed advertisement and performing additional analysis. This makes it possible to provide optimal advertisements according to the emotional state of each passenger and provide feedback on the effectiveness of the advertisements in real time.
[0359] The "imaging means" is a device for capturing facial images of passengers.
[0360] "Voice collection means" refers to equipment for collecting passenger voices.
[0361] The "communication means" is a function for transmitting data obtained from the imaging means and the sound collecting means to a server in real time.
[0362] "Analysis means" refers to a device or program that analyzes the received data and determines the passenger's gender, age, purpose of the ride, and emotional state.
[0363] The "advertisement generation means" is a device or program for generating an optimal advertisement based on the passenger information and emotional state determined by the analysis means.
[0364] The "display means" is a device for displaying the advertisement generated by the advertisement generation means.
[0365] The "QR code generating means" is a function for generating a QR code for providing a coupon or link related to the advertisement.
[0366] The "feedback collection means" is a function for obtaining passengers' emotional feedback regarding the displayed advertisements and for performing additional analysis.
[0367] The "image analysis means" is a device or program for detecting the age and gender of passengers using image data acquired from the imaging means.
[0368] The "voice analysis means" is a device or program for identifying the purpose of riding and the emotional state of the passenger using the voice data obtained from the voice collection means.
[0369] The "generative AI model" is an artificial intelligence model that selects advertising materials and generates advertising content based on the passenger's gender, age, purpose of riding, and emotional state.
[0370] System Configuration
[0371] The present invention includes the following major components:
[0372] 1. Terminal installed in the taxi: equipped with imaging means, audio collection means, display means, and emotion engine.
[0373] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means.
[0374] 3. Network: Data communication between the terminal and the server.
[0375] Hardware and software used
[0376] Hardware:
[0377] Device imaging means (camera): High-resolution camera
[0378] Audio collection means (microphone): High-sensitivity microphone
[0379] Display means (display): Tablet display
[0380] software:
[0381] Face Recognition Algorithm: Image Processing Library
[0382] Speech Recognition System: Speech Recognition API
[0383] Emotion Engine: Emotion Analysis Tool
[0384] Ad Generation AI: Generative AI Model
[0385] System Operation
[0386] Collecting passenger information
[0387] When a passenger gets into a taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image. At the same time, the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time as soon as the passenger gets into the taxi.
[0388] Data analysis
[0389] The server runs a facial recognition algorithm based on the received image data to estimate the passenger's gender and age. It also applies a voice recognition system to the audio data to analyze the passenger's purpose. It also uses an emotion engine to determine the passenger's emotional state based on their facial expressions and voice. The results of these analyses are stored in a database.
[0390] Ad Generation
[0391] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses a generative AI model to create advertisements for passengers. The advertisements include text information and graphics tailored to the passenger's gender, age, purpose of the ride, and emotional state. QR codes containing special coupons and related links are also generated and embedded in the advertisement content.
[0392] Advertisement display and coupon offer
[0393] The advertising data sent from the server is returned to the terminal and provided to passengers via a display. The advertisement contains a QR code, which passengers can scan to receive special coupons or link information. In addition, real-time emotional feedback from passengers regarding the displayed advertisement is collected again, and the content of the advertisement is dynamically adjusted to provide a more effective advertising experience.
[0394] Specific examples
[0395] Scenario: A man in his 30s gets into a taxi for a business meeting and feels a little nervous.
[0396] 1. Information collection during the ride:
[0397] The device "uses a camera to capture passengers' faces and a microphone to record the purpose of the ride."
[0398] Device: "Sends the captured image and audio data to the server"
[0399] 2. Data Analysis:
[0400] Server: "From the received image data, a facial recognition algorithm determines the gender as male and the age as in their 30s."
[0401] Server: "The voice recognition system converts the voice data into text and extracts the keyword 'business negotiations'."
[0402] Server: "Using an emotion engine to detect tension levels from passengers' facial expressions and voices."
[0403] 3. Ad generation:
[0404] Server: "Based on the analysis results, select the most suitable business-related advertising materials."
[0405] The server "generates ads that include text and images that soften the product's features and characteristics, incorporating tension-reducing content."
[0406] Server: Generates QR codes containing special coupons and links and embeds them in advertising content.
[0407] 4. Advertising and Coupon Offers:
[0408] Server: "Sends the generated advertising data to the device"
[0409] Terminal: "Display advertisements on the display and place QR codes in an easily visible location."
[0410] User: "Scan the QR code with your smartphone to get a special coupon."
[0411] The device "recaptures the passenger's facial expression in response to the displayed advertisement and sends emotional feedback to the server."
[0412] Examples of prompts:
[0413] Prompt statement:
[0414] Please generate an ad based on the following criteria:
[0415] Gender: Male
[0416] Age: 30s
[0417] Purpose of the ride: Business negotiations
[0418] Emotional state: Tension
[0419] Ads should include text and images with tension-busting content, as well as a QR code with a special coupon or relevant link."
[0420] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0421] Step 1:
[0422] Gathering information when riding
[0423] Specific operation: The device automatically activates the camera (photography means) and microphone (audio collection means). The camera captures the passenger's facial image, and the microphone records the passenger's conversation.
[0424] Input: Passenger face image, passenger voice data
[0425] Data processing / calculation: capturing facial images, collecting voice data
[0426] Output: Acquired facial image data and audio data
[0427] Step 2:
[0428] Sending data
[0429] Specific operation: The device sends the collected facial image data and voice data to the server in real time. Data communication is performed via the network.
[0430] Input: Facial image data, audio data
[0431] Data processing / calculation: Data encryption and transmission
[0432] Output: Facial image data and audio data sent to the server
[0433] Step 3:
[0434] Data analysis
[0435] What it does: The server analyzes the data it receives. It uses a facial recognition algorithm to estimate gender and age, a voice recognition system to convert the audio data into text, and an emotion engine to determine the passenger's emotional state.
[0436] Input: Received facial image data and audio data
[0437] Data processing / calculation: facial recognition, voice recognition, emotion analysis
[0438] Output: Gender, age, purpose of ride, emotional state
[0439] Step 4:
[0440] Ad Generation
[0441] What it does: The server generates the best ads based on the analysis results, uses generative AI models to create text information and graphics, and creates QR codes with special offers and links.
[0442] Input: Gender, age, purpose of ride, emotional state
[0443] Data processing / calculation: Generating advertising content, generating QR codes
[0444] Output: Advertisement data (text, images, QR code, etc.)
[0445] Step 5:
[0446] Advertisement display and coupon offer
[0447] Specific operation: The advertising data sent from the server is returned to the device and displayed on the screen. The user can then scan the QR code with their smartphone to receive a special coupon.
[0448] Input: Ad data
[0449] Data processing / calculation: Display of advertising data
[0450] Output: Displayed ad content, user-acquired reward coupon
[0451] Step 6:
[0452] Gathering feedback
[0453] How it works: The device recaptures passenger responses after the ad is displayed and sends the data to the server, which performs additional analysis and dynamically adjusts the ad content.
[0454] Input: Recollected facial image data, emotion feedback
[0455] Data processing / calculation: Feedback analysis, dynamic adjustment of advertising content
[0456] Output: Adjusted advertising data
[0457] (Application example 2)
[0458] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0459] Conventional systems have been able to display advertisements based on passengers' gender, age, and purpose of riding, but no systems have taken into account passengers' emotional state or security risks. This makes it difficult to respond quickly when a passenger exhibits suspicious behavior or is considered a dangerous person. It is also not possible to dynamically change advertisements based on the passenger's emotional state. Therefore, the present invention aims to solve these problems and provide a system that evaluates security risks while taking passengers' emotional state into account, and displays appropriate advertisements and generates warnings.
[0460] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an emotion engine that determines the emotional state of passengers, a warning generation means that evaluates security risks based on the emotional state and generates a warning, and a database comparison means that compares the emotional state with a dangerous person database. This makes it possible to analyze the emotional state of passengers in real time, detect dangerous people early, and send warnings to the driver or security services in real time as necessary.
[0461] "Imaging means" means a camera or other image capture device for capturing facial images of passengers.
[0462] "Audio collection means" means a microphone or other audio capture device for collecting passenger audio.
[0463] The "analysis means" refers to an algorithm or system that analyzes the data obtained from the imaging means and audio collection means and determines the gender, age, and purpose of the passenger.
[0464] The "advertising generation means" is a program or system for generating optimal advertisements based on passenger information determined by the analysis means.
[0465] The "display means" refers to a display or screen for displaying the advertisements generated by the advertisement generating means to passengers.
[0466] A "QR code generator" is a program or system for generating a QR code to provide a coupon or link related to an advertisement.
[0467] The "emotion engine" is an algorithm or system for determining the emotional state of passengers based on image data and audio data acquired from imaging means.
[0468] An "alert generator" is a program or system for assessing security risks based on emotional states and generating alerts.
[0469] The "database comparison means" is a program or system that compares the information with a database of dangerous people and sends a warning in real time if there is a match.
[0470] The present invention provides a system that evaluates security risks while taking into account the emotional state of passengers, and displays appropriate advertisements and generates warnings. Specific embodiments of the system are described below.
[0471] System Configuration
[0472] The system of the present invention includes the following major components:
[0473] 1. Device:
[0474] Imaging means: A camera installed inside the taxi captures facial images of passengers.
[0475] Voice collection method: Passenger voices are collected using microphones installed inside the taxi.
[0476] Display means: A display for showing advertising or warning messages.
[0477] Emotion engine: Determines the emotional state of passengers based on acquired image and audio data.
[0478] 2. Server:
[0479] Analysis method: Analyzes image and audio data sent from the terminal to determine the passenger's gender, age, and purpose of the ride.
[0480] Advertisement generation means: Generates an optimal advertisement based on the determined information.
[0481] QR code generator: Generates a QR code to provide coupons or links related to the advertisement.
[0482] Warning generation means: Evaluates security risks based on emotional states and generates warnings.
[0483] Database matching means: Matches with a database of dangerous people and sends a real-time alert if there is a match.
[0484] System Operation
[0485] 1. Collecting passenger information:
[0486] When a passenger gets into the taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image, and the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time.
[0487] 2. Data Analysis:
[0488] The server runs a facial recognition algorithm based on the received image data to estimate the passenger's gender and age. It also applies a voice recognition system to the audio data to analyze the passenger's purpose. It also uses an emotion engine to determine the passenger's emotional state based on their facial expressions and voice. If the passenger's emotional state is determined to be dangerous, the warning generation means generates a warning in real time and, if necessary, automatically notifies security services.
[0489] 3. Ad generation:
[0490] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses ad generation AI to create ads for passengers. The ads include text information and graphics tailored to the passenger's gender, age, purpose of the ride, and emotional state. QR codes containing special coupons and related links are also generated and embedded in the advertising content.
[0491] 4. Advertising and Coupon Offers:
[0492] The advertisement data sent from the server is returned to the terminal and provided to passengers through a display. The advertisement contains a QR code, which passengers can scan to receive special coupons and link information.
[0493] Examples of specific examples and prompts
[0494] For example, if a man in his 30s gets into a taxi for a business meeting and feels a little nervous, the system will operate as follows.
[0495] Examples:
[0496] As passengers board, a camera captures their faces and a microphone records the purpose of their trip.
[0497] The server determines the gender as male and the age as in their 30s from the image data, and extracts the keyword "business negotiation" from the voice data.
[0498] The emotion engine detects tension from passengers' facial expressions and voice.
[0499] The server selects the best advertising materials for the business and generates an advertisement containing text and images that will calm the nerves.
[0500] Generate QR codes containing special coupons and links and embed them in advertising content.
[0501] The generated advertisement data is transmitted to the terminal, and the advertisement is displayed on the display.
[0502] Passengers scan the QR code to get reward coupons.
[0503] The passenger's facial expression in response to the displayed advertisement is again captured and the emotional feedback is sent to the server.
[0504] Example prompt sentence:
[0505] "I want to detect passengers exhibiting abnormal behavior and send warnings. In particular, I want real-time alerts to be sent when the emotion is 'angry' or 'fear'."
[0506] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0507] Processing Steps
[0508] Step 1: Acquire passenger facial images and voice
[0509] Input: Passenger boarding the taxi
[0510] Specific behavior:
[0511] The terminal automatically activates the imaging means (camera) and captures the passenger's facial image.
[0512] At the same time, the audio collection means (microphone) starts to collect passengers' conversations.
[0513] Output: Acquired facial image data and audio data
[0514] Step 2: Sending face image and voice data
[0515] Input: Acquired facial image data and voice data
[0516] Specific behavior:
[0517] The terminal transmits the face image data and the voice data to the server.
[0518] Output: Facial image data and voice data sent to the server
[0519] Step 3: Analyze facial images and audio data
[0520] Input: Facial image data and voice data sent to the server
[0521] Specific behavior:
[0522] The server runs a facial recognition algorithm based on the received facial image data to estimate the passenger's gender and age.
[0523] The server applies a voice recognition system to the received voice data and analyzes the purpose of the ride.
[0524] The server uses an emotion engine to determine the passenger's emotional state from their facial expressions and voice.
[0525] Output: Analysis results of gender, age, purpose of ride, emotional state
[0526] Step 4: Assess security risks
[0527] Input: Gender, age, purpose of ride, emotional state analysis results
[0528] Specific behavior:
[0529] The server assesses security risks based on emotional states.
[0530] The server compares the analysis results with a database of dangerous people.
[0531] Output: Security risk assessment results and risky person matching results
[0532] Step 5: Generate Ads
[0533] Input: Analysis results and security risk assessment results
[0534] Specific behavior:
[0535] Based on the analysis results, the server selects the most suitable advertising materials and uses advertising generation AI to create advertisements for passengers.
[0536] The generated advertisements include text information and graphics tailored to the passenger's gender, age, purpose of the trip, and emotional state.
[0537] The server generates a QR code containing a special coupon or related link and embeds it in the advertising content.
[0538] Output: Generated advertising data (text, graphics, QR code)
[0539] Step 6: Generate and send an alert
[0540] Input: Security risk assessment results and risky person matching results
[0541] Specific behavior:
[0542] If the server has a high security risk, a warning generating means generates a warning message.
[0543] The server sends warning messages to designated security services and drivers in real time as needed.
[0544] Output: Generated warning messages and sent warnings
[0545] Step 7: Displaying Advertisements and Warnings
[0546] Input: Generated ad data and warning message
[0547] Specific behavior:
[0548] The terminal receives the generated advertisement data and displays it on the display.
[0549] The device will display warning messages to the driver as needed.
[0550] Output: Advertisements and warnings shown on the display
[0551] Step 8: Collect passenger feedback
[0552] Input: Passenger facial expression while viewing advertisement
[0553] Specific behavior:
[0554] The device then recaptures the passenger's real-time facial expression in response to the displayed advertisement.
[0555] The terminal transmits the acquired feedback data to the server.
[0556] Output: Feedback data
[0557] Step 9: Dynamically adjust ads
[0558] Input: Feedback data
[0559] Specific behavior:
[0560] The server analyzes the feedback data and dynamically adjusts the content of the advertisements.
[0561] Output: Adjusted advertising data
[0562] In this way, the system can assess passengers' emotional state and security risks in real time, providing them with the optimal advertising experience and security.
[0563] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0564] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0565] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0566] [Second embodiment]
[0567] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0568] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0569] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0570] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0571] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0572] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0573] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0574] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0575] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0576] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0577] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0578] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0579] System Configuration
[0580] The system of the present invention includes the following major components:
[0581] 1. Terminal installed in the taxi: Equipped with imaging means, audio collection means, and display means.
[0582] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means.
[0583] 3. Network: Data communication between the terminal and the server.
[0584] System Operation
[0585] Collecting passenger information
[0586] When a passenger gets into a taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image. At the same time, the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time as soon as the passenger gets into the taxi.
[0587] Data analysis
[0588] The server runs a facial recognition algorithm on the received image data to estimate the passenger's gender and age, and applies a voice recognition system to the audio data to analyze the passenger's purpose. The results of these analyses are stored in a database.
[0589] Ad Generation
[0590] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses ad generation AI to create ads for passengers. The ads contain text information and graphics tailored to the passenger's interests and goals. QR codes containing special coupons and related links are also generated and embedded in the advertising content.
[0591] Advertisement display and coupon offer
[0592] The advertisement data sent from the server is returned to the terminal and provided to passengers through a display. The advertisement contains a QR code, which passengers can scan to receive special coupons and link information.
[0593] Specific program processing
[0594] Program processing explanation
[0595] Scenario: A man in his 30s gets into a taxi for a business meeting.
[0596] 1. Information collection during the ride:
[0597] The device "uses a camera to capture passengers' faces and a microphone to record the purpose of the ride."
[0598] Device: "Sends the captured image and audio data to the server"
[0599] 2. Data Analysis:
[0600] Server: "From the received image data, a facial recognition algorithm estimates the gender as male and the age as being in their 30s."
[0601] Server: "The voice recognition system converts the voice data into text and extracts the keyword 'business negotiations'."
[0602] 3. Ad generation:
[0603] Server: "Based on the analysis results, select business-related advertisements from the advertising materials."
[0604] Server: Generates ads that combine business-related videos and text.
[0605] Server: Embed a QR code containing a special coupon or link in the ad.
[0606] 4. Advertising and Coupon Offers:
[0607] Server: "Sends the generated advertising data to the device"
[0608] Device: "The advertisement will be displayed on the tablet screen, and the QR code will be placed in an easily visible position."
[0609] User: "Scan the QR code with your smartphone and receive a special coupon."
[0610] This will improve passengers' in-cab experience by providing tailored advertising, while also enabling advertisers to effectively target their ads and maximize their advertising effectiveness.
[0611] The processing flow will be explained below.
[0612] Program processing steps
[0613] Specific scenario: A man in his 30s gets into a taxi for a business meeting.
[0614] Step 1:
[0615] The device: "The moment a passenger gets into the taxi, the camera activates and captures the passenger's face."
[0616] The device photographs the passenger's face from multiple angles to obtain clear image data.
[0617] Step 2:
[0618] "At the same time, the microphone collects passenger conversations and records voice data in real time."
[0619] The device records the conversation audio in high quality and performs pre-processing to remove noise.
[0620] Step 3:
[0621] Terminal: "Sends collected image data and audio data to the server."
[0622] The terminal encrypts the image and audio data and transmits it to the server using a secure communication protocol.
[0623] Step 4:
[0624] Server: Based on the received image data, it runs a facial recognition algorithm to estimate the passenger's gender and age.
[0625] The server uses a deep learning model to classify the gender as male and the age as in their 30s from the image.
[0626] Step 5:
[0627] Server: "Apply a voice recognition system to the voice data and analyze the purpose of the ride."
[0628] The server uses a speech recognition API to convert the speech into text and extract the keyword "business negotiation."
[0629] Step 6:
[0630] Server: "Save the analysis results in the database"
[0631] The server stores data such as passenger ID, estimated gender, age, and purpose of ride in a database.
[0632] Step 7:
[0633] Server: "Based on the analysis results, we will create optimal ads using ad generation AI."
[0634] The server selects business-related videos and text from advertising materials aimed at businessmen in their 30s.
[0635] Step 8:
[0636] Server: Generate a QR code containing a special coupon or link for the ad and embed it in the ad content.
[0637] The server uses a QR code generation algorithm and integrates it into the advertisement.
[0638] Step 9:
[0639] Server: "Sends the generated advertising data to the taxi terminal."
[0640] The server re-encrypts the advertisement data and transmits it to the terminal.
[0641] Step 10:
[0642] Terminal: "Decodes the received advertising data and displays it on the tablet display."
[0643] The terminal places advertisements in a location that is easy for passengers to see.
[0644] Step 11:
[0645] The terminal "displays the QR code included in the advertisement in front of the passenger."
[0646] Place your device in the center of the screen so that the QR code is easy to scan.
[0647] Step 12:
[0648] User: "Scan the QR code with your smartphone to get a special coupon."
[0649] Users can scan the QR code using their smartphone's camera app to receive the coupon information.
[0650] The above are the specific processing steps of the program of the present invention. In this way, the processing at each step works together to provide passengers with optimal targeted advertising and interactive experiences.
[0651] Example 1
[0652] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0653] Conventional taxi ride systems lacked a method for providing appropriate and effective advertisements to passengers. Furthermore, technology for generating and displaying advertisements tailored to passengers' interests in real time was underdeveloped. As a result, the effectiveness of advertisements was low, and passenger convenience was difficult to improve. The present invention aims to solve these problems and provide more effective and personalized advertisements to passengers.
[0654] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0655] In this invention, the server includes an imaging means for capturing facial images of passengers, an audio collection means for capturing passenger voices, an analysis means for analyzing data obtained from the imaging means and the audio collection means to determine the passenger's gender, age, and purpose of the ride, an advertisement generation means for generating an optimal advertisement based on the passenger information determined by the analysis means, a display means for displaying the advertisement generated by the advertisement generation means, a QR code generation means for providing benefits and links related to the advertisement, a communication means for encrypting data acquired by the imaging means and the audio collection means and transmitting the encrypted data to the server, an image recognition means for executing the facial recognition algorithm to estimate the passenger's gender and age, a speech conversion means for converting speech data into text using the speech recognition system and extracting important keywords, an advertising material selection means for creating optimal advertising content from advertising materials using a generative AI model, and a QR code generation means for generating a QR code containing a benefit coupon or link destination information. This allows passengers to receive advertisements tailored to their needs, improving their in-taxi experience.
[0656] "Imaging means" is a hardware or software component for capturing facial images of passengers.
[0657] "Voice collection means" is a hardware or software component for collecting passenger voices.
[0658] "Analysis means" is a general term for hardware or software for analyzing data obtained from the imaging means and audio collection means and determining the gender, age, and purpose of the passenger's ride.
[0659] The "advertising generation means" is a hardware or software component for generating optimal advertisements based on passenger information determined by the analysis means.
[0660] The "display means" refers to a display device or software for visually presenting the advertisements generated by the advertisement generating means to passengers.
[0661] A "QR code generator" is a hardware or software component that generates a QR code to provide a special offer or link related to an advertisement.
[0662] The "communication means" is a hardware or software function for encrypting data acquired by the imaging means and the audio collecting means and transmitting the data to the server.
[0663] "Image Recognition Means" means a hardware or software component that runs a facial recognition algorithm and estimates the gender and age of a passenger.
[0664] "Speech conversion means" refers to hardware or software functionality that converts voice data into text using a voice recognition system and extracts important keywords.
[0665] "Advertising material selection means" means a hardware or software function that uses a generative AI model to select and create optimal advertising content from advertising materials.
[0666] The system of the present invention mainly comprises a terminal installed in a taxi, a server that performs data analysis and advertisement generation, and a network that communicates between these components. The operation of the system proceeds as follows.
[0667] System Configuration
[0668] 1. Terminal
[0669] Imaging means: Includes a high-resolution camera for capturing facial images of passengers. The camera is automatically activated when a passenger gets into the taxi and captures facial images. As a specific example, a Sony IMX586 camera is used.
[0670] Audio collection means: Includes a microphone for collecting passengers' voices. The microphone is activated when passengers board the vehicle and records their conversations. As a specific example, a Rode NT-USB microphone is used.
[0671] Display means: includes a high-resolution touchscreen display for displaying advertising content. As a specific example, a 16-inch high-resolution touchscreen display is used.
[0672] Communication means: Includes a communication module for encrypting data acquired by the imaging means and audio collection means and transmitting it to a server. As a specific example, a 4G LTE module is used.
[0673] 2. Server
[0674] Analysis means: includes a software component for analyzing data obtained from the imaging means and audio collection means and determining the gender, age, and purpose of the passenger.
[0675] Image Recognition: Implements a facial recognition algorithm to estimate gender and age. For example, the OpenCV library is used.
[0676] Speech conversion method: A speech recognition system is used to convert the audio data into text and extract important keywords. Specific examples include DeepSpeech and Amazon Transcribe.
[0677] Ad generation methods:
[0678] Advertising material selection means: Based on passenger information determined by the analysis means, the optimal advertising material is selected and advertising content is created using a generative AI model. As a specific example, GPT-4 is used.
[0679] QR code generation method: Generates a QR code containing special coupons and link information. As a concrete example, the Libqrencode library is used.
[0680] 3. Network
[0681] This includes a network for fast and secure data communication between the terminals installed in the taxi and the server. For example, 4G / 5G networks are used.
[0682] Specific examples
[0683] The following is a specific example of the system's operation.
[0684] Scenario: A man in his 30s gets into a taxi for a business meeting.
[0685] 1. Terminal: When a passenger gets into a taxi, the terminal automatically activates the camera to capture the passenger's face image, and simultaneously activates the microphone to record the passenger's conversation.
[0686] 2. Terminal: The acquired facial image and audio data are encrypted and sent to the server.
[0687] 3. Server: Analyzes the received facial image and estimates the passenger's gender as "male" and age as "30s." The voice data is converted into text using a voice recognition system, and the keyword "business negotiation" is extracted.
[0688] 4. Server: Based on the analysis results, the generative AI model generates optimal business-related advertising content. For example, the prompt sentence should include keywords such as "men in their 30s" and "business negotiations" to generate business-related advertising content.
[0689] 5. Server: Embeds a QR code containing a special coupon or link information into the generated advertising content and sends it to the device.
[0690] 6. Terminal: Display the advertisement on the display inside the taxi and place the QR code in a location that is easy to see.
[0691] 7. User: Scans the QR code displayed on the screen with a smartphone to receive special coupons and link information.
[0692] This system will enable passengers to receive advertisements tailored to their needs, improving their in-cab experience, while also enabling advertisers to target their ads more effectively, maximizing their advertising impact.
[0693] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0694] Program processing
[0695] Step 1:
[0696] Device: When a passenger gets into the taxi, the device automatically activates its camera to capture the passenger's facial image, and simultaneously activates its microphone to begin recording the passenger's conversation.
[0697] Input: Passenger riding behavior
[0698] Data processing: Images are taken with a camera and image data (JPEG format) is generated. Audio is collected with a microphone and audio data (WAV format) is generated.
[0699] Output: Facial image data and audio data
[0700] Step 2:
[0701] Terminal: The acquired facial image data and voice data are end-to-end encrypted using SSL and sent to the server.
[0702] Input: Facial image data and audio data
[0703] Data processing: SSL encryption is applied and data communication is converted into secure data.
[0704] Output: Encrypted facial image data and audio data
[0705] Step 3:
[0706] Server: Based on the received facial image data, a facial recognition algorithm (OpenCV library) is executed to estimate the passenger's gender and age. The results are stored in a database.
[0707] Input: Encrypted facial image data
[0708] Data processing: The data is decrypted and facial recognition algorithms are used to generate gender and age data.
[0709] Output: Gender and age information
[0710] Step 4:
[0711] Server: The voice data is converted into text using a speech recognition system (DeepSpeech or Amazon Transcribe), and the purpose of the ride is analyzed using a natural language processing (NLP) algorithm. The analysis results are stored in a database.
[0712] Input: Encrypted audio data
[0713] Data processing: The data is decoded and generated into text data using a speech recognition system, after which keywords are extracted using NLP algorithms.
[0714] Output: Keywords for trip purpose
[0715] Step 5:
[0716] Server: Based on the analysis results (gender, age, purpose of ride), a generative AI model (GPT-4) is used to generate the optimal advertisement. Examples of prompt sentences include keywords such as "male in his 30s" and "business negotiation."
[0717] Input: Gender, age, purpose of ride keywords
[0718] Data processing: Input prompt text into the generative AI model to generate advertising content.
[0719] Output: Ad content data
[0720] Step 6:
[0721] Server: Generates a QR code containing a special coupon or link information in the generated advertising content and embeds it in the advertising content.
[0722] Input: Ad content data
[0723] Data processing: Use a QR code generation library to generate a QR code containing special coupon information and link information. This can then be incorporated into advertising content.
[0724] Output: QR code embedded advertising content data
[0725] Step 7:
[0726] Server: Sends advertising content data with an embedded QR code to the terminal.
[0727] Input: QR code embedded advertising content data
[0728] Data processing: The advertising content data is encrypted and prepared for transmission.
[0729] Output: Encrypted advertising content data
[0730] Step 8:
[0731] Terminal: Display the received advertising content data on the display and place the QR code in an easily visible location.
[0732] Input: Encrypted ad content data
[0733] Data processing: Decrypting the data and optimizing it for display.
[0734] Output: Advertising content and QR code displayed on the display
[0735] Step 9:
[0736] User: Scans the QR code displayed on the screen with a smartphone to receive special coupons and link information.
[0737] Input: QR code displayed on the screen
[0738] Data processing: Scan the QR code with your smartphone to obtain special coupons and link information.
[0739] Output: Bonus coupons and link information saved on your smartphone
[0740] This ensures smooth operation of the entire system, improves passengers' in-cab experience by allowing them to receive advertisements tailored to their needs, and enables advertisers to effectively target their advertisements, maximizing their advertising effectiveness.
[0741] (Application example 1)
[0742] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0743] Conventional in-taxi advertising systems have the challenge of making it difficult to provide advertisements tailored to passengers' interests and needs in real time. Furthermore, because advertisements are displayed on fixed monitors, there is no guarantee that passengers will see them. In such situations, advertising effectiveness is not fully realized, resulting in an unsatisfactory experience for both advertisers and passengers.
[0744] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0745] In this invention, the server includes an imaging means for capturing facial images of passengers, an audio collecting means for collecting passenger voices, an analysis means for analyzing data obtained from the imaging means and the audio collecting means and determining the gender, age, and purpose of the passenger's ride, an advertisement generating means for generating an optimal advertisement based on the passenger information determined by the analysis means, a display means for displaying the advertisement generated by the advertisement generating means, a QR code generating means for providing coupons and links related to the advertisement, and a display means for identifying that the display means is a smart glasses display. This allows advertisements tailored to individual passenger needs to be provided in real time and the advertisements to be displayed in a location that is easily visible through the smart glasses.
[0746] "Imaging means" refers to a function or device used to capture facial images of passengers.
[0747] "Voice collection means" refers to the functionality or device used to collect passenger voices.
[0748] The "analysis means" is a function or device that analyzes the data obtained from the imaging means and the audio collection means and determines the gender, age, and purpose of the passenger.
[0749] The "advertisement generation means" is a function or device that generates optimal advertisements based on passenger information determined by the analysis means.
[0750] The "display means" refers to a function or device that displays the advertisement generated by the advertisement generation means, and includes the display of the smart glasses.
[0751] The term "QR code generating means" refers to a function or device that generates a QR code for providing a coupon or link related to an advertisement.
[0752] System Configuration
[0753] An embodiment of the present invention includes the following major components:
[0754] 1. Terminal: Equipped with an imaging means (camera), an audio collection means (microphone), and a display means (display) for displaying advertisements, all of which are installed in the smart glasses.
[0755] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means for providing coupons and links.
[0756] 3. Network: Data communication between the terminal and the server.
[0757] System Operation
[0758] Collecting passenger information
[0759] When a user puts on the smart glasses, the camera automatically captures facial images and the microphone starts recording the user's conversation, and these data are sent to the server in real time.
[0760] Data analysis
[0761] The server analyzes the received image and voice data, using a facial recognition algorithm to estimate the user's gender and age, and a voice recognition system to analyze the user's intent based on the voice data.
[0762] Ad Generation
[0763] Based on the analysis results, the server generates the most suitable advertisement, which includes text information and graphics tailored to the user's interests and goals, as well as a QR code containing special coupons and related links.
[0764] Advertisement display and coupon offer
[0765] The generated advertising data is displayed on the smart glasses display, and users can receive special coupons and link information by scanning the QR code.
[0766] Specific program processing explanation
[0767] The server uses a facial recognition library (e.g., OpenCV) and a voice recognition library (e.g., SpeechRecognition) to analyze the user's facial image and voice data. This allows the server to extract the user's gender, age, and purpose, and then uses ad generation AI to create optimal ads. It also generates QR codes containing coupons and link information, and incorporates them into the ad content.
[0768] Specific examples
[0769] Scenario: A woman in her twenties puts on smart glasses while out shopping.
[0770] 1. The camera in the smart glasses captures the woman's face, and the microphone collects the audio of her saying "shopping."
[0771] 2. The server uses facial recognition to determine that the user is a woman in her 20s, and uses voice analysis to detect the keyword "shopping."
[0772] 3. Generate ads for the latest fashion and cosmetics for women and embed QR codes containing special coupons and shop links into the ads.
[0773] 4. Advertisements will be displayed on the smart glasses display, and users can scan the QR code to receive reward coupons.
[0774] Prompt Sentence Examples
[0775] Give your generative AI model the following inputs:
[0776] User information: Female in her 20s
[0777] Keywords: Shopping
[0778] Generated advertising content: Latest fashion, cosmetics advertising
[0779] QR code content: Bonus coupon, shop link QR code
[0780] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0781] Step 1:
[0782] The device detects that the user has put on the smart glasses. At this time, the camera is activated and an image of the user's face is captured. The captured image is saved in a data format (e.g., JPEG format) and sent to the server.
[0783] Input: User's face image data
[0784] Output: JPEG format face image file
[0785] Step 2:
[0786] The device collects the user's voice using the microphone in the smart glasses. The collected voice data is saved as an audio file (e.g., WAV format) and sent to the server.
[0787] Input: User's voice data
[0788] Output: WAV format audio file
[0789] Step 3:
[0790] The server receives the facial image data and uses a facial recognition algorithm (e.g., OpenCV) to estimate the user's gender and age. During this process, image processing operations are performed to obtain the estimated gender and age.
[0791] Input: JPEG format face image file
[0792] Output: User's gender and age (e.g., female in her 20s)
[0793] Step 4:
[0794] The server receives the voice data and converts the voice data into text data using a voice recognition system (e.g., SpeechRecognition). Keywords (e.g., "shopping") are extracted from the converted text data.
[0795] Input: WAV format audio file
[0796] Output: Text data and extracted keywords
[0797] Step 5:
[0798] The server uses ad generation AI to generate optimal ads based on user information (gender, age, keywords) obtained through facial and voice recognition. During this process, ad materials are selected and combined.
[0799] Input: User information (gender, age, keywords)
[0800] Output: Generated advertising data (text information, graphics, QR code)
[0801] Step 6:
[0802] The server transmits the generated advertisement data to the terminal, and the terminal displays the received advertisement data on the display of the smart glasses.
[0803] Input: Generated ad data
[0804] Output: Advertisement displayed on the smart glasses display
[0805] Step 7:
[0806] Users scan the QR code displayed on the smart glasses display with their smartphone to receive reward coupons and shop links.
[0807] Input: QR code displayed on the screen
[0808] Output: Special coupons and shop links acquired by the user's smartphone
[0809] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0810] System Configuration
[0811] The present invention includes the following major components:
[0812] 1. Terminal installed in the taxi: equipped with imaging means, audio collection means, display means, and emotion engine.
[0813] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means.
[0814] 3. Network: Data communication between the terminal and the server.
[0815] System Operation
[0816] Collecting passenger information
[0817] When a passenger gets into a taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image. At the same time, the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time as soon as the passenger gets into the taxi.
[0818] Data analysis
[0819] The server runs a facial recognition algorithm based on the received image data to estimate the passenger's gender and age. It also applies a voice recognition system to the audio data to analyze the passenger's purpose. It also uses an emotion engine to determine the passenger's emotional state based on their facial expressions and voice. The results of these analyses are stored in a database.
[0820] Ad Generation
[0821] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses ad generation AI to create ads for passengers. The ads include text information and graphics tailored to the passenger's gender, age, purpose of the ride, and emotional state. QR codes containing special coupons and related links are also generated and embedded in the advertising content.
[0822] Advertisement display and coupon offer
[0823] The advertising data sent from the server is returned to the terminal and provided to passengers via a display. The advertisement contains a QR code, which passengers can scan to receive special coupons or link information. In addition, real-time emotional feedback from passengers regarding the displayed advertisement is collected again, and the content of the advertisement is dynamically adjusted to provide a more effective advertising experience.
[0824] Specific program processing
[0825] Program processing explanation
[0826] Scenario: A man in his 30s gets into a taxi for a business meeting and feels a little nervous.
[0827] 1. Information collection during the ride:
[0828] The device "uses a camera to capture passengers' faces and a microphone to record the purpose of the ride."
[0829] Device: "Sends the captured image and audio data to the server"
[0830] 2. Data Analysis:
[0831] Server: "From the received image data, a facial recognition algorithm determines the gender as male and the age as in their 30s."
[0832] Server: "The voice recognition system converts the voice data into text and extracts the keyword 'business negotiations'."
[0833] Server: "Using an emotion engine to detect tension levels from passengers' facial expressions and voices."
[0834] 3. Ad generation:
[0835] Server: "Based on the analysis results, select the most suitable business-related advertising materials."
[0836] The server "generates ads containing text and images that soften products and services and incorporates content that reduces tension."
[0837] Server: Generates QR codes containing special coupons and links and embeds them in advertising content.
[0838] 4. Advertising and Coupon Offers:
[0839] Server: "Sends the generated advertising data to the device"
[0840] Device: "The advertisement will be displayed on the tablet screen, and the QR code will be placed in an easily visible position."
[0841] User: "Scan the QR code with your smartphone to get a special coupon."
[0842] The device "recaptures the passenger's facial expression in response to the displayed advertisement and sends emotional feedback to the server."
[0843] This allows passengers to see ads that reflect their real-time emotional state, providing them with an optimal advertising experience, while enabling advertisers to achieve more effective targeting and interactive advertising methods.
[0844] The processing flow will be explained below.
[0845] Program processing steps
[0846] Specific scenario: A man in his 30s is getting into a taxi for a business meeting and is feeling a bit nervous.
[0847] Step 1:
[0848] The moment a passenger gets into the taxi, the camera activates and captures the passenger's facial image.
[0849] The device photographs the passenger's face from multiple angles to obtain clear image data.
[0850] Step 2:
[0851] "At the same time, the microphone collects passenger conversations and records voice data in real time."
[0852] The device records passengers' conversations in high quality and performs pre-processing to remove noise.
[0853] Step 3:
[0854] The device "encodes the collected image and audio data and transmits it to a server using a secure communication protocol."
[0855] The terminal encrypts the data to ensure secure communication.
[0856] Step 4:
[0857] Server: Based on the received image data, it runs a facial recognition algorithm and determines the passenger's gender as male and age as in their 30s.
[0858] The server uses a deep learning model to estimate the gender and age of passengers from facial images.
[0859] Step 5:
[0860] Server: "Apply a voice recognition system to the voice data and analyze the purpose of the ride."
[0861] The server uses a speech recognition API to convert the speech into text data and extract the keyword "business negotiation."
[0862] Step 6:
[0863] Server: "Uses an emotion engine to determine the passenger's emotional state from their facial expressions and voice."
[0864] The server performs facial expression analysis and voice tone analysis to detect the passenger's state of tension.
[0865] Step 7:
[0866] Server: "Save the analysis results in the database"
[0867] The server stores passengers' gender, age, purpose of the ride, and emotional state data in a database.
[0868] Step 8:
[0869] Server: "Based on the analysis results, we will use ad generation AI to create ads for passengers."
[0870] The server selects advertising materials that are business-related and de-escalating, and generates customized advertisements.
[0871] Step 9:
[0872] Server: Generates QR codes containing special coupons and links and embeds them in advertising content.
[0873] The server uses a QR code generation algorithm and binds it to the advertisement.
[0874] Step 10:
[0875] Server: "Sends the generated advertising data to the taxi terminal."
[0876] The server encodes the advertisement data and transmits it to the terminal.
[0877] Step 11:
[0878] Terminal: "Decodes the received advertising data and displays it on the tablet display."
[0879] The terminal displays advertising content in real time and places the advertisements in a location that is easy for passengers to see.
[0880] Step 12:
[0881] The terminal "displays the QR code included in the advertisement in front of the passenger."
[0882] Place your device in the center of the screen so that the QR code is easy to scan.
[0883] Step 13:
[0884] User: "Scan the QR code with your smartphone to get a special coupon."
[0885] Users scan the QR code with their smartphone's camera app to obtain coupon and link information.
[0886] Step 14:
[0887] The device "recaptures the passenger's facial expressions in response to the displayed advertisement and obtains emotional feedback."
[0888] The device then takes another photograph of the passenger's face and analyzes their emotional state.
[0889] Step 15:
[0890] The device "sends emotional feedback to the server and readjusts the advertising content."
[0891] The terminal transmits the obtained emotional feedback to the server and dynamically updates the advertisement content.
[0892] This allows passengers to see ads that are tailored to their real-time emotional state, providing a more satisfying advertising experience for passengers and enabling advertisers to deliver more targeted and effective ads.
[0893] Example 2
[0894] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0895] Conventional in-taxi advertising systems do not optimize advertisements based on the passenger's gender, age, and purpose of the ride, nor do they optimize advertisements based on the passenger's current emotional state. This makes it difficult to effectively deliver advertisements that are most appropriate for each passenger. It is also difficult to grasp the effectiveness of advertisements in real time and dynamically adjust them. To solve these problems, the present invention aims to provide optimal advertisements based on detailed passenger information, including the passenger's emotional state, and to provide feedback on the effectiveness of advertisements in real time.
[0896] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0897] In this invention, the server includes an analysis means for analyzing the received data and determining the gender, age, purpose of riding, and emotional state of the passenger, an advertisement generation means for generating an optimal advertisement based on the passenger information and emotional state determined by the analysis means, and a feedback collection means for acquiring passenger emotional feedback on the displayed advertisement and performing additional analysis. This makes it possible to provide optimal advertisements according to the emotional state of each passenger and provide feedback on the effectiveness of the advertisements in real time.
[0898] The "imaging means" is a device for capturing facial images of passengers.
[0899] "Voice collection means" refers to equipment for collecting passenger voices.
[0900] The "communication means" is a function for transmitting data obtained from the imaging means and the sound collecting means to a server in real time.
[0901] "Analysis means" refers to a device or program that analyzes the received data and determines the passenger's gender, age, purpose of the ride, and emotional state.
[0902] The "advertisement generation means" is a device or program for generating an optimal advertisement based on the passenger information and emotional state determined by the analysis means.
[0903] The "display means" is a device for displaying the advertisement generated by the advertisement generation means.
[0904] The "QR code generating means" is a function for generating a QR code for providing a coupon or link related to the advertisement.
[0905] The "feedback collection means" is a function for obtaining passengers' emotional feedback regarding the displayed advertisements and for performing additional analysis.
[0906] The "image analysis means" is a device or program for detecting the age and gender of passengers using image data acquired from the imaging means.
[0907] The "voice analysis means" is a device or program for identifying the purpose of riding and the emotional state of the passenger using the voice data obtained from the voice collection means.
[0908] The "generative AI model" is an artificial intelligence model that selects advertising materials and generates advertising content based on the passenger's gender, age, purpose of riding, and emotional state.
[0909] System Configuration
[0910] The present invention includes the following major components:
[0911] 1. Terminal installed in the taxi: equipped with imaging means, audio collection means, display means, and emotion engine.
[0912] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means.
[0913] 3. Network: Data communication between the terminal and the server.
[0914] Hardware and software used
[0915] Hardware:
[0916] Device imaging means (camera): High-resolution camera
[0917] Audio collection means (microphone): High-sensitivity microphone
[0918] Display means (display): Tablet display
[0919] software:
[0920] Face Recognition Algorithm: Image Processing Library
[0921] Speech Recognition System: Speech Recognition API
[0922] Emotion Engine: Emotion Analysis Tool
[0923] Ad Generation AI: Generative AI Model
[0924] System Operation
[0925] Collecting passenger information
[0926] When a passenger gets into a taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image. At the same time, the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time as soon as the passenger gets into the taxi.
[0927] Data analysis
[0928] The server runs a facial recognition algorithm based on the received image data to estimate the passenger's gender and age. It also applies a voice recognition system to the audio data to analyze the passenger's purpose. It also uses an emotion engine to determine the passenger's emotional state based on their facial expressions and voice. The results of these analyses are stored in a database.
[0929] Ad Generation
[0930] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses a generative AI model to create advertisements for passengers. The advertisements include text information and graphics tailored to the passenger's gender, age, purpose of the ride, and emotional state. QR codes containing special coupons and related links are also generated and embedded in the advertisement content.
[0931] Advertisement display and coupon offer
[0932] The advertising data sent from the server is returned to the terminal and provided to passengers via a display. The advertisement contains a QR code, which passengers can scan to receive special coupons or link information. In addition, real-time emotional feedback from passengers regarding the displayed advertisement is collected again, and the content of the advertisement is dynamically adjusted to provide a more effective advertising experience.
[0933] Specific examples
[0934] Scenario: A man in his 30s gets into a taxi for a business meeting and feels a little nervous.
[0935] 1. Information collection during the ride:
[0936] The device "uses a camera to capture passengers' faces and a microphone to record the purpose of the ride."
[0937] Device: "Sends the captured image and audio data to the server"
[0938] 2. Data Analysis:
[0939] Server: "From the received image data, a facial recognition algorithm determines the gender as male and the age as in their 30s."
[0940] Server: "The voice recognition system converts the voice data into text and extracts the keyword 'business negotiations'."
[0941] Server: "Using an emotion engine to detect tension levels from passengers' facial expressions and voices."
[0942] 3. Ad generation:
[0943] Server: "Based on the analysis results, select the most suitable business-related advertising materials."
[0944] The server "generates ads that include text and images that soften the product's features and characteristics, incorporating tension-reducing content."
[0945] Server: Generates QR codes containing special coupons and links and embeds them in advertising content.
[0946] 4. Advertising and Coupon Offers:
[0947] Server: "Sends the generated advertising data to the device"
[0948] Terminal: "Display advertisements on the display and place QR codes in an easily visible location."
[0949] User: "Scan the QR code with your smartphone to get a special coupon."
[0950] The device "recaptures the passenger's facial expression in response to the displayed advertisement and sends emotional feedback to the server."
[0951] Examples of prompts:
[0952] Prompt statement:
[0953] Please generate an ad based on the following criteria:
[0954] Gender: Male
[0955] Age: 30s
[0956] Purpose of the ride: Business negotiations
[0957] Emotional state: Tension
[0958] Ads should include text and images with tension-busting content, as well as a QR code with a special coupon or relevant link."
[0959] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0960] Step 1:
[0961] Gathering information when riding
[0962] Specific operation: The device automatically activates the camera (photography means) and microphone (audio collection means). The camera captures the passenger's facial image, and the microphone records the passenger's conversation.
[0963] Input: Passenger face image, passenger voice data
[0964] Data processing / calculation: capturing facial images, collecting voice data
[0965] Output: Acquired facial image data and audio data
[0966] Step 2:
[0967] Sending data
[0968] Specific operation: The device sends the collected facial image data and voice data to the server in real time. Data communication is performed via the network.
[0969] Input: Facial image data, audio data
[0970] Data processing / calculation: Data encryption and transmission
[0971] Output: Facial image data and audio data sent to the server
[0972] Step 3:
[0973] Data analysis
[0974] What it does: The server analyzes the data it receives. It uses a facial recognition algorithm to estimate gender and age, a voice recognition system to convert the audio data into text, and an emotion engine to determine the passenger's emotional state.
[0975] Input: Received facial image data and audio data
[0976] Data processing / calculation: facial recognition, voice recognition, emotion analysis
[0977] Output: Gender, age, purpose of ride, emotional state
[0978] Step 4:
[0979] Ad Generation
[0980] What it does: The server generates the best ads based on the analysis results, uses generative AI models to create text information and graphics, and creates QR codes with special offers and links.
[0981] Input: Gender, age, purpose of ride, emotional state
[0982] Data processing / calculation: Generating advertising content, generating QR codes
[0983] Output: Advertisement data (text, images, QR code, etc.)
[0984] Step 5:
[0985] Advertisement display and coupon offer
[0986] Specific operation: The advertising data sent from the server is returned to the device and displayed on the screen. The user can then scan the QR code with their smartphone to receive a special coupon.
[0987] Input: Ad data
[0988] Data processing / calculation: Display of advertising data
[0989] Output: Displayed ad content, user-acquired reward coupon
[0990] Step 6:
[0991] Gathering feedback
[0992] How it works: The device recaptures passenger responses after the ad is displayed and sends the data to the server, which performs additional analysis and dynamically adjusts the ad content.
[0993] Input: Recollected facial image data, emotion feedback
[0994] Data processing / calculation: Feedback analysis, dynamic adjustment of advertising content
[0995] Output: Adjusted advertising data
[0996] (Application example 2)
[0997] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0998] Conventional systems have been able to display advertisements based on passengers' gender, age, and purpose of riding, but no systems have taken into account passengers' emotional state or security risks. This makes it difficult to respond quickly when a passenger exhibits suspicious behavior or is considered a dangerous person. It is also not possible to dynamically change advertisements based on the passenger's emotional state. Therefore, the present invention aims to solve these problems and provide a system that evaluates security risks while taking passengers' emotional state into account, and displays appropriate advertisements and generates warnings.
[0999] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an emotion engine that determines the emotional state of passengers, a warning generation means that evaluates security risks based on the emotional state and generates a warning, and a database comparison means that compares the emotional state with a dangerous person database. This makes it possible to analyze the emotional state of passengers in real time, detect dangerous people early, and send warnings to the driver or security services in real time as necessary.
[1000] "Imaging means" means a camera or other image capture device for capturing facial images of passengers.
[1001] "Audio collection means" means a microphone or other audio capture device for collecting passenger audio.
[1002] The "analysis means" refers to an algorithm or system that analyzes the data obtained from the imaging means and audio collection means and determines the gender, age, and purpose of the passenger.
[1003] The "advertising generation means" is a program or system for generating optimal advertisements based on passenger information determined by the analysis means.
[1004] The "display means" refers to a display or screen for displaying the advertisements generated by the advertisement generating means to passengers.
[1005] A "QR code generator" is a program or system for generating a QR code to provide a coupon or link related to an advertisement.
[1006] The "emotion engine" is an algorithm or system for determining the emotional state of passengers based on image data and audio data acquired from imaging means.
[1007] An "alert generator" is a program or system for assessing security risks based on emotional states and generating alerts.
[1008] The "database comparison means" is a program or system that compares the information with a database of dangerous people and sends a warning in real time if there is a match.
[1009] The present invention provides a system that evaluates security risks while taking into account the emotional state of passengers, and displays appropriate advertisements and generates warnings. Specific embodiments of the system are described below.
[1010] System Configuration
[1011] The system of the present invention includes the following major components:
[1012] 1. Device:
[1013] Imaging means: A camera installed inside the taxi captures facial images of passengers.
[1014] Voice collection method: Passenger voices are collected using microphones installed inside the taxi.
[1015] Display means: A display for showing advertising or warning messages.
[1016] Emotion engine: Determines the emotional state of passengers based on acquired image and audio data.
[1017] 2. Server:
[1018] Analysis method: Analyzes image and audio data sent from the terminal to determine the passenger's gender, age, and purpose of the ride.
[1019] Advertisement generation means: Generates an optimal advertisement based on the determined information.
[1020] QR code generator: Generates a QR code to provide coupons or links related to the advertisement.
[1021] Warning generation means: Evaluates security risks based on emotional states and generates warnings.
[1022] Database matching means: Matches with a database of dangerous people and sends a real-time alert if there is a match.
[1023] System Operation
[1024] 1. Collecting passenger information:
[1025] When a passenger gets into the taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image, and the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time.
[1026] 2. Data Analysis:
[1027] The server runs a facial recognition algorithm based on the received image data to estimate the passenger's gender and age. It also applies a voice recognition system to the audio data to analyze the passenger's purpose. It also uses an emotion engine to determine the passenger's emotional state based on their facial expressions and voice. If the passenger's emotional state is determined to be dangerous, the warning generation means generates a warning in real time and, if necessary, automatically notifies security services.
[1028] 3. Ad generation:
[1029] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses ad generation AI to create ads for passengers. The ads include text information and graphics tailored to the passenger's gender, age, purpose of the ride, and emotional state. QR codes containing special coupons and related links are also generated and embedded in the advertising content.
[1030] 4. Advertising and Coupon Offers:
[1031] The advertisement data sent from the server is returned to the terminal and provided to passengers through a display. The advertisement contains a QR code, which passengers can scan to receive special coupons and link information.
[1032] Examples of specific examples and prompts
[1033] For example, if a man in his 30s gets into a taxi for a business meeting and feels a little nervous, the system will operate as follows.
[1034] Examples:
[1035] As passengers board, a camera captures their faces and a microphone records the purpose of their trip.
[1036] The server determines the gender as male and the age as in their 30s from the image data, and extracts the keyword "business negotiation" from the voice data.
[1037] The emotion engine detects tension from passengers' facial expressions and voice.
[1038] The server selects the best advertising materials for the business and generates an advertisement containing text and images that will calm the nerves.
[1039] Generate QR codes containing special coupons and links and embed them in advertising content.
[1040] The generated advertisement data is transmitted to the terminal, and the advertisement is displayed on the display.
[1041] Passengers scan the QR code to get reward coupons.
[1042] The passenger's facial expression in response to the displayed advertisement is again captured and the emotional feedback is sent to the server.
[1043] Example prompt sentence:
[1044] "I want to detect passengers exhibiting abnormal behavior and send warnings. In particular, I want real-time alerts to be sent when the emotion is 'angry' or 'fear'."
[1045] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1046] Processing Steps
[1047] Step 1: Acquire passenger facial images and voice
[1048] Input: Passenger boarding the taxi
[1049] Specific behavior:
[1050] The terminal automatically activates the imaging means (camera) and captures the passenger's facial image.
[1051] At the same time, the audio collection means (microphone) starts to collect passengers' conversations.
[1052] Output: Acquired facial image data and audio data
[1053] Step 2: Sending face image and voice data
[1054] Input: Acquired facial image data and voice data
[1055] Specific behavior:
[1056] The terminal transmits the face image data and the voice data to the server.
[1057] Output: Facial image data and voice data sent to the server
[1058] Step 3: Analyze facial images and audio data
[1059] Input: Facial image data and voice data sent to the server
[1060] Specific behavior:
[1061] The server runs a facial recognition algorithm based on the received facial image data to estimate the passenger's gender and age.
[1062] The server applies a voice recognition system to the received voice data and analyzes the purpose of the ride.
[1063] The server uses an emotion engine to determine the passenger's emotional state from their facial expressions and voice.
[1064] Output: Analysis results of gender, age, purpose of ride, emotional state
[1065] Step 4: Assess security risks
[1066] Input: Gender, age, purpose of ride, emotional state analysis results
[1067] Specific behavior:
[1068] The server assesses security risks based on emotional states.
[1069] The server compares the analysis results with a database of dangerous people.
[1070] Output: Security risk assessment results and risky person matching results
[1071] Step 5: Generate Ads
[1072] Input: Analysis results and security risk assessment results
[1073] Specific behavior:
[1074] Based on the analysis results, the server selects the most suitable advertising materials and uses advertising generation AI to create advertisements for passengers.
[1075] The generated advertisements include text information and graphics tailored to the passenger's gender, age, purpose of the trip, and emotional state.
[1076] The server generates a QR code containing a special coupon or related link and embeds it in the advertising content.
[1077] Output: Generated advertising data (text, graphics, QR code)
[1078] Step 6: Generate and send an alert
[1079] Input: Security risk assessment results and risky person matching results
[1080] Specific behavior:
[1081] If the server has a high security risk, a warning generating means generates a warning message.
[1082] The server sends warning messages to designated security services and drivers in real time as needed.
[1083] Output: Generated warning messages and sent warnings
[1084] Step 7: Displaying Advertisements and Warnings
[1085] Input: Generated ad data and warning message
[1086] Specific behavior:
[1087] The terminal receives the generated advertisement data and displays it on the display.
[1088] The device will display warning messages to the driver as needed.
[1089] Output: Advertisements and warnings shown on the display
[1090] Step 8: Collect passenger feedback
[1091] Input: Passenger facial expression while viewing advertisement
[1092] Specific behavior:
[1093] The device then recaptures the passenger's real-time facial expression in response to the displayed advertisement.
[1094] The terminal transmits the acquired feedback data to the server.
[1095] Output: Feedback data
[1096] Step 9: Dynamically adjust ads
[1097] Input: Feedback data
[1098] Specific behavior:
[1099] The server analyzes the feedback data and dynamically adjusts the content of the advertisements.
[1100] Output: Adjusted advertising data
[1101] In this way, the system can assess passengers' emotional state and security risks in real time, providing them with the optimal advertising experience and security.
[1102] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1103] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1104] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1105] [Third embodiment]
[1106] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1107] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1108] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1109] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1110] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1111] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1112] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1113] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1114] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1115] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1116] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1117] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1118] System Configuration
[1119] The system of the present invention includes the following major components:
[1120] 1. Terminal installed in the taxi: Equipped with imaging means, audio collection means, and display means.
[1121] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means.
[1122] 3. Network: Data communication between the terminal and the server.
[1123] System Operation
[1124] Collecting passenger information
[1125] When a passenger gets into a taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image. At the same time, the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time as soon as the passenger gets into the taxi.
[1126] Data analysis
[1127] The server runs a facial recognition algorithm on the received image data to estimate the passenger's gender and age, and applies a voice recognition system to the audio data to analyze the passenger's purpose. The results of these analyses are stored in a database.
[1128] Ad Generation
[1129] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses ad generation AI to create ads for passengers. The ads contain text information and graphics tailored to the passenger's interests and goals. QR codes containing special coupons and related links are also generated and embedded in the advertising content.
[1130] Advertisement display and coupon offer
[1131] The advertisement data sent from the server is returned to the terminal and provided to passengers through a display. The advertisement contains a QR code, which passengers can scan to receive special coupons and link information.
[1132] Specific program processing
[1133] Program processing explanation
[1134] Scenario: A man in his 30s gets into a taxi for a business meeting.
[1135] 1. Information collection during the ride:
[1136] The device "uses a camera to capture passengers' faces and a microphone to record the purpose of the ride."
[1137] Device: "Sends the captured image and audio data to the server"
[1138] 2. Data Analysis:
[1139] Server: "From the received image data, a facial recognition algorithm estimates the gender as male and the age as being in their 30s."
[1140] Server: "The voice recognition system converts the voice data into text and extracts the keyword 'business negotiations'."
[1141] 3. Ad generation:
[1142] Server: "Based on the analysis results, select business-related advertisements from the advertising materials."
[1143] Server: Generates ads that combine business-related videos and text.
[1144] Server: Embed a QR code containing a special coupon or link in the ad.
[1145] 4. Advertising and Coupon Offers:
[1146] Server: "Sends the generated advertising data to the device"
[1147] Device: "The advertisement will be displayed on the tablet screen, and the QR code will be placed in an easily visible position."
[1148] User: "Scan the QR code with your smartphone and receive a special coupon."
[1149] This will improve passengers' in-cab experience by providing tailored advertising, while also enabling advertisers to effectively target their ads and maximize their advertising effectiveness.
[1150] The processing flow will be explained below.
[1151] Program processing steps
[1152] Specific scenario: A man in his 30s gets into a taxi for a business meeting.
[1153] Step 1:
[1154] The device: "The moment a passenger gets into the taxi, the camera activates and captures the passenger's face."
[1155] The device photographs the passenger's face from multiple angles to obtain clear image data.
[1156] Step 2:
[1157] "At the same time, the microphone collects passenger conversations and records voice data in real time."
[1158] The device records the conversation audio in high quality and performs pre-processing to remove noise.
[1159] Step 3:
[1160] Terminal: "Sends collected image data and audio data to the server."
[1161] The terminal encrypts the image and audio data and transmits it to the server using a secure communication protocol.
[1162] Step 4:
[1163] Server: Based on the received image data, it runs a facial recognition algorithm to estimate the passenger's gender and age.
[1164] The server uses a deep learning model to classify the gender as male and the age as in their 30s from the image.
[1165] Step 5:
[1166] Server: "Apply a voice recognition system to the voice data and analyze the purpose of the ride."
[1167] The server uses a speech recognition API to convert the speech into text and extract the keyword "business negotiation."
[1168] Step 6:
[1169] Server: "Save the analysis results in the database"
[1170] The server stores data such as passenger ID, estimated gender, age, and purpose of ride in a database.
[1171] Step 7:
[1172] Server: "Based on the analysis results, we will create optimal ads using ad generation AI."
[1173] The server selects business-related videos and text from advertising materials aimed at businessmen in their 30s.
[1174] Step 8:
[1175] Server: Generate a QR code containing a special coupon or link for the ad and embed it in the ad content.
[1176] The server uses a QR code generation algorithm and integrates it into the advertisement.
[1177] Step 9:
[1178] Server: "Sends the generated advertising data to the taxi terminal."
[1179] The server re-encrypts the advertisement data and transmits it to the terminal.
[1180] Step 10:
[1181] Terminal: "Decodes the received advertising data and displays it on the tablet display."
[1182] The terminal places advertisements in a location that is easy for passengers to see.
[1183] Step 11:
[1184] The terminal "displays the QR code included in the advertisement in front of the passenger."
[1185] Place your device in the center of the screen so that the QR code is easy to scan.
[1186] Step 12:
[1187] User: "Scan the QR code with your smartphone to get a special coupon."
[1188] Users can scan the QR code using their smartphone's camera app to receive the coupon information.
[1189] The above are the specific processing steps of the program of the present invention. In this way, the processing at each step works together to provide passengers with optimal targeted advertising and interactive experiences.
[1190] Example 1
[1191] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1192] Conventional taxi ride systems lacked a method for providing appropriate and effective advertisements to passengers. Furthermore, technology for generating and displaying advertisements tailored to passengers' interests in real time was underdeveloped. As a result, the effectiveness of advertisements was low, and passenger convenience was difficult to improve. The present invention aims to solve these problems and provide more effective and personalized advertisements to passengers.
[1193] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1194] In this invention, the server includes an imaging means for capturing facial images of passengers, an audio collection means for capturing passenger voices, an analysis means for analyzing data obtained from the imaging means and the audio collection means to determine the passenger's gender, age, and purpose of the ride, an advertisement generation means for generating an optimal advertisement based on the passenger information determined by the analysis means, a display means for displaying the advertisement generated by the advertisement generation means, a QR code generation means for providing benefits and links related to the advertisement, a communication means for encrypting data acquired by the imaging means and the audio collection means and transmitting the encrypted data to the server, an image recognition means for executing the facial recognition algorithm to estimate the passenger's gender and age, a speech conversion means for converting speech data into text using the speech recognition system and extracting important keywords, an advertising material selection means for creating optimal advertising content from advertising materials using a generative AI model, and a QR code generation means for generating a QR code containing a benefit coupon or link destination information. This allows passengers to receive advertisements tailored to their needs, improving their in-taxi experience.
[1195] "Imaging means" is a hardware or software component for capturing facial images of passengers.
[1196] "Voice collection means" is a hardware or software component for collecting passenger voices.
[1197] "Analysis means" is a general term for hardware or software for analyzing data obtained from the imaging means and audio collection means and determining the gender, age, and purpose of the passenger's ride.
[1198] The "advertising generation means" is a hardware or software component for generating optimal advertisements based on passenger information determined by the analysis means.
[1199] The "display means" refers to a display device or software for visually presenting the advertisements generated by the advertisement generating means to passengers.
[1200] A "QR code generator" is a hardware or software component that generates a QR code to provide a special offer or link related to an advertisement.
[1201] The "communication means" is a hardware or software function for encrypting data acquired by the imaging means and the audio collecting means and transmitting the data to the server.
[1202] "Image Recognition Means" means a hardware or software component that runs a facial recognition algorithm and estimates the gender and age of a passenger.
[1203] "Speech conversion means" refers to hardware or software functionality that converts voice data into text using a voice recognition system and extracts important keywords.
[1204] "Advertising material selection means" means a hardware or software function that uses a generative AI model to select and create optimal advertising content from advertising materials.
[1205] The system of the present invention mainly comprises a terminal installed in a taxi, a server that performs data analysis and advertisement generation, and a network that communicates between these components. The operation of the system proceeds as follows.
[1206] System Configuration
[1207] 1. Terminal
[1208] Imaging means: Includes a high-resolution camera for capturing facial images of passengers. The camera is automatically activated when a passenger gets into the taxi and captures facial images. As a specific example, a Sony IMX586 camera is used.
[1209] Audio collection means: Includes a microphone for collecting passengers' voices. The microphone is activated when passengers board the vehicle and records their conversations. As a specific example, a Rode NT-USB microphone is used.
[1210] Display means: includes a high-resolution touchscreen display for displaying advertising content. As a specific example, a 16-inch high-resolution touchscreen display is used.
[1211] Communication means: Includes a communication module for encrypting data acquired by the imaging means and audio collection means and transmitting it to a server. As a specific example, a 4G LTE module is used.
[1212] 2. Server
[1213] Analysis means: includes a software component for analyzing data obtained from the imaging means and audio collection means and determining the gender, age, and purpose of the passenger.
[1214] Image Recognition: Implements a facial recognition algorithm to estimate gender and age. For example, the OpenCV library is used.
[1215] Speech conversion method: A speech recognition system is used to convert the audio data into text and extract important keywords. Specific examples include DeepSpeech and Amazon Transcribe.
[1216] Ad generation methods:
[1217] Advertising material selection means: Based on passenger information determined by the analysis means, the optimal advertising material is selected and advertising content is created using a generative AI model. As a specific example, GPT-4 is used.
[1218] QR code generation method: Generates a QR code containing special coupons and link information. As a concrete example, the Libqrencode library is used.
[1219] 3. Network
[1220] This includes a network for fast and secure data communication between the terminals installed in the taxi and the server. For example, 4G / 5G networks are used.
[1221] Specific examples
[1222] The following is a specific example of the system's operation.
[1223] Scenario: A man in his 30s gets into a taxi for a business meeting.
[1224] 1. Terminal: When a passenger gets into a taxi, the terminal automatically activates the camera to capture the passenger's face image, and simultaneously activates the microphone to record the passenger's conversation.
[1225] 2. Terminal: The acquired facial image and audio data are encrypted and sent to the server.
[1226] 3. Server: Analyzes the received facial image and estimates the passenger's gender as "male" and age as "30s." The voice data is converted into text using a voice recognition system, and the keyword "business negotiation" is extracted.
[1227] 4. Server: Based on the analysis results, the generative AI model generates optimal business-related advertising content. For example, the prompt sentence should include keywords such as "men in their 30s" and "business negotiations" to generate business-related advertising content.
[1228] 5. Server: Embeds a QR code containing a special coupon or link information into the generated advertising content and sends it to the device.
[1229] 6. Terminal: Display the advertisement on the display inside the taxi and place the QR code in a location that is easy to see.
[1230] 7. User: Scans the QR code displayed on the screen with a smartphone to receive special coupons and link information.
[1231] This system will enable passengers to receive advertisements tailored to their needs, improving their in-cab experience, while also enabling advertisers to target their ads more effectively, maximizing their advertising impact.
[1232] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1233] Program processing
[1234] Step 1:
[1235] Device: When a passenger gets into the taxi, the device automatically activates its camera to capture the passenger's facial image, and simultaneously activates its microphone to begin recording the passenger's conversation.
[1236] Input: Passenger riding behavior
[1237] Data processing: Images are taken with a camera and image data (JPEG format) is generated. Audio is collected with a microphone and audio data (WAV format) is generated.
[1238] Output: Facial image data and audio data
[1239] Step 2:
[1240] Terminal: The acquired facial image data and voice data are end-to-end encrypted using SSL and sent to the server.
[1241] Input: Facial image data and audio data
[1242] Data processing: SSL encryption is applied and data communication is converted into secure data.
[1243] Output: Encrypted facial image data and audio data
[1244] Step 3:
[1245] Server: Based on the received facial image data, a facial recognition algorithm (OpenCV library) is executed to estimate the passenger's gender and age. The results are stored in a database.
[1246] Input: Encrypted facial image data
[1247] Data processing: The data is decrypted and facial recognition algorithms are used to generate gender and age data.
[1248] Output: Gender and age information
[1249] Step 4:
[1250] Server: The voice data is converted into text using a speech recognition system (DeepSpeech or Amazon Transcribe), and the purpose of the ride is analyzed using a natural language processing (NLP) algorithm. The analysis results are stored in a database.
[1251] Input: Encrypted audio data
[1252] Data processing: The data is decoded and generated into text data using a speech recognition system, after which keywords are extracted using NLP algorithms.
[1253] Output: Keywords for trip purpose
[1254] Step 5:
[1255] Server: Based on the analysis results (gender, age, purpose of ride), a generative AI model (GPT-4) is used to generate the optimal advertisement. Examples of prompt sentences include keywords such as "male in his 30s" and "business negotiation."
[1256] Input: Gender, age, purpose of ride keywords
[1257] Data processing: Input prompt text into the generative AI model to generate advertising content.
[1258] Output: Ad content data
[1259] Step 6:
[1260] Server: Generates a QR code containing a special coupon or link information in the generated advertising content and embeds it in the advertising content.
[1261] Input: Ad content data
[1262] Data processing: Use a QR code generation library to generate a QR code containing special coupon information and link information. This can then be incorporated into advertising content.
[1263] Output: QR code embedded advertising content data
[1264] Step 7:
[1265] Server: Sends advertising content data with an embedded QR code to the terminal.
[1266] Input: QR code embedded advertising content data
[1267] Data processing: The advertising content data is encrypted and prepared for transmission.
[1268] Output: Encrypted advertising content data
[1269] Step 8:
[1270] Terminal: Display the received advertising content data on the display and place the QR code in an easily visible location.
[1271] Input: Encrypted ad content data
[1272] Data processing: Decrypting the data and optimizing it for display.
[1273] Output: Advertising content and QR code displayed on the display
[1274] Step 9:
[1275] User: Scans the QR code displayed on the screen with a smartphone to receive special coupons and link information.
[1276] Input: QR code displayed on the screen
[1277] Data processing: Scan the QR code with your smartphone to obtain special coupons and link information.
[1278] Output: Bonus coupons and link information saved on your smartphone
[1279] This ensures smooth operation of the entire system, improves passengers' in-cab experience by allowing them to receive advertisements tailored to their needs, and enables advertisers to effectively target their advertisements, maximizing their advertising effectiveness.
[1280] (Application example 1)
[1281] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1282] Conventional in-taxi advertising systems have the challenge of making it difficult to provide advertisements tailored to passengers' interests and needs in real time. Furthermore, because advertisements are displayed on fixed monitors, there is no guarantee that passengers will see them. In such situations, advertising effectiveness is not fully realized, resulting in an unsatisfactory experience for both advertisers and passengers.
[1283] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1284] In this invention, the server includes an imaging means for capturing facial images of passengers, an audio collecting means for collecting passenger voices, an analysis means for analyzing data obtained from the imaging means and the audio collecting means and determining the gender, age, and purpose of the passenger's ride, an advertisement generating means for generating an optimal advertisement based on the passenger information determined by the analysis means, a display means for displaying the advertisement generated by the advertisement generating means, a QR code generating means for providing coupons and links related to the advertisement, and a display means for identifying that the display means is a smart glasses display. This allows advertisements tailored to individual passenger needs to be provided in real time and the advertisements to be displayed in a location that is easily visible through the smart glasses.
[1285] "Imaging means" refers to a function or device used to capture facial images of passengers.
[1286] "Voice collection means" refers to the functionality or device used to collect passenger voices.
[1287] The "analysis means" is a function or device that analyzes the data obtained from the imaging means and the audio collection means and determines the gender, age, and purpose of the passenger.
[1288] The "advertisement generation means" is a function or device that generates optimal advertisements based on passenger information determined by the analysis means.
[1289] The "display means" refers to a function or device that displays the advertisement generated by the advertisement generation means, and includes the display of the smart glasses.
[1290] The term "QR code generating means" refers to a function or device that generates a QR code for providing a coupon or link related to an advertisement.
[1291] System Configuration
[1292] An embodiment of the present invention includes the following major components:
[1293] 1. Terminal: Equipped with an imaging means (camera), an audio collection means (microphone), and a display means (display) for displaying advertisements, all of which are installed in the smart glasses.
[1294] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means for providing coupons and links.
[1295] 3. Network: Data communication between the terminal and the server.
[1296] System Operation
[1297] Collecting passenger information
[1298] When a user puts on the smart glasses, the camera automatically captures facial images and the microphone starts recording the user's conversation, and these data are sent to the server in real time.
[1299] Data analysis
[1300] The server analyzes the received image and voice data, using a facial recognition algorithm to estimate the user's gender and age, and a voice recognition system to analyze the user's intent based on the voice data.
[1301] Ad Generation
[1302] Based on the analysis results, the server generates the most suitable advertisement, which includes text information and graphics tailored to the user's interests and goals, as well as a QR code containing special coupons and related links.
[1303] Advertisement display and coupon offer
[1304] The generated advertising data is displayed on the smart glasses display, and users can receive special coupons and link information by scanning the QR code.
[1305] Specific program processing explanation
[1306] The server uses a facial recognition library (e.g., OpenCV) and a voice recognition library (e.g., SpeechRecognition) to analyze the user's facial image and voice data. This allows the server to extract the user's gender, age, and purpose, and then uses ad generation AI to create optimal ads. It also generates QR codes containing coupons and link information, and incorporates them into the ad content.
[1307] Specific examples
[1308] Scenario: A woman in her twenties puts on smart glasses while out shopping.
[1309] 1. The camera in the smart glasses captures the woman's face, and the microphone collects the audio of her saying "shopping."
[1310] 2. The server uses facial recognition to determine that the user is a woman in her 20s, and uses voice analysis to detect the keyword "shopping."
[1311] 3. Generate ads for the latest fashion and cosmetics for women and embed QR codes containing special coupons and shop links into the ads.
[1312] 4. Advertisements will be displayed on the smart glasses display, and users can scan the QR code to receive reward coupons.
[1313] Prompt Sentence Examples
[1314] Give your generative AI model the following inputs:
[1315] User information: Female in her 20s
[1316] Keywords: Shopping
[1317] Generated advertising content: Latest fashion, cosmetics advertising
[1318] QR code content: Bonus coupon, shop link QR code
[1319] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1320] Step 1:
[1321] The device detects that the user has put on the smart glasses. At this time, the camera is activated and an image of the user's face is captured. The captured image is saved in a data format (e.g., JPEG format) and sent to the server.
[1322] Input: User's face image data
[1323] Output: JPEG format face image file
[1324] Step 2:
[1325] The device collects the user's voice using the microphone in the smart glasses. The collected voice data is saved as an audio file (e.g., WAV format) and sent to the server.
[1326] Input: User's voice data
[1327] Output: WAV format audio file
[1328] Step 3:
[1329] The server receives the facial image data and uses a facial recognition algorithm (e.g., OpenCV) to estimate the user's gender and age. During this process, image processing operations are performed to obtain the estimated gender and age.
[1330] Input: JPEG format face image file
[1331] Output: User's gender and age (e.g., female in her 20s)
[1332] Step 4:
[1333] The server receives the voice data and converts the voice data into text data using a voice recognition system (e.g., SpeechRecognition). Keywords (e.g., "shopping") are extracted from the converted text data.
[1334] Input: WAV format audio file
[1335] Output: Text data and extracted keywords
[1336] Step 5:
[1337] The server uses ad generation AI to generate optimal ads based on user information (gender, age, keywords) obtained through facial and voice recognition. During this process, ad materials are selected and combined.
[1338] Input: User information (gender, age, keywords)
[1339] Output: Generated advertising data (text information, graphics, QR code)
[1340] Step 6:
[1341] The server transmits the generated advertisement data to the terminal, and the terminal displays the received advertisement data on the display of the smart glasses.
[1342] Input: Generated ad data
[1343] Output: Advertisement displayed on the smart glasses display
[1344] Step 7:
[1345] Users scan the QR code displayed on the smart glasses display with their smartphone to receive reward coupons and shop links.
[1346] Input: QR code displayed on the screen
[1347] Output: Special coupons and shop links acquired by the user's smartphone
[1348] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1349] System Configuration
[1350] The present invention includes the following major components:
[1351] 1. Terminal installed in the taxi: equipped with imaging means, audio collection means, display means, and emotion engine.
[1352] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means.
[1353] 3. Network: Data communication between the terminal and the server.
[1354] System Operation
[1355] Collecting passenger information
[1356] When a passenger gets into a taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image. At the same time, the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time as soon as the passenger gets into the taxi.
[1357] Data analysis
[1358] The server runs a facial recognition algorithm based on the received image data to estimate the passenger's gender and age. It also applies a voice recognition system to the audio data to analyze the passenger's purpose. It also uses an emotion engine to determine the passenger's emotional state based on their facial expressions and voice. The results of these analyses are stored in a database.
[1359] Ad Generation
[1360] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses ad generation AI to create ads for passengers. The ads include text information and graphics tailored to the passenger's gender, age, purpose of the ride, and emotional state. QR codes containing special coupons and related links are also generated and embedded in the advertising content.
[1361] Advertisement display and coupon offer
[1362] The advertising data sent from the server is returned to the terminal and provided to passengers via a display. The advertisement contains a QR code, which passengers can scan to receive special coupons or link information. In addition, real-time emotional feedback from passengers regarding the displayed advertisement is collected again, and the content of the advertisement is dynamically adjusted to provide a more effective advertising experience.
[1363] Specific program processing
[1364] Program processing explanation
[1365] Scenario: A man in his 30s gets into a taxi for a business meeting and feels a little nervous.
[1366] 1. Information collection during the ride:
[1367] The device "uses a camera to capture passengers' faces and a microphone to record the purpose of the ride."
[1368] Device: "Sends the captured image and audio data to the server"
[1369] 2. Data Analysis:
[1370] Server: "From the received image data, a facial recognition algorithm determines the gender as male and the age as in their 30s."
[1371] Server: "The voice recognition system converts the voice data into text and extracts the keyword 'business negotiations'."
[1372] Server: "Using an emotion engine to detect tension levels from passengers' facial expressions and voices."
[1373] 3. Ad generation:
[1374] Server: "Based on the analysis results, select the most suitable business-related advertising materials."
[1375] The server "generates ads containing text and images that soften products and services and incorporates content that reduces tension."
[1376] Server: Generates QR codes containing special coupons and links and embeds them in advertising content.
[1377] 4. Advertising and Coupon Offers:
[1378] Server: "Sends the generated advertising data to the device"
[1379] Device: "The advertisement will be displayed on the tablet screen, and the QR code will be placed in an easily visible position."
[1380] User: "Scan the QR code with your smartphone to get a special coupon."
[1381] The device "recaptures the passenger's facial expression in response to the displayed advertisement and sends emotional feedback to the server."
[1382] This allows passengers to see ads that reflect their real-time emotional state, providing them with an optimal advertising experience, while enabling advertisers to achieve more effective targeting and interactive advertising methods.
[1383] The processing flow will be explained below.
[1384] Program processing steps
[1385] Specific scenario: A man in his 30s is getting into a taxi for a business meeting and is feeling a bit nervous.
[1386] Step 1:
[1387] The moment a passenger gets into the taxi, the camera activates and captures the passenger's facial image.
[1388] The device photographs the passenger's face from multiple angles to obtain clear image data.
[1389] Step 2:
[1390] "At the same time, the microphone collects passenger conversations and records voice data in real time."
[1391] The device records passengers' conversations in high quality and performs pre-processing to remove noise.
[1392] Step 3:
[1393] The device "encodes the collected image and audio data and transmits it to a server using a secure communication protocol."
[1394] The terminal encrypts the data to ensure secure communication.
[1395] Step 4:
[1396] Server: Based on the received image data, it runs a facial recognition algorithm and determines the passenger's gender as male and age as in their 30s.
[1397] The server uses a deep learning model to estimate the gender and age of passengers from facial images.
[1398] Step 5:
[1399] Server: "Apply a voice recognition system to the voice data and analyze the purpose of the ride."
[1400] The server uses a speech recognition API to convert the speech into text data and extract the keyword "business negotiation."
[1401] Step 6:
[1402] Server: "Uses an emotion engine to determine the passenger's emotional state from their facial expressions and voice."
[1403] The server performs facial expression analysis and voice tone analysis to detect the passenger's state of tension.
[1404] Step 7:
[1405] Server: "Save the analysis results in the database"
[1406] The server stores passengers' gender, age, purpose of the ride, and emotional state data in a database.
[1407] Step 8:
[1408] Server: "Based on the analysis results, we will use ad generation AI to create ads for passengers."
[1409] The server selects advertising materials that are business-related and de-escalating, and generates customized advertisements.
[1410] Step 9:
[1411] Server: Generates QR codes containing special coupons and links and embeds them in advertising content.
[1412] The server uses a QR code generation algorithm and binds it to the advertisement.
[1413] Step 10:
[1414] Server: "Sends the generated advertising data to the taxi terminal."
[1415] The server encodes the advertisement data and transmits it to the terminal.
[1416] Step 11:
[1417] Terminal: "Decodes the received advertising data and displays it on the tablet display."
[1418] The terminal displays advertising content in real time and places the advertisements in a location that is easy for passengers to see.
[1419] Step 12:
[1420] The terminal "displays the QR code included in the advertisement in front of the passenger."
[1421] Place your device in the center of the screen so that the QR code is easy to scan.
[1422] Step 13:
[1423] User: "Scan the QR code with your smartphone to get a special coupon."
[1424] Users scan the QR code with their smartphone's camera app to obtain coupon and link information.
[1425] Step 14:
[1426] The device "recaptures the passenger's facial expressions in response to the displayed advertisement and obtains emotional feedback."
[1427] The device then takes another photograph of the passenger's face and analyzes their emotional state.
[1428] Step 15:
[1429] The device "sends emotional feedback to the server and readjusts the advertising content."
[1430] The terminal transmits the obtained emotional feedback to the server and dynamically updates the advertisement content.
[1431] This allows passengers to see ads that are tailored to their real-time emotional state, providing a more satisfying advertising experience for passengers and enabling advertisers to deliver more targeted and effective ads.
[1432] Example 2
[1433] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1434] Conventional in-taxi advertising systems do not optimize advertisements based on the passenger's gender, age, and purpose of the ride, nor do they optimize advertisements based on the passenger's current emotional state. This makes it difficult to effectively deliver advertisements that are most appropriate for each passenger. It is also difficult to grasp the effectiveness of advertisements in real time and dynamically adjust them. To solve these problems, the present invention aims to provide optimal advertisements based on detailed passenger information, including the passenger's emotional state, and to provide feedback on the effectiveness of advertisements in real time.
[1435] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1436] In this invention, the server includes an analysis means for analyzing the received data and determining the gender, age, purpose of riding, and emotional state of the passenger, an advertisement generation means for generating an optimal advertisement based on the passenger information and emotional state determined by the analysis means, and a feedback collection means for acquiring passenger emotional feedback on the displayed advertisement and performing additional analysis. This makes it possible to provide optimal advertisements according to the emotional state of each passenger and provide feedback on the effectiveness of the advertisements in real time.
[1437] The "imaging means" is a device for capturing facial images of passengers.
[1438] "Voice collection means" refers to equipment for collecting passenger voices.
[1439] The "communication means" is a function for transmitting data obtained from the imaging means and the sound collecting means to a server in real time.
[1440] "Analysis means" refers to a device or program that analyzes the received data and determines the passenger's gender, age, purpose of the ride, and emotional state.
[1441] The "advertisement generation means" is a device or program for generating an optimal advertisement based on the passenger information and emotional state determined by the analysis means.
[1442] The "display means" is a device for displaying the advertisement generated by the advertisement generation means.
[1443] The "QR code generating means" is a function for generating a QR code for providing a coupon or link related to the advertisement.
[1444] The "feedback collection means" is a function for obtaining passengers' emotional feedback regarding the displayed advertisements and for performing additional analysis.
[1445] The "image analysis means" is a device or program for detecting the age and gender of passengers using image data acquired from the imaging means.
[1446] The "voice analysis means" is a device or program for identifying the purpose of riding and the emotional state of the passenger using the voice data obtained from the voice collection means.
[1447] The "generative AI model" is an artificial intelligence model that selects advertising materials and generates advertising content based on the passenger's gender, age, purpose of riding, and emotional state.
[1448] System Configuration
[1449] The present invention includes the following major components:
[1450] 1. Terminal installed in the taxi: equipped with imaging means, audio collection means, display means, and emotion engine.
[1451] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means.
[1452] 3. Network: Data communication between the terminal and the server.
[1453] Hardware and software used
[1454] Hardware:
[1455] Device imaging means (camera): High-resolution camera
[1456] Audio collection means (microphone): High-sensitivity microphone
[1457] Display means (display): Tablet display
[1458] software:
[1459] Face Recognition Algorithm: Image Processing Library
[1460] Speech Recognition System: Speech Recognition API
[1461] Emotion Engine: Emotion Analysis Tool
[1462] Ad Generation AI: Generative AI Model
[1463] System Operation
[1464] Collecting passenger information
[1465] When a passenger gets into a taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image. At the same time, the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time as soon as the passenger gets into the taxi.
[1466] Data analysis
[1467] The server runs a facial recognition algorithm based on the received image data to estimate the passenger's gender and age. It also applies a voice recognition system to the audio data to analyze the passenger's purpose. It also uses an emotion engine to determine the passenger's emotional state based on their facial expressions and voice. The results of these analyses are stored in a database.
[1468] Ad Generation
[1469] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses a generative AI model to create advertisements for passengers. The advertisements include text information and graphics tailored to the passenger's gender, age, purpose of the ride, and emotional state. QR codes containing special coupons and related links are also generated and embedded in the advertisement content.
[1470] Advertisement display and coupon offer
[1471] The advertising data sent from the server is returned to the terminal and provided to passengers via a display. The advertisement contains a QR code, which passengers can scan to receive special coupons or link information. In addition, real-time emotional feedback from passengers regarding the displayed advertisement is collected again, and the content of the advertisement is dynamically adjusted to provide a more effective advertising experience.
[1472] Specific examples
[1473] Scenario: A man in his 30s gets into a taxi for a business meeting and feels a little nervous.
[1474] 1. Information collection during the ride:
[1475] The device "uses a camera to capture passengers' faces and a microphone to record the purpose of the ride."
[1476] Device: "Sends the captured image and audio data to the server"
[1477] 2. Data Analysis:
[1478] Server: "From the received image data, a facial recognition algorithm determines the gender as male and the age as in their 30s."
[1479] Server: "The voice recognition system converts the voice data into text and extracts the keyword 'business negotiations'."
[1480] Server: "Using an emotion engine to detect tension levels from passengers' facial expressions and voices."
[1481] 3. Ad generation:
[1482] Server: "Based on the analysis results, select the most suitable business-related advertising materials."
[1483] The server "generates ads that include text and images that soften the product's features and characteristics, incorporating tension-reducing content."
[1484] Server: Generates QR codes containing special coupons and links and embeds them in advertising content.
[1485] 4. Advertising and Coupon Offers:
[1486] Server: "Sends the generated advertising data to the device"
[1487] Terminal: "Display advertisements on the display and place QR codes in an easily visible location."
[1488] User: "Scan the QR code with your smartphone to get a special coupon."
[1489] The device "recaptures the passenger's facial expression in response to the displayed advertisement and sends emotional feedback to the server."
[1490] Examples of prompts:
[1491] Prompt statement:
[1492] Please generate an ad based on the following criteria:
[1493] Gender: Male
[1494] Age: 30s
[1495] Purpose of the ride: Business negotiations
[1496] Emotional state: Tension
[1497] Ads should include text and images with tension-busting content, as well as a QR code with a special coupon or relevant link."
[1498] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1499] Step 1:
[1500] Gathering information when riding
[1501] Specific operation: The device automatically activates the camera (photography means) and microphone (audio collection means). The camera captures the passenger's facial image, and the microphone records the passenger's conversation.
[1502] Input: Passenger face image, passenger voice data
[1503] Data processing / calculation: capturing facial images, collecting voice data
[1504] Output: Acquired facial image data and audio data
[1505] Step 2:
[1506] Sending data
[1507] Specific operation: The device sends the collected facial image data and voice data to the server in real time. Data communication is performed via the network.
[1508] Input: Facial image data, audio data
[1509] Data processing / calculation: Data encryption and transmission
[1510] Output: Facial image data and audio data sent to the server
[1511] Step 3:
[1512] Data analysis
[1513] What it does: The server analyzes the data it receives. It uses a facial recognition algorithm to estimate gender and age, a voice recognition system to convert the audio data into text, and an emotion engine to determine the passenger's emotional state.
[1514] Input: Received facial image data and audio data
[1515] Data processing / calculation: facial recognition, voice recognition, emotion analysis
[1516] Output: Gender, age, purpose of ride, emotional state
[1517] Step 4:
[1518] Ad Generation
[1519] What it does: The server generates the best ads based on the analysis results, uses generative AI models to create text information and graphics, and creates QR codes with special offers and links.
[1520] Input: Gender, age, purpose of ride, emotional state
[1521] Data processing / calculation: Generating advertising content, generating QR codes
[1522] Output: Advertisement data (text, images, QR code, etc.)
[1523] Step 5:
[1524] Advertisement display and coupon offer
[1525] Specific operation: The advertising data sent from the server is returned to the device and displayed on the screen. The user can then scan the QR code with their smartphone to receive a special coupon.
[1526] Input: Ad data
[1527] Data processing / calculation: Display of advertising data
[1528] Output: Displayed ad content, user-acquired reward coupon
[1529] Step 6:
[1530] Gathering feedback
[1531] How it works: The device recaptures passenger responses after the ad is displayed and sends the data to the server, which performs additional analysis and dynamically adjusts the ad content.
[1532] Input: Recollected facial image data, emotion feedback
[1533] Data processing / calculation: Feedback analysis, dynamic adjustment of advertising content
[1534] Output: Adjusted advertising data
[1535] (Application example 2)
[1536] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1537] Conventional systems have been able to display advertisements based on passengers' gender, age, and purpose of riding, but no systems have taken into account passengers' emotional state or security risks. This makes it difficult to respond quickly when a passenger exhibits suspicious behavior or is considered a dangerous person. It is also not possible to dynamically change advertisements based on the passenger's emotional state. Therefore, the present invention aims to solve these problems and provide a system that evaluates security risks while taking passengers' emotional state into account, and displays appropriate advertisements and generates warnings.
[1538] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an emotion engine that determines the emotional state of passengers, a warning generation means that evaluates security risks based on the emotional state and generates a warning, and a database comparison means that compares the emotional state with a dangerous person database. This makes it possible to analyze the emotional state of passengers in real time, detect dangerous people early, and send warnings to the driver or security services in real time as necessary.
[1539] "Imaging means" means a camera or other image capture device for capturing facial images of passengers.
[1540] "Audio collection means" means a microphone or other audio capture device for collecting passenger audio.
[1541] The "analysis means" refers to an algorithm or system that analyzes the data obtained from the imaging means and audio collection means and determines the gender, age, and purpose of the passenger.
[1542] The "advertising generation means" is a program or system for generating optimal advertisements based on passenger information determined by the analysis means.
[1543] The "display means" refers to a display or screen for displaying the advertisements generated by the advertisement generating means to passengers.
[1544] A "QR code generator" is a program or system for generating a QR code to provide a coupon or link related to an advertisement.
[1545] The "emotion engine" is an algorithm or system for determining the emotional state of passengers based on image data and audio data acquired from imaging means.
[1546] An "alert generator" is a program or system for assessing security risks based on emotional states and generating alerts.
[1547] The "database comparison means" is a program or system that compares the information with a database of dangerous people and sends a warning in real time if there is a match.
[1548] The present invention provides a system that evaluates security risks while taking into account the emotional state of passengers, and displays appropriate advertisements and generates warnings. Specific embodiments of the system are described below.
[1549] System Configuration
[1550] The system of the present invention includes the following major components:
[1551] 1. Device:
[1552] Imaging means: A camera installed inside the taxi captures facial images of passengers.
[1553] Voice collection method: Passenger voices are collected using microphones installed inside the taxi.
[1554] Display means: A display for showing advertising or warning messages.
[1555] Emotion engine: Determines the emotional state of passengers based on acquired image and audio data.
[1556] 2. Server:
[1557] Analysis method: Analyzes image and audio data sent from the terminal to determine the passenger's gender, age, and purpose of the ride.
[1558] Advertisement generation means: Generates an optimal advertisement based on the determined information.
[1559] QR code generator: Generates a QR code to provide coupons or links related to the advertisement.
[1560] Warning generation means: Evaluates security risks based on emotional states and generates warnings.
[1561] Database matching means: Matches with a database of dangerous people and sends a real-time alert if there is a match.
[1562] System Operation
[1563] 1. Collecting passenger information:
[1564] When a passenger gets into the taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image, and the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time.
[1565] 2. Data Analysis:
[1566] The server runs a facial recognition algorithm based on the received image data to estimate the passenger's gender and age. It also applies a voice recognition system to the audio data to analyze the passenger's purpose. It also uses an emotion engine to determine the passenger's emotional state based on their facial expressions and voice. If the passenger's emotional state is determined to be dangerous, the warning generation means generates a warning in real time and, if necessary, automatically notifies security services.
[1567] 3. Ad generation:
[1568] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses ad generation AI to create ads for passengers. The ads include text information and graphics tailored to the passenger's gender, age, purpose of the ride, and emotional state. QR codes containing special coupons and related links are also generated and embedded in the advertising content.
[1569] 4. Advertising and Coupon Offers:
[1570] The advertisement data sent from the server is returned to the terminal and provided to passengers through a display. The advertisement contains a QR code, which passengers can scan to receive special coupons and link information.
[1571] Examples of specific examples and prompts
[1572] For example, if a man in his 30s gets into a taxi for a business meeting and feels a little nervous, the system will operate as follows.
[1573] Examples:
[1574] As passengers board, a camera captures their faces and a microphone records the purpose of their trip.
[1575] The server determines the gender as male and the age as in their 30s from the image data, and extracts the keyword "business negotiation" from the voice data.
[1576] The emotion engine detects tension from passengers' facial expressions and voice.
[1577] The server selects the best advertising materials for the business and generates an advertisement containing text and images that will calm the nerves.
[1578] Generate QR codes containing special coupons and links and embed them in advertising content.
[1579] The generated advertisement data is transmitted to the terminal, and the advertisement is displayed on the display.
[1580] Passengers scan the QR code to get reward coupons.
[1581] The passenger's facial expression in response to the displayed advertisement is again captured and the emotional feedback is sent to the server.
[1582] Example prompt sentence:
[1583] "I want to detect passengers exhibiting abnormal behavior and send warnings. In particular, I want real-time alerts to be sent when the emotion is 'angry' or 'fear'."
[1584] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1585] Processing Steps
[1586] Step 1: Acquire passenger facial images and voice
[1587] Input: Passenger boarding the taxi
[1588] Specific behavior:
[1589] The terminal automatically activates the imaging means (camera) and captures the passenger's facial image.
[1590] At the same time, the audio collection means (microphone) starts to collect passengers' conversations.
[1591] Output: Acquired facial image data and audio data
[1592] Step 2: Sending face image and voice data
[1593] Input: Acquired facial image data and voice data
[1594] Specific behavior:
[1595] The terminal transmits the face image data and the voice data to the server.
[1596] Output: Facial image data and voice data sent to the server
[1597] Step 3: Analyze facial images and audio data
[1598] Input: Facial image data and voice data sent to the server
[1599] Specific behavior:
[1600] The server runs a facial recognition algorithm based on the received facial image data to estimate the passenger's gender and age.
[1601] The server applies a voice recognition system to the received voice data and analyzes the purpose of the ride.
[1602] The server uses an emotion engine to determine the passenger's emotional state from their facial expressions and voice.
[1603] Output: Analysis results of gender, age, purpose of ride, emotional state
[1604] Step 4: Assess security risks
[1605] Input: Gender, age, purpose of ride, emotional state analysis results
[1606] Specific behavior:
[1607] The server assesses security risks based on emotional states.
[1608] The server compares the analysis results with a database of dangerous people.
[1609] Output: Security risk assessment results and risky person matching results
[1610] Step 5: Generate Ads
[1611] Input: Analysis results and security risk assessment results
[1612] Specific behavior:
[1613] Based on the analysis results, the server selects the most suitable advertising materials and uses advertising generation AI to create advertisements for passengers.
[1614] The generated advertisements include text information and graphics tailored to the passenger's gender, age, purpose of the trip, and emotional state.
[1615] The server generates a QR code containing a special coupon or related link and embeds it in the advertising content.
[1616] Output: Generated advertising data (text, graphics, QR code)
[1617] Step 6: Generate and send an alert
[1618] Input: Security risk assessment results and risky person matching results
[1619] Specific behavior:
[1620] If the server has a high security risk, a warning generating means generates a warning message.
[1621] The server sends warning messages to designated security services and drivers in real time as needed.
[1622] Output: Generated warning messages and sent warnings
[1623] Step 7: Displaying Advertisements and Warnings
[1624] Input: Generated ad data and warning message
[1625] Specific behavior:
[1626] The terminal receives the generated advertisement data and displays it on the display.
[1627] The device will display warning messages to the driver as needed.
[1628] Output: Advertisements and warnings shown on the display
[1629] Step 8: Collect passenger feedback
[1630] Input: Passenger facial expression while viewing advertisement
[1631] Specific behavior:
[1632] The device then recaptures the passenger's real-time facial expression in response to the displayed advertisement.
[1633] The terminal transmits the acquired feedback data to the server.
[1634] Output: Feedback data
[1635] Step 9: Dynamically adjust ads
[1636] Input: Feedback data
[1637] Specific behavior:
[1638] The server analyzes the feedback data and dynamically adjusts the content of the advertisements.
[1639] Output: Adjusted advertising data
[1640] In this way, the system can assess passengers' emotional state and security risks in real time, providing them with the optimal advertising experience and security.
[1641] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1642] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1643] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1644] [Fourth embodiment]
[1645] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1646] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1647] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1648] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1649] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1650] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1651] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1652] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1653] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1654] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1655] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1656] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1657] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1658] System Configuration
[1659] The system of the present invention includes the following major components:
[1660] 1. Terminal installed in the taxi: Equipped with imaging means, audio collection means, and display means.
[1661] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means.
[1662] 3. Network: Data communication between the terminal and the server.
[1663] System Operation
[1664] Collecting passenger information
[1665] When a passenger gets into a taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image. At the same time, the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time as soon as the passenger gets into the taxi.
[1666] Data analysis
[1667] The server runs a facial recognition algorithm on the received image data to estimate the passenger's gender and age, and applies a voice recognition system to the audio data to analyze the passenger's purpose. The results of these analyses are stored in a database.
[1668] Ad Generation
[1669] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses ad generation AI to create ads for passengers. The ads contain text information and graphics tailored to the passenger's interests and goals. QR codes containing special coupons and related links are also generated and embedded in the advertising content.
[1670] Advertisement display and coupon offer
[1671] The advertisement data sent from the server is returned to the terminal and provided to passengers through a display. The advertisement contains a QR code, which passengers can scan to receive special coupons and link information.
[1672] Specific program processing
[1673] Program processing explanation
[1674] Scenario: A man in his 30s gets into a taxi for a business meeting.
[1675] 1. Information collection during the ride:
[1676] The device "uses a camera to capture passengers' faces and a microphone to record the purpose of the ride."
[1677] Device: "Sends the captured image and audio data to the server"
[1678] 2. Data Analysis:
[1679] Server: "From the received image data, a facial recognition algorithm estimates the gender as male and the age as being in their 30s."
[1680] Server: "The voice recognition system converts the voice data into text and extracts the keyword 'business negotiations'."
[1681] 3. Ad generation:
[1682] Server: "Based on the analysis results, select business-related advertisements from the advertising materials."
[1683] Server: Generates ads that combine business-related videos and text.
[1684] Server: Embed a QR code containing a special coupon or link in the ad.
[1685] 4. Advertising and Coupon Offers:
[1686] Server: "Sends the generated advertising data to the device"
[1687] Device: "The advertisement will be displayed on the tablet screen, and the QR code will be placed in an easily visible position."
[1688] User: "Scan the QR code with your smartphone and receive a special coupon."
[1689] This will improve passengers' in-cab experience by providing tailored advertising, while also enabling advertisers to effectively target their ads and maximize their advertising effectiveness.
[1690] The processing flow will be explained below.
[1691] Program processing steps
[1692] Specific scenario: A man in his 30s gets into a taxi for a business meeting.
[1693] Step 1:
[1694] The device: "The moment a passenger gets into the taxi, the camera activates and captures the passenger's face."
[1695] The device photographs the passenger's face from multiple angles to obtain clear image data.
[1696] Step 2:
[1697] "At the same time, the microphone collects passenger conversations and records voice data in real time."
[1698] The device records the conversation audio in high quality and performs pre-processing to remove noise.
[1699] Step 3:
[1700] Terminal: "Sends collected image data and audio data to the server."
[1701] The terminal encrypts the image and audio data and transmits it to the server using a secure communication protocol.
[1702] Step 4:
[1703] Server: Based on the received image data, it runs a facial recognition algorithm to estimate the passenger's gender and age.
[1704] The server uses a deep learning model to classify the gender as male and the age as in their 30s from the image.
[1705] Step 5:
[1706] Server: "Apply a voice recognition system to the voice data and analyze the purpose of the ride."
[1707] The server uses a speech recognition API to convert the speech into text and extract the keyword "business negotiation."
[1708] Step 6:
[1709] Server: "Save the analysis results in the database"
[1710] The server stores data such as passenger ID, estimated gender, age, and purpose of ride in a database.
[1711] Step 7:
[1712] Server: "Based on the analysis results, we will create optimal ads using ad generation AI."
[1713] The server selects business-related videos and text from advertising materials aimed at businessmen in their 30s.
[1714] Step 8:
[1715] Server: Generate a QR code containing a special coupon or link for the ad and embed it in the ad content.
[1716] The server uses a QR code generation algorithm and integrates it into the advertisement.
[1717] Step 9:
[1718] Server: "Sends the generated advertising data to the taxi terminal."
[1719] The server re-encrypts the advertisement data and transmits it to the terminal.
[1720] Step 10:
[1721] Terminal: "Decodes the received advertising data and displays it on the tablet display."
[1722] The terminal places advertisements in a location that is easy for passengers to see.
[1723] Step 11:
[1724] The terminal "displays the QR code included in the advertisement in front of the passenger."
[1725] Place your device in the center of the screen so that the QR code is easy to scan.
[1726] Step 12:
[1727] User: "Scan the QR code with your smartphone to get a special coupon."
[1728] Users can scan the QR code using their smartphone's camera app to receive the coupon information.
[1729] The above are the specific processing steps of the program of the present invention. In this way, the processing at each step works together to provide passengers with optimal targeted advertising and interactive experiences.
[1730] Example 1
[1731] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1732] Conventional taxi ride systems lacked a method for providing appropriate and effective advertisements to passengers. Furthermore, technology for generating and displaying advertisements tailored to passengers' interests in real time was underdeveloped. As a result, the effectiveness of advertisements was low, and passenger convenience was difficult to improve. The present invention aims to solve these problems and provide more effective and personalized advertisements to passengers.
[1733] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1734] In this invention, the server includes an imaging means for capturing facial images of passengers, an audio collection means for capturing passenger voices, an analysis means for analyzing data obtained from the imaging means and the audio collection means to determine the passenger's gender, age, and purpose of the ride, an advertisement generation means for generating an optimal advertisement based on the passenger information determined by the analysis means, a display means for displaying the advertisement generated by the advertisement generation means, a QR code generation means for providing benefits and links related to the advertisement, a communication means for encrypting data acquired by the imaging means and the audio collection means and transmitting the encrypted data to the server, an image recognition means for executing the facial recognition algorithm to estimate the passenger's gender and age, a speech conversion means for converting speech data into text using the speech recognition system and extracting important keywords, an advertising material selection means for creating optimal advertising content from advertising materials using a generative AI model, and a QR code generation means for generating a QR code containing a benefit coupon or link destination information. This allows passengers to receive advertisements tailored to their needs, improving their in-taxi experience.
[1735] "Imaging means" is a hardware or software component for capturing facial images of passengers.
[1736] "Voice collection means" is a hardware or software component for collecting passenger voices.
[1737] "Analysis means" is a general term for hardware or software for analyzing data obtained from the imaging means and audio collection means and determining the gender, age, and purpose of the passenger's ride.
[1738] The "advertising generation means" is a hardware or software component for generating optimal advertisements based on passenger information determined by the analysis means.
[1739] The "display means" refers to a display device or software for visually presenting the advertisements generated by the advertisement generating means to passengers.
[1740] A "QR code generator" is a hardware or software component that generates a QR code to provide a special offer or link related to an advertisement.
[1741] The "communication means" is a hardware or software function for encrypting data acquired by the imaging means and the audio collecting means and transmitting the data to the server.
[1742] "Image Recognition Means" means a hardware or software component that runs a facial recognition algorithm and estimates the gender and age of a passenger.
[1743] "Speech conversion means" refers to hardware or software functionality that converts voice data into text using a voice recognition system and extracts important keywords.
[1744] "Advertising material selection means" means a hardware or software function that uses a generative AI model to select and create optimal advertising content from advertising materials.
[1745] The system of the present invention mainly comprises a terminal installed in a taxi, a server that performs data analysis and advertisement generation, and a network that communicates between these components. The operation of the system proceeds as follows.
[1746] System Configuration
[1747] 1. Terminal
[1748] Imaging means: Includes a high-resolution camera for capturing facial images of passengers. The camera is automatically activated when a passenger gets into the taxi and captures facial images. As a specific example, a Sony IMX586 camera is used.
[1749] Audio collection means: Includes a microphone for collecting passengers' voices. The microphone is activated when passengers board the vehicle and records their conversations. As a specific example, a Rode NT-USB microphone is used.
[1750] Display means: includes a high-resolution touchscreen display for displaying advertising content. As a specific example, a 16-inch high-resolution touchscreen display is used.
[1751] Communication means: Includes a communication module for encrypting data acquired by the imaging means and audio collection means and transmitting it to a server. As a specific example, a 4G LTE module is used.
[1752] 2. Server
[1753] Analysis means: includes a software component for analyzing data obtained from the imaging means and audio collection means and determining the gender, age, and purpose of the passenger.
[1754] Image Recognition: Implements a facial recognition algorithm to estimate gender and age. For example, the OpenCV library is used.
[1755] Speech conversion method: A speech recognition system is used to convert the audio data into text and extract important keywords. Specific examples include DeepSpeech and Amazon Transcribe.
[1756] Ad generation methods:
[1757] Advertising material selection means: Based on passenger information determined by the analysis means, the optimal advertising material is selected and advertising content is created using a generative AI model. As a specific example, GPT-4 is used.
[1758] QR code generation method: Generates a QR code containing special coupons and link information. As a concrete example, the Libqrencode library is used.
[1759] 3. Network
[1760] This includes a network for fast and secure data communication between the terminals installed in the taxi and the server. For example, 4G / 5G networks are used.
[1761] Specific examples
[1762] The following is a specific example of the system's operation.
[1763] Scenario: A man in his 30s gets into a taxi for a business meeting.
[1764] 1. Terminal: When a passenger gets into a taxi, the terminal automatically activates the camera to capture the passenger's face image, and simultaneously activates the microphone to record the passenger's conversation.
[1765] 2. Terminal: The acquired facial image and audio data are encrypted and sent to the server.
[1766] 3. Server: Analyzes the received facial image and estimates the passenger's gender as "male" and age as "30s." The voice data is converted into text using a voice recognition system, and the keyword "business negotiation" is extracted.
[1767] 4. Server: Based on the analysis results, the generative AI model generates optimal business-related advertising content. For example, the prompt sentence should include keywords such as "men in their 30s" and "business negotiations" to generate business-related advertising content.
[1768] 5. Server: Embeds a QR code containing a special coupon or link information into the generated advertising content and sends it to the device.
[1769] 6. Terminal: Display the advertisement on the display inside the taxi and place the QR code in a location that is easy to see.
[1770] 7. User: Scans the QR code displayed on the screen with a smartphone to receive special coupons and link information.
[1771] This system will enable passengers to receive advertisements tailored to their needs, improving their in-cab experience, while also enabling advertisers to target their ads more effectively, maximizing their advertising impact.
[1772] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1773] Program processing
[1774] Step 1:
[1775] Device: When a passenger gets into the taxi, the device automatically activates its camera to capture the passenger's facial image, and simultaneously activates its microphone to begin recording the passenger's conversation.
[1776] Input: Passenger riding behavior
[1777] Data processing: Images are taken with a camera and image data (JPEG format) is generated. Audio is collected with a microphone and audio data (WAV format) is generated.
[1778] Output: Facial image data and audio data
[1779] Step 2:
[1780] Terminal: The acquired facial image data and voice data are end-to-end encrypted using SSL and sent to the server.
[1781] Input: Facial image data and audio data
[1782] Data processing: SSL encryption is applied and data communication is converted into secure data.
[1783] Output: Encrypted facial image data and audio data
[1784] Step 3:
[1785] Server: Based on the received facial image data, a facial recognition algorithm (OpenCV library) is executed to estimate the passenger's gender and age. The results are stored in a database.
[1786] Input: Encrypted facial image data
[1787] Data processing: The data is decrypted and facial recognition algorithms are used to generate gender and age data.
[1788] Output: Gender and age information
[1789] Step 4:
[1790] Server: The voice data is converted into text using a speech recognition system (DeepSpeech or Amazon Transcribe), and the purpose of the ride is analyzed using a natural language processing (NLP) algorithm. The analysis results are stored in a database.
[1791] Input: Encrypted audio data
[1792] Data processing: The data is decoded and generated into text data using a speech recognition system, after which keywords are extracted using NLP algorithms.
[1793] Output: Keywords for trip purpose
[1794] Step 5:
[1795] Server: Based on the analysis results (gender, age, purpose of ride), a generative AI model (GPT-4) is used to generate the optimal advertisement. Examples of prompt sentences include keywords such as "male in his 30s" and "business negotiation."
[1796] Input: Gender, age, purpose of ride keywords
[1797] Data processing: Input prompt text into the generative AI model to generate advertising content.
[1798] Output: Ad content data
[1799] Step 6:
[1800] Server: Generates a QR code containing a special coupon or link information in the generated advertising content and embeds it in the advertising content.
[1801] Input: Ad content data
[1802] Data processing: Use a QR code generation library to generate a QR code containing special coupon information and link information. This can then be incorporated into advertising content.
[1803] Output: QR code embedded advertising content data
[1804] Step 7:
[1805] Server: Sends advertising content data with an embedded QR code to the terminal.
[1806] Input: QR code embedded advertising content data
[1807] Data processing: The advertising content data is encrypted and prepared for transmission.
[1808] Output: Encrypted advertising content data
[1809] Step 8:
[1810] Terminal: Display the received advertising content data on the display and place the QR code in an easily visible location.
[1811] Input: Encrypted ad content data
[1812] Data processing: Decrypting the data and optimizing it for display.
[1813] Output: Advertising content and QR code displayed on the display
[1814] Step 9:
[1815] User: Scans the QR code displayed on the screen with a smartphone to receive special coupons and link information.
[1816] Input: QR code displayed on the screen
[1817] Data processing: Scan the QR code with your smartphone to obtain special coupons and link information.
[1818] Output: Bonus coupons and link information saved on your smartphone
[1819] This ensures smooth operation of the entire system, improves passengers' in-cab experience by allowing them to receive advertisements tailored to their needs, and enables advertisers to effectively target their advertisements, maximizing their advertising effectiveness.
[1820] (Application example 1)
[1821] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1822] Conventional in-taxi advertising systems have the challenge of making it difficult to provide advertisements tailored to passengers' interests and needs in real time. Furthermore, because advertisements are displayed on fixed monitors, there is no guarantee that passengers will see them. In such situations, advertising effectiveness is not fully realized, resulting in an unsatisfactory experience for both advertisers and passengers.
[1823] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1824] In this invention, the server includes an imaging means for capturing facial images of passengers, an audio collecting means for collecting passenger voices, an analysis means for analyzing data obtained from the imaging means and the audio collecting means and determining the gender, age, and purpose of the passenger's ride, an advertisement generating means for generating an optimal advertisement based on the passenger information determined by the analysis means, a display means for displaying the advertisement generated by the advertisement generating means, a QR code generating means for providing coupons and links related to the advertisement, and a display means for identifying that the display means is a smart glasses display. This allows advertisements tailored to individual passenger needs to be provided in real time and the advertisements to be displayed in a location that is easily visible through the smart glasses.
[1825] "Imaging means" refers to a function or device used to capture facial images of passengers.
[1826] "Voice collection means" refers to the functionality or device used to collect passenger voices.
[1827] The "analysis means" is a function or device that analyzes the data obtained from the imaging means and the audio collection means and determines the gender, age, and purpose of the passenger.
[1828] The "advertisement generation means" is a function or device that generates optimal advertisements based on passenger information determined by the analysis means.
[1829] The "display means" refers to a function or device that displays the advertisement generated by the advertisement generation means, and includes the display of the smart glasses.
[1830] The term "QR code generating means" refers to a function or device that generates a QR code for providing a coupon or link related to an advertisement.
[1831] System Configuration
[1832] An embodiment of the present invention includes the following major components:
[1833] 1. Terminal: Equipped with an imaging means (camera), an audio collection means (microphone), and a display means (display) for displaying advertisements, all of which are installed in the smart glasses.
[1834] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means for providing coupons and links.
[1835] 3. Network: Data communication between the terminal and the server.
[1836] System Operation
[1837] Collecting passenger information
[1838] When a user puts on the smart glasses, the camera automatically captures facial images and the microphone starts recording the user's conversation, and these data are sent to the server in real time.
[1839] Data analysis
[1840] The server analyzes the received image and voice data, using a facial recognition algorithm to estimate the user's gender and age, and a voice recognition system to analyze the user's intent based on the voice data.
[1841] Ad Generation
[1842] Based on the analysis results, the server generates the most suitable advertisement, which includes text information and graphics tailored to the user's interests and goals, as well as a QR code containing special coupons and related links.
[1843] Advertisement display and coupon offer
[1844] The generated advertising data is displayed on the smart glasses display, and users can receive special coupons and link information by scanning the QR code.
[1845] Specific program processing explanation
[1846] The server uses a facial recognition library (e.g., OpenCV) and a voice recognition library (e.g., SpeechRecognition) to analyze the user's facial image and voice data. This allows the server to extract the user's gender, age, and purpose, and then uses ad generation AI to create optimal ads. It also generates QR codes containing coupons and link information, and incorporates them into the ad content.
[1847] Specific examples
[1848] Scenario: A woman in her twenties puts on smart glasses while out shopping.
[1849] 1. The camera in the smart glasses captures the woman's face, and the microphone collects the audio of her saying "shopping."
[1850] 2. The server uses facial recognition to determine that the user is a woman in her 20s, and uses voice analysis to detect the keyword "shopping."
[1851] 3. Generate ads for the latest fashion and cosmetics for women and embed QR codes containing special coupons and shop links into the ads.
[1852] 4. Advertisements will be displayed on the smart glasses display, and users can scan the QR code to receive reward coupons.
[1853] Prompt Sentence Examples
[1854] Give your generative AI model the following inputs:
[1855] User information: Female in her 20s
[1856] Keywords: Shopping
[1857] Generated advertising content: Latest fashion, cosmetics advertising
[1858] QR code content: Bonus coupon, shop link QR code
[1859] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1860] Step 1:
[1861] The device detects that the user has put on the smart glasses. At this time, the camera is activated and an image of the user's face is captured. The captured image is saved in a data format (e.g., JPEG format) and sent to the server.
[1862] Input: User's face image data
[1863] Output: JPEG format face image file
[1864] Step 2:
[1865] The device collects the user's voice using the microphone in the smart glasses. The collected voice data is saved as an audio file (e.g., WAV format) and sent to the server.
[1866] Input: User's voice data
[1867] Output: WAV format audio file
[1868] Step 3:
[1869] The server receives the facial image data and uses a facial recognition algorithm (e.g., OpenCV) to estimate the user's gender and age. During this process, image processing operations are performed to obtain the estimated gender and age.
[1870] Input: JPEG format face image file
[1871] Output: User's gender and age (e.g., female in her 20s)
[1872] Step 4:
[1873] The server receives the voice data and converts the voice data into text data using a voice recognition system (e.g., SpeechRecognition). Keywords (e.g., "shopping") are extracted from the converted text data.
[1874] Input: WAV format audio file
[1875] Output: Text data and extracted keywords
[1876] Step 5:
[1877] The server uses ad generation AI to generate optimal ads based on user information (gender, age, keywords) obtained through facial and voice recognition. During this process, ad materials are selected and combined.
[1878] Input: User information (gender, age, keywords)
[1879] Output: Generated advertising data (text information, graphics, QR code)
[1880] Step 6:
[1881] The server transmits the generated advertisement data to the terminal, and the terminal displays the received advertisement data on the display of the smart glasses.
[1882] Input: Generated ad data
[1883] Output: Advertisement displayed on the smart glasses display
[1884] Step 7:
[1885] Users scan the QR code displayed on the smart glasses display with their smartphone to receive reward coupons and shop links.
[1886] Input: QR code displayed on the screen
[1887] Output: Special coupons and shop links acquired by the user's smartphone
[1888] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1889] System Configuration
[1890] The present invention includes the following major components:
[1891] 1. Terminal installed in the taxi: equipped with imaging means, audio collection means, display means, and emotion engine.
[1892] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means.
[1893] 3. Network: Data communication between the terminal and the server.
[1894] System Operation
[1895] Collecting passenger information
[1896] When a passenger gets into a taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image. At the same time, the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time as soon as the passenger gets into the taxi.
[1897] Data analysis
[1898] The server runs a facial recognition algorithm based on the received image data to estimate the passenger's gender and age. It also applies a voice recognition system to the audio data to analyze the passenger's purpose. It also uses an emotion engine to determine the passenger's emotional state based on their facial expressions and voice. The results of these analyses are stored in a database.
[1899] Ad Generation
[1900] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses ad generation AI to create ads for passengers. The ads include text information and graphics tailored to the passenger's gender, age, purpose of the ride, and emotional state. QR codes containing special coupons and related links are also generated and embedded in the advertising content.
[1901] Advertisement display and coupon offer
[1902] The advertising data sent from the server is returned to the terminal and provided to passengers via a display. The advertisement contains a QR code, which passengers can scan to receive special coupons or link information. In addition, real-time emotional feedback from passengers regarding the displayed advertisement is collected again, and the content of the advertisement is dynamically adjusted to provide a more effective advertising experience.
[1903] Specific program processing
[1904] Program processing explanation
[1905] Scenario: A man in his 30s gets into a taxi for a business meeting and feels a little nervous.
[1906] 1. Information collection during the ride:
[1907] The device "uses a camera to capture passengers' faces and a microphone to record the purpose of the ride."
[1908] Device: "Sends the captured image and audio data to the server"
[1909] 2. Data Analysis:
[1910] Server: "From the received image data, a facial recognition algorithm determines the gender as male and the age as in their 30s."
[1911] Server: "The voice recognition system converts the voice data into text and extracts the keyword 'business negotiations'."
[1912] Server: "Using an emotion engine to detect tension levels from passengers' facial expressions and voices."
[1913] 3. Ad generation:
[1914] Server: "Based on the analysis results, select the most suitable business-related advertising materials."
[1915] The server "generates ads containing text and images that soften products and services and incorporates content that reduces tension."
[1916] Server: Generates QR codes containing special coupons and links and embeds them in advertising content.
[1917] 4. Advertising and Coupon Offers:
[1918] Server: "Sends the generated advertising data to the device"
[1919] Device: "The advertisement will be displayed on the tablet screen, and the QR code will be placed in an easily visible position."
[1920] User: "Scan the QR code with your smartphone to get a special coupon."
[1921] The device "recaptures the passenger's facial expression in response to the displayed advertisement and sends emotional feedback to the server."
[1922] This allows passengers to see ads that reflect their real-time emotional state, providing them with an optimal advertising experience, while enabling advertisers to achieve more effective targeting and interactive advertising methods.
[1923] The processing flow will be explained below.
[1924] Program processing steps
[1925] Specific scenario: A man in his 30s is getting into a taxi for a business meeting and is feeling a bit nervous.
[1926] Step 1:
[1927] The moment a passenger gets into the taxi, the camera activates and captures the passenger's facial image.
[1928] The device photographs the passenger's face from multiple angles to obtain clear image data.
[1929] Step 2:
[1930] "At the same time, the microphone collects passenger conversations and records voice data in real time."
[1931] The device records passengers' conversations in high quality and performs pre-processing to remove noise.
[1932] Step 3:
[1933] The device "encodes the collected image and audio data and transmits it to a server using a secure communication protocol."
[1934] The terminal encrypts the data to ensure secure communication.
[1935] Step 4:
[1936] Server: Based on the received image data, it runs a facial recognition algorithm and determines the passenger's gender as male and age as in their 30s.
[1937] The server uses a deep learning model to estimate the gender and age of passengers from facial images.
[1938] Step 5:
[1939] Server: "Apply a voice recognition system to the voice data and analyze the purpose of the ride."
[1940] The server uses a speech recognition API to convert the speech into text data and extract the keyword "business negotiation."
[1941] Step 6:
[1942] Server: "Uses an emotion engine to determine the passenger's emotional state from their facial expressions and voice."
[1943] The server performs facial expression analysis and voice tone analysis to detect the passenger's state of tension.
[1944] Step 7:
[1945] Server: "Save the analysis results in the database"
[1946] The server stores passengers' gender, age, purpose of the ride, and emotional state data in a database.
[1947] Step 8:
[1948] Server: "Based on the analysis results, we will use ad generation AI to create ads for passengers."
[1949] The server selects advertising materials that are business-related and de-escalating, and generates customized advertisements.
[1950] Step 9:
[1951] Server: Generates QR codes containing special coupons and links and embeds them in advertising content.
[1952] The server uses a QR code generation algorithm and binds it to the advertisement.
[1953] Step 10:
[1954] Server: "Sends the generated advertising data to the taxi terminal."
[1955] The server encodes the advertisement data and transmits it to the terminal.
[1956] Step 11:
[1957] Terminal: "Decodes the received advertising data and displays it on the tablet display."
[1958] The terminal displays advertising content in real time and places the advertisements in a location that is easy for passengers to see.
[1959] Step 12:
[1960] The terminal "displays the QR code included in the advertisement in front of the passenger."
[1961] Place your device in the center of the screen so that the QR code is easy to scan.
[1962] Step 13:
[1963] User: "Scan the QR code with your smartphone to get a special coupon."
[1964] Users scan the QR code with their smartphone's camera app to obtain coupon and link information.
[1965] Step 14:
[1966] The device "recaptures the passenger's facial expressions in response to the displayed advertisement and obtains emotional feedback."
[1967] The device then takes another photograph of the passenger's face and analyzes their emotional state.
[1968] Step 15:
[1969] The device "sends emotional feedback to the server and readjusts the advertising content."
[1970] The terminal transmits the obtained emotional feedback to the server and dynamically updates the advertisement content.
[1971] This allows passengers to see ads that are tailored to their real-time emotional state, providing a more satisfying advertising experience for passengers and enabling advertisers to deliver more targeted and effective ads.
[1972] Example 2
[1973] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1974] Conventional in-taxi advertising systems do not optimize advertisements based on the passenger's gender, age, and purpose of the ride, nor do they optimize advertisements based on the passenger's current emotional state. This makes it difficult to effectively deliver advertisements that are most appropriate for each passenger. It is also difficult to grasp the effectiveness of advertisements in real time and dynamically adjust them. To solve these problems, the present invention aims to provide optimal advertisements based on detailed passenger information, including the passenger's emotional state, and to provide feedback on the effectiveness of advertisements in real time.
[1975] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1976] In this invention, the server includes an analysis means for analyzing the received data and determining the gender, age, purpose of riding, and emotional state of the passenger, an advertisement generation means for generating an optimal advertisement based on the passenger information and emotional state determined by the analysis means, and a feedback collection means for acquiring passenger emotional feedback on the displayed advertisement and performing additional analysis. This makes it possible to provide optimal advertisements according to the emotional state of each passenger and provide feedback on the effectiveness of the advertisements in real time.
[1977] The "imaging means" is a device for capturing facial images of passengers.
[1978] "Voice collection means" refers to equipment for collecting passenger voices.
[1979] The "communication means" is a function for transmitting data obtained from the imaging means and the sound collecting means to a server in real time.
[1980] "Analysis means" refers to a device or program that analyzes the received data and determines the passenger's gender, age, purpose of the ride, and emotional state.
[1981] The "advertisement generation means" is a device or program for generating an optimal advertisement based on the passenger information and emotional state determined by the analysis means.
[1982] The "display means" is a device for displaying the advertisement generated by the advertisement generation means.
[1983] The "QR code generating means" is a function for generating a QR code for providing a coupon or link related to the advertisement.
[1984] The "feedback collection means" is a function for obtaining passengers' emotional feedback regarding the displayed advertisements and for performing additional analysis.
[1985] The "image analysis means" is a device or program for detecting the age and gender of passengers using image data acquired from the imaging means.
[1986] The "voice analysis means" is a device or program for identifying the purpose of riding and the emotional state of the passenger using the voice data obtained from the voice collection means.
[1987] The "generative AI model" is an artificial intelligence model that selects advertising materials and generates advertising content based on the passenger's gender, age, purpose of riding, and emotional state.
[1988] System Configuration
[1989] The present invention includes the following major components:
[1990] 1. Terminal installed in the taxi: equipped with imaging means, audio collection means, display means, and emotion engine.
[1991] 2. Server: Equipped with data analysis means, advertisement generation means, and QR code generation means.
[1992] 3. Network: Data communication between the terminal and the server.
[1993] Hardware and software used
[1994] Hardware:
[1995] Device imaging means (camera): High-resolution camera
[1996] Audio collection means (microphone): High-sensitivity microphone
[1997] Display means (display): Tablet display
[1998] software:
[1999] Face Recognition Algorithm: Image Processing Library
[2000] Speech Recognition System: Speech Recognition API
[2001] Emotion Engine: Emotion Analysis Tool
[2002] Ad Generation AI: Generative AI Model
[2003] System Operation
[2004] Collecting passenger information
[2005] When a passenger gets into a taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image. At the same time, the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time as soon as the passenger gets into the taxi.
[2006] Data analysis
[2007] The server runs a facial recognition algorithm based on the received image data to estimate the passenger's gender and age. It also applies a voice recognition system to the audio data to analyze the passenger's purpose. It also uses an emotion engine to determine the passenger's emotional state based on their facial expressions and voice. The results of these analyses are stored in a database.
[2008] Ad Generation
[2009] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses a generative AI model to create advertisements for passengers. The advertisements include text information and graphics tailored to the passenger's gender, age, purpose of the ride, and emotional state. QR codes containing special coupons and related links are also generated and embedded in the advertisement content.
[2010] Advertisement display and coupon offer
[2011] The advertising data sent from the server is returned to the terminal and provided to passengers via a display. The advertisement contains a QR code, which passengers can scan to receive special coupons or link information. In addition, real-time emotional feedback from passengers regarding the displayed advertisement is collected again, and the content of the advertisement is dynamically adjusted to provide a more effective advertising experience.
[2012] Specific examples
[2013] Scenario: A man in his 30s gets into a taxi for a business meeting and feels a little nervous.
[2014] 1. Information collection during the ride:
[2015] The device "uses a camera to capture passengers' faces and a microphone to record the purpose of the ride."
[2016] Device: "Sends the captured image and audio data to the server"
[2017] 2. Data Analysis:
[2018] Server: "From the received image data, a facial recognition algorithm determines the gender as male and the age as in their 30s."
[2019] Server: "The voice recognition system converts the voice data into text and extracts the keyword 'business negotiations'."
[2020] Server: "Using an emotion engine to detect tension levels from passengers' facial expressions and voices."
[2021] 3. Ad generation:
[2022] Server: "Based on the analysis results, select the most suitable business-related advertising materials."
[2023] The server "generates ads that include text and images that soften the product's features and characteristics, incorporating tension-reducing content."
[2024] Server: Generates QR codes containing special coupons and links and embeds them in advertising content.
[2025] 4. Advertising and Coupon Offers:
[2026] Server: "Sends the generated advertising data to the device"
[2027] Terminal: "Display advertisements on the display and place QR codes in an easily visible location."
[2028] User: "Scan the QR code with your smartphone to get a special coupon."
[2029] The device "recaptures the passenger's facial expression in response to the displayed advertisement and sends emotional feedback to the server."
[2030] Examples of prompts:
[2031] Prompt statement:
[2032] Please generate an ad based on the following criteria:
[2033] Gender: Male
[2034] Age: 30s
[2035] Purpose of the ride: Business negotiations
[2036] Emotional state: Tension
[2037] Ads should include text and images with tension-busting content, as well as a QR code with a special coupon or relevant link."
[2038] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2039] Step 1:
[2040] Gathering information when riding
[2041] Specific operation: The device automatically activates the camera (photography means) and microphone (audio collection means). The camera captures the passenger's facial image, and the microphone records the passenger's conversation.
[2042] Input: Passenger face image, passenger voice data
[2043] Data processing / calculation: capturing facial images, collecting voice data
[2044] Output: Acquired facial image data and audio data
[2045] Step 2:
[2046] Sending data
[2047] Specific operation: The device sends the collected facial image data and voice data to the server in real time. Data communication is performed via the network.
[2048] Input: Facial image data, audio data
[2049] Data processing / calculation: Data encryption and transmission
[2050] Output: Facial image data and audio data sent to the server
[2051] Step 3:
[2052] Data analysis
[2053] What it does: The server analyzes the data it receives. It uses a facial recognition algorithm to estimate gender and age, a voice recognition system to convert the audio data into text, and an emotion engine to determine the passenger's emotional state.
[2054] Input: Received facial image data and audio data
[2055] Data processing / calculation: facial recognition, voice recognition, emotion analysis
[2056] Output: Gender, age, purpose of ride, emotional state
[2057] Step 4:
[2058] Ad Generation
[2059] What it does: The server generates the best ads based on the analysis results, uses generative AI models to create text information and graphics, and creates QR codes with special offers and links.
[2060] Input: Gender, age, purpose of ride, emotional state
[2061] Data processing / calculation: Generating advertising content, generating QR codes
[2062] Output: Advertisement data (text, images, QR code, etc.)
[2063] Step 5:
[2064] Advertisement display and coupon offer
[2065] Specific operation: The advertising data sent from the server is returned to the device and displayed on the screen. The user can then scan the QR code with their smartphone to receive a special coupon.
[2066] Input: Ad data
[2067] Data processing / calculation: Display of advertising data
[2068] Output: Displayed ad content, user-acquired reward coupon
[2069] Step 6:
[2070] Gathering feedback
[2071] How it works: The device recaptures passenger responses after the ad is displayed and sends the data to the server, which performs additional analysis and dynamically adjusts the ad content.
[2072] Input: Recollected facial image data, emotion feedback
[2073] Data processing / calculation: Feedback analysis, dynamic adjustment of advertising content
[2074] Output: Adjusted advertising data
[2075] (Application example 2)
[2076] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2077] Conventional systems have been able to display advertisements based on passengers' gender, age, and purpose of riding, but no systems have taken into account passengers' emotional state or security risks. This makes it difficult to respond quickly when a passenger exhibits suspicious behavior or is considered a dangerous person. It is also not possible to dynamically change advertisements based on the passenger's emotional state. Therefore, the present invention aims to solve these problems and provide a system that evaluates security risks while taking passengers' emotional state into account, and displays appropriate advertisements and generates warnings.
[2078] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes an emotion engine that determines the emotional state of passengers, a warning generation means that evaluates security risks based on the emotional state and generates a warning, and a database comparison means that compares the emotional state with a dangerous person database. This makes it possible to analyze the emotional state of passengers in real time, detect dangerous people early, and send warnings to the driver or security services in real time as necessary.
[2079] "Imaging means" means a camera or other image capture device for capturing facial images of passengers.
[2080] "Audio collection means" means a microphone or other audio capture device for collecting passenger audio.
[2081] The "analysis means" refers to an algorithm or system that analyzes the data obtained from the imaging means and audio collection means and determines the gender, age, and purpose of the passenger.
[2082] The "advertising generation means" is a program or system for generating optimal advertisements based on passenger information determined by the analysis means.
[2083] The "display means" refers to a display or screen for displaying the advertisements generated by the advertisement generating means to passengers.
[2084] A "QR code generator" is a program or system for generating a QR code to provide a coupon or link related to an advertisement.
[2085] The "emotion engine" is an algorithm or system for determining the emotional state of passengers based on image data and audio data acquired from imaging means.
[2086] An "alert generator" is a program or system for assessing security risks based on emotional states and generating alerts.
[2087] The "database comparison means" is a program or system that compares the information with a database of dangerous people and sends a warning in real time if there is a match.
[2088] The present invention provides a system that evaluates security risks while taking into account the emotional state of passengers, and displays appropriate advertisements and generates warnings. Specific embodiments of the system are described below.
[2089] System Configuration
[2090] The system of the present invention includes the following major components:
[2091] 1. Device:
[2092] Imaging means: A camera installed inside the taxi captures facial images of passengers.
[2093] Voice collection method: Passenger voices are collected using microphones installed inside the taxi.
[2094] Display means: A display for showing advertising or warning messages.
[2095] Emotion engine: Determines the emotional state of passengers based on acquired image and audio data.
[2096] 2. Server:
[2097] Analysis method: Analyzes image and audio data sent from the terminal to determine the passenger's gender, age, and purpose of the ride.
[2098] Advertisement generation means: Generates an optimal advertisement based on the determined information.
[2099] QR code generator: Generates a QR code to provide coupons or links related to the advertisement.
[2100] Warning generation means: Evaluates security risks based on emotional states and generates warnings.
[2101] Database matching means: Matches with a database of dangerous people and sends a real-time alert if there is a match.
[2102] System Operation
[2103] 1. Collecting passenger information:
[2104] When a passenger gets into the taxi, the device automatically activates the imaging means (camera) to capture the passenger's facial image, and the audio collection means (microphone) starts recording the passenger's conversation. This data is sent to the server in real time.
[2105] 2. Data Analysis:
[2106] The server runs a facial recognition algorithm based on the received image data to estimate the passenger's gender and age. It also applies a voice recognition system to the audio data to analyze the passenger's purpose. It also uses an emotion engine to determine the passenger's emotional state based on their facial expressions and voice. If the passenger's emotional state is determined to be dangerous, the warning generation means generates a warning in real time and, if necessary, automatically notifies security services.
[2107] 3. Ad generation:
[2108] Based on the analysis results, the server selects the most suitable advertising materials from available sources and uses ad generation AI to create ads for passengers. The ads include text information and graphics tailored to the passenger's gender, age, purpose of the ride, and emotional state. QR codes containing special coupons and related links are also generated and embedded in the advertising content.
[2109] 4. Advertising and Coupon Offers:
[2110] The advertisement data sent from the server is returned to the terminal and provided to passengers through a display. The advertisement contains a QR code, which passengers can scan to receive special coupons and link information.
[2111] Examples of specific examples and prompts
[2112] For example, if a man in his 30s gets into a taxi for a business meeting and feels a little nervous, the system will operate as follows.
[2113] Examples:
[2114] As passengers board, a camera captures their faces and a microphone records the purpose of their trip.
[2115] The server determines the gender as male and the age as in their 30s from the image data, and extracts the keyword "business negotiation" from the voice data.
[2116] The emotion engine detects tension from passengers' facial expressions and voice.
[2117] The server selects the best advertising materials for the business and generates an advertisement containing text and images that will calm the nerves.
[2118] Generate QR codes containing special coupons and links and embed them in advertising content.
[2119] The generated advertisement data is transmitted to the terminal, and the advertisement is displayed on the display.
[2120] Passengers scan the QR code to get reward coupons.
[2121] The passenger's facial expression in response to the displayed advertisement is again captured and the emotional feedback is sent to the server.
[2122] Example prompt sentence:
[2123] "I want to detect passengers exhibiting abnormal behavior and send warnings. In particular, I want real-time alerts to be sent when the emotion is 'angry' or 'fear'."
[2124] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2125] Processing Steps
[2126] Step 1: Acquire passenger facial images and voice
[2127] Input: Passenger boarding the taxi
[2128] Specific behavior:
[2129] The terminal automatically activates the imaging means (camera) and captures the passenger's facial image.
[2130] At the same time, the audio collection means (microphone) starts to collect passengers' conversations.
[2131] Output: Acquired facial image data and audio data
[2132] Step 2: Sending face image and voice data
[2133] Input: Acquired facial image data and voice data
[2134] Specific behavior:
[2135] The terminal transmits the face image data and the voice data to the server.
[2136] Output: Facial image data and voice data sent to the server
[2137] Step 3: Analyze facial images and audio data
[2138] Input: Facial image data and voice data sent to the server
[2139] Specific behavior:
[2140] The server runs a facial recognition algorithm based on the received facial image data to estimate the passenger's gender and age.
[2141] The server applies a voice recognition system to the received voice data and analyzes the purpose of the ride.
[2142] The server uses an emotion engine to determine the passenger's emotional state from their facial expressions and voice.
[2143] Output: Analysis results of gender, age, purpose of ride, emotional state
[2144] Step 4: Assess security risks
[2145] Input: Gender, age, purpose of ride, emotional state analysis results
[2146] Specific behavior:
[2147] The server assesses security risks based on emotional states.
[2148] The server compares the analysis results with a database of dangerous people.
[2149] Output: Security risk assessment results and risky person matching results
[2150] Step 5: Generate Ads
[2151] Input: Analysis results and security risk assessment results
[2152] Specific behavior:
[2153] Based on the analysis results, the server selects the most suitable advertising materials and uses advertising generation AI to create advertisements for passengers.
[2154] The generated advertisements include text information and graphics tailored to the passenger's gender, age, purpose of the trip, and emotional state.
[2155] The server generates a QR code containing a special coupon or related link and embeds it in the advertising content.
[2156] Output: Generated advertising data (text, graphics, QR code)
[2157] Step 6: Generate and send an alert
[2158] Input: Security risk assessment results and risky person matching results
[2159] Specific behavior:
[2160] If the server has a high security risk, a warning generating means generates a warning message.
[2161] The server sends warning messages to designated security services and drivers in real time as needed.
[2162] Output: Generated warning messages and sent warnings
[2163] Step 7: Displaying Advertisements and Warnings
[2164] Input: Generated ad data and warning message
[2165] Specific behavior:
[2166] The terminal receives the generated advertisement data and displays it on the display.
[2167] The device will display warning messages to the driver as needed.
[2168] Output: Advertisements and warnings shown on the display
[2169] Step 8: Collect passenger feedback
[2170] Input: Passenger facial expression while viewing advertisement
[2171] Specific behavior:
[2172] The device then recaptures the passenger's real-time facial expression in response to the displayed advertisement.
[2173] The terminal transmits the acquired feedback data to the server.
[2174] Output: Feedback data
[2175] Step 9: Dynamically adjust ads
[2176] Input: Feedback data
[2177] Specific behavior:
[2178] The server analyzes the feedback data and dynamically adjusts the content of the advertisements.
[2179] Output: Adjusted advertising data
[2180] In this way, the system can assess passengers' emotional state and security risks in real time, providing them with the optimal advertising experience and security.
[2181] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2182] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2183] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2184] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2185] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2186] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2187] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2188] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2189] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2190] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2191] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2192] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2193] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2194] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2195] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2196] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2197] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2198] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2199] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2200] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2201] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2202] The following is further disclosed regarding the above embodiment.
[2203] (Claim 1)
[2204] imaging means for acquiring facial images of passengers;
[2205] a voice collection means for collecting voices of passengers;
[2206] an analysis means for analyzing the data obtained from the imaging means and the voice collecting means and determining the gender, age, and purpose of the passenger;
[2207] an advertisement generating means for generating an optimal advertisement based on the passenger information determined by the analyzing means;
[2208] a display means for displaying the advertisement generated by the advertisement generation means;
[2209] a QR code generating means for providing a coupon or link related to the advertisement;
[2210] A system including:
[2211] (Claim 2)
[2212] The system described in claim 1, characterized in that the analysis means includes an image analysis means that detects the age and gender of passengers using image data acquired from the imaging means, and a voice analysis means that identifies the purpose of the ride using voice data acquired from the voice collection means.
[2213] (Claim 3)
[2214] 2. The system according to claim 1, wherein the advertisement generating means comprises a generating means for selecting advertisement material according to the gender, age, and purpose of riding of a passenger, and generating advertisement content.
[2215] "Example 1"
[2216] (Claim 1)
[2217] imaging means for acquiring facial images of passengers;
[2218] a voice collection means for collecting voices of passengers;
[2219] an analysis means for analyzing the data obtained from the imaging means and the voice collecting means and determining the gender, age, and purpose of the passenger;
[2220] an advertisement generating means for generating an optimal advertisement based on the passenger information determined by the analyzing means;
[2221] a display means for displaying the advertisement generated by the advertisement generation means;
[2222] a QR code generating means for providing a reward or link related to the advertisement;
[2223] a communication means for encrypting the data acquired by the imaging means and the sound collecting means and transmitting the encrypted data to a server;
[2224] image recognition means for executing the face recognition algorithm to estimate gender and age;
[2225] a speech conversion means for converting speech data into text using the speech recognition system and extracting important keywords;
[2226] An advertising material selection method that uses a generative AI model to create optimal advertising content from advertising materials;
[2227] A system including: a QR code generating means for generating a QR code including a special coupon and link destination information.
[2228] (Claim 2)
[2229] The system described in claim 1, characterized in that the analysis means includes an image analysis means that detects the age and gender of passengers using image data acquired from the imaging means, and a voice analysis means that identifies the purpose of the ride using voice data acquired from the voice collection means.
[2230] (Claim 3)
[2231] The system according to claim 1, characterized in that the advertising generation means has a generation means for selecting advertising material according to the passenger's gender, age, and purpose of riding, and generating advertising content using a generative AI model.
[2232] "Application Example 1"
[2233] (Claim 1)
[2234] imaging means for acquiring facial images of passengers;
[2235] a voice collection means for collecting voices of passengers;
[2236] an analysis means for analyzing the data obtained from the imaging means and the voice collecting means and determining the gender, age, and purpose of the passenger;
[2237] an advertisement generating means for generating an optimal advertisement based on the passenger information determined by the analyzing means;
[2238] a display means for displaying the advertisement generated by the advertisement generation means;
[2239] a QR code generating means for providing a coupon or link related to the advertisement;
[2240] a display means for identifying that the display means is a display of smart glasses;
[2241] A system including:
[2242] (Claim 2)
[2243] The system described in claim 1, characterized in that the analysis means includes an image analysis means that detects the age and gender of passengers using image data acquired from the imaging means, and a voice analysis means that identifies the purpose of the ride using voice data acquired from the voice collection means.
[2244] (Claim 3)
[2245] The system described in claim 1, characterized in that the advertising generation means has a generation means for selecting advertising material according to the passenger's gender, age, and purpose of riding, generating advertising content, and generating QR codes including bonus coupons and shop links.
[2246] "Example 2: Combining Emotion Engines"
[2247] (Claim 1)
[2248] imaging means for acquiring facial images of passengers;
[2249] a voice collection means for collecting voices of passengers;
[2250] a communication means for transmitting data obtained from the imaging means and the sound collecting means to a server in real time;
[2251] analysis means for analyzing the received data to determine the passenger's gender, age, purpose of the ride, and emotional state;
[2252] an advertisement generating means for generating an optimal advertisement based on the passenger information and emotional state determined by the analyzing means;
[2253] a display means for displaying the advertisement generated by the advertisement generation means;
[2254] a QR code generating means for providing a coupon or link related to the advertisement;
[2255] a feedback collection means for acquiring passengers' emotional feedback on the displayed advertisements and for further analysis;
[2256] A system including:
[2257] (Claim 2)
[2258] The system described in claim 1, characterized in that the analysis means includes an image analysis means that detects the age and gender of the passenger using image data acquired from the imaging means, and an audio analysis means that identifies the purpose of the ride and emotional state using audio data acquired from the audio collection means.
[2259] (Claim 3)
[2260] The system according to claim 1, characterized in that the advertising generation means has a generative AI model that selects advertising materials based on the passenger's gender, age, purpose of riding, and emotional state, and generates advertising content.
[2261] "Application example 2 when combining emotion engines"
[2262] (Claim 1)
[2263] imaging means for acquiring facial images of passengers;
[2264] a voice collection means for collecting voices of passengers;
[2265] an analysis means for analyzing the data obtained from the imaging means and the voice collecting means and determining the gender, age, and purpose of the passenger;
[2266] an advertisement generating means for generating an optimal advertisement based on the passenger information determined by the analyzing means;
[2267] a display means for displaying the advertisement generated by the advertisement generation means;
[2268] a QR code generating means for providing a coupon or link related to the advertisement;
[2269] a system including an emotion engine that determines the emotional state of a passenger based on image data and audio data acquired from the imaging means;
[2270] an alert generation means for evaluating a security risk based on the emotional state and generating an alert;
[2271] a database matching means for matching with a dangerous person database and sending a real-time warning if there is a match;
[2272] A system including:
[2273] (Claim 2)
[2274] The system according to claim 1, characterized in that the analysis means comprises: an image analysis means for detecting the age and gender of the passenger using image data acquired from the imaging means; an audio analysis means for identifying the purpose of the ride using audio data acquired from the audio collection means; and an emotion analysis means for analyzing the emotional state of the passenger.
[2275] (Claim 3)
[2276] The system described in claim 1, characterized in that the advertising generation means has a generation means for selecting advertising material and generating advertising content based on the passenger's gender, age, and purpose of riding, and further has a means for assessing security risks based on the passenger's real-time emotional state and comparing it with a database of dangerous people. [Explanation of symbols]
[2277] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. imaging means for acquiring facial images of passengers; a voice collection means for collecting voices of passengers; an analysis means for analyzing the data obtained from the imaging means and the voice collecting means and determining the gender, age, and purpose of the passenger; an advertisement generating means for generating an optimal advertisement based on the passenger information determined by the analyzing means; a display means for displaying the advertisement generated by the advertisement generation means; a QR code generating means for providing a coupon or link related to the advertisement; A system including:
2. The system described in claim 1, characterized in that the analysis means includes an image analysis means that detects the age and gender of passengers using image data acquired from the imaging means, and a voice analysis means that identifies the purpose of the ride using voice data obtained from the voice collection means.
3. 2. The system according to claim 1, wherein said advertisement generating means comprises generating means for selecting advertisement material according to the gender, age and purpose of the passenger and generating advertisement content.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A