System

The system addresses the challenge of mismatched keywords in e-commerce by generating and modifying product images for search, improving accuracy and user satisfaction through visual search and database updates.

JP2026021068APending Publication Date: 2026-02-10SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024122750
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-02-10

Smart Images

  • Figure 2026021068000001_ABST
    Figure 2026021068000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for extracting an input search keyword; generation means for generating an image drawing based on the extracted search keyword; display means for displaying the generated image drawing on a user terminal; reception means for receiving an image drawing corrected by a user; conversion means for converting the corrected image drawing into a vector space; search means for searching for a similar product based on the converted vector; and transmission means for transmitting a search result to the user terminal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Current e-commerce systems have a problem where product searches fail if the user's search keywords do not match the descriptions in the product database. This causes problems such as users being unable to effectively find the products they are looking for, resulting in a poor user experience. Furthermore, if the search keywords are ambiguous or ambiguous, it is difficult to obtain appropriate search results. [Means for solving the problem]

[0005] The present invention provides a system that generates an image based on search keywords entered by a user and uses the image to perform a product search. Specifically, the system includes means for extracting the entered search keywords, means for generating an image based on the extracted search keywords, means for displaying the generated image on a user terminal, means for receiving an image modified by the user, means for converting the modified image into a vector space, means for searching for similar products based on the converted vectors, and means for transmitting search results to the user terminal. Furthermore, the system also provides means for receiving user feedback and updating the product database to improve the accuracy of the database and the quality of the search results. This allows users to efficiently search for and obtain desired products even if the keywords are not listed on the product.

[0006] A "search keyword" is text information that a user enters to search for a desired product.

[0007] The "generation means" is a part that executes functions and algorithms to generate an image based on search keywords.

[0008] The "display means" is a part that has the function of displaying the generated image diagram on the user terminal.

[0009] The "receiving means" is a part that executes the function of receiving an image diagram modified by a user.

[0010] The "conversion means" is a part that has the function of converting the corrected image into a vector space.

[0011] The "search means" is a part that has the function of searching for similar products based on the converted vector.

[0012] The "transmission means" is a part that has a function of transmitting search results to the user terminal.

[0013] The "editing means" is a part that has a function for correcting the image displayed on the user terminal.

[0014] The "update means" is a part that has the function of receiving user feedback and updating the product database. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] This invention is a system that generates an image of a product associated with a user's search keywords and uses that image to search for products. This system is realized by linking a server and a user terminal.

[0037] System configuration

[0038] The system consists of the following main components:

[0039] 1. Server

[0040] 2. User Device

[0041] 3. Generative AI Models

[0042] 4. Product Database

[0043] Program processing

[0044] (1) Enter keywords and submit

[0045] User: Enters a keyword describing the desired product into the device's search bar (e.g., user enters "red sneakers").

[0046] Terminal: Generates a request to send the entered keyword to the server and sends it to the server.

[0047] (2) Creating an image

[0048] Server: Extract keywords from the incoming request.

[0049] Server: Passes the extracted keywords to a generative AI model (e.g., image generation model).

[0050] Generative AI model: Generates image images of products associated with keywords and returns them to the server.

[0051] Server: Sends the generated image to the user's device.

[0052] (3) Display and edit the image

[0053] Terminal: Displays the image received from the server on the user interface.

[0054] User: Check the displayed image and make any necessary modifications (e.g., change the color, add / delete items, etc.) (e.g., the user changes the color of the sneakers to dark red).

[0055] Terminal: Generate new image data that reflects the modifications and send it to the server.

[0056] (4) Product search

[0057] Server: Converts the received corrected image into vector space.

[0058] Server: Based on the converted vector, it compares it with the image vectors in the product database and extracts similar products.

[0059] Server: Lists products with high matching scores and sends this information to the user's device.

[0060] (5) Search result presentation and feedback

[0061] Terminal: The search results received from the server are displayed in a user interface (e.g., three red sneakers are displayed as options).

[0062] User: Review search results and enter satisfaction feedback (e.g., "As expected," "Disappointing," etc.) into the device.

[0063] Device: Sends feedback information to the server.

[0064] (6) Database Improvement

[0065] Server: Analyzes the received user feedback.

[0066] Server: Based on the feedback, consider ways to update and improve the product database (e.g., add new product data, modify existing data, or tune the generative AI model or search algorithm).

[0067] In this way, the present invention goes beyond the limitations of conventional keyword searches and provides a means for users to intuitively and effectively search for the products they want. Through specific steps, we achieve improved search quality and increased user satisfaction.

[0068] The processing flow will be explained below.

[0069] Step 1:

[0070] The user enters a keyword describing the desired product into the device's search bar and presses the search button (e.g., the user enters "red sneakers").

[0071] Step 2:

[0072] The terminal generates a request to send the input keyword to the server, and sends it to the server.

[0073] Step 3:

[0074] The server extracts search keywords from the incoming request and passes them as input to the generative AI model.

[0075] Step 4:

[0076] The generative AI model generates an image of a product associated with the search keyword and returns that image to the server.

[0077] Step 5:

[0078] The server transmits the generated image to the user's terminal.

[0079] Step 6:

[0080] The terminal displays the image received from the server on the user interface.

[0081] Step 7:

[0082] The user checks the displayed image, makes corrections as necessary (e.g., the user changes the color of the sneakers to dark red), and confirms the corrected image on the terminal.

[0083] Step 8:

[0084] The terminal sends the corrected image to the server.

[0085] Step 9:

[0086] The server converts the corrected image into a vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN).

[0087] Step 10:

[0088] The server then uses the converted vector to calculate the similarity between the image vectors in the product database, using techniques such as cosine similarity and Euclidean distance.

[0089] Step 11:

[0090] The server lists similar products and sends the information to the user's terminal.

[0091] Step 12:

[0092] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[0093] Step 13:

[0094] The user reviews the search results and enters their satisfaction feedback (e.g., "As expected," "Disappointed," etc.).

[0095] Step 14:

[0096] The terminal transmits the user's feedback information to the server.

[0097] Step 15:

[0098] The server analyzes the feedback it receives and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[0099] Example 1

[0100] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0101] Conventional keyword search systems have difficulty reflecting the specific image of the product a user is looking for, resulting in low search accuracy. Furthermore, they lack the ability for users to visually modify their image or provide feedback, which often results in search results that do not meet user expectations. Furthermore, there is a lack of a way to continuously improve the product database using user feedback.

[0102] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0103] In this invention, the server includes means for extracting input search keywords, generating means for generating an image based on the extracted search keywords, display means for displaying the generated image on a user terminal, receiving means for receiving an image modified by the user, conversion means for converting the modified image into a vector space, search means for searching for similar products based on the converted vectors, transmission means for transmitting search results to the user terminal, update means for receiving user feedback and updating the product database, and means for passing the input search keywords to the generative AI model as a prompt sentence and receiving an image generated based on the prompt sentence. This allows users to search for products based on a visual image, and enables search accuracy to be improved through corrections and feedback.

[0104] The "means for extracting input search keywords" refers to a method or device for extracting search keywords input by a user to a terminal.

[0105] The "means for generating an image diagram based on the extracted search keywords" refers to a method or device for generating a related image diagram based on the extracted keywords.

[0106] The "display means for displaying the generated image diagram on the user terminal" refers to a method or device for displaying the generated image diagram on the screen of the user's terminal.

[0107] The "receiving means for receiving an image diagram modified by a user" refers to a method or device for receiving image diagram data modified by a user.

[0108] The "conversion means for converting the corrected image into a vector space" is a method or device for converting the image corrected by the user into a numerical vector.

[0109] The "search means for searching for similar products based on the converted vectors" refers to a method or device for searching a database for similar products using the converted vector data.

[0110] The "transmission means for transmitting search results to the user terminal" refers to a method or device for transmitting information about the searched products to the user terminal.

[0111] The "means for receiving user feedback and updating the product database" refers to a method or device for receiving feedback information from users and updating the contents of the product database.

[0112] "Means for passing search keywords input to a generative AI model as a prompt sentence and receiving an image diagram generated based on the prompt sentence" refers to a method or device for passing search keywords as a prompt sentence to a generative AI model and receiving the resulting image diagram.

[0113] This invention is a system that generates product image images associated with a user's search keywords and uses the images to search for products. This system is implemented by a server and a user terminal, and utilizes a generative AI model and a product database.

[0114] Hardware and software used

[0115] 1. Server:

[0116] Use a high performance computer server.

[0117] The software uses DeepAI's generative AI model and TensorFlow for image generation and data processing.

[0118] The search algorithm uses Elasticsearch and other search engines.

[0119] 2. User Device:

[0120] A personal computer or smartphone is used for user operation.

[0121] The user interface is built using technologies such as HTML5, JavaScript, and CSS.

[0122] Specific processing of the program

[0123] Enter keywords and send

[0124] The user enters a keyword that describes the desired product into the search bar of the device. For example, the user enters "red sneakers." In response to this input, the device sends a request including the corresponding keyword to the server.

[0125] Generate image diagrams

[0126] The server receives the request sent from the device and extracts the search keywords. The extracted keywords are passed to the generative AI model as a prompt. The generative AI model generates an image based on the prompt, for example, "red sneakers," and returns it to the server. The server then sends the generated image to the user's device.

[0127] Display and edit images

[0128] The terminal displays the image received from the server on the user interface. The user can check the displayed image and modify it as necessary. For example, the user may change the color of the sneakers to dark red. Once the modification is made, the terminal transmits the modified image data to the server.

[0129] Product search

[0130] The server converts the corrected image into vector space using machine learning libraries such as TensorFlow. The converted vector is compared with other image vectors in a product database. For example, the server uses Elasticsearch to calculate similarity and extract similar products. The server then lists products with high similarity and sends the search results to the user's device.

[0131] Search results and feedback

[0132] The terminal displays the search result product list received from the server on a user interface. For example, three red sneaker options are displayed. The user can review the search results and enter feedback regarding satisfaction. The terminal then sends this feedback information to the server.

[0133] Database improvements

[0134] The server analyzes the feedback received from users. A machine learning library (e.g., scikit-learn) is used for the analysis. The server updates the product database based on the analysis results and also tunes the generative AI model and search algorithm. This makes it possible to continuously improve search accuracy.

[0135] Examples of concrete examples and prompts

[0136] As a concrete example, consider the case where a user enters the search keyword "red sneakers." The following prompt is passed to the generative AI model:

[0137] "Generate an image of red sneakers."

[0138] "Create an image of a pair of dark red sneakers."

[0139] Based on this prompt, the generative AI model generates an image, which the user can then review and modify, and then search for and present more similar products.

[0140] In this way, the present invention provides users with the ability to visually search for products and provides a means for continually improving the search system based on feedback.

[0141] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0142] Step 1:

[0143] A user enters a keyword that describes the desired product into the search bar of the device. For example, the user enters "red sneakers." The device captures the entered keyword and generates an HTTP request to send to the server. The input is the text data "red sneakers," and the output is the request data to the server.

[0144] Step 2:

[0145] The server receives the HTTP request sent from the device. Next, the server extracts the keyword "red sneakers" from the request. It generates a prompt based on this extracted keyword and sends it to the generative AI model. The input is the HTTP request data, and the output is the prompt to the generative AI model.

[0146] Step 3:

[0147] The generative AI model generates an image based on the received prompt "red sneakers." In this process, the generative AI model uses its internal algorithm to compose relevant image data. The generated image is then returned to the server. The input is the prompt, and the output is the generated image.

[0148] Step 4:

[0149] The server receives the image diagram returned from the generative AI model. It encodes the received image diagram in an appropriate format and generates an HTTP response to send to the user device. The input is the generated image diagram, and the output is an HTTP response containing the image diagram data to the user device.

[0150] Step 5:

[0151] The terminal receives the HTTP response from the server, decodes the image data, and displays it on the user interface. The user checks the displayed image and makes any necessary corrections. For example, the user changes the color of the sneakers to dark red. The input is the response from the server, and the output is the corrected image.

[0152] Step 6:

[0153] The terminal generates image data modified by the user and generates an HTTP request to send it to the server. The input is the modified image, and the output is an HTTP request including the modified image data to the server.

[0154] Step 7:

[0155] The server receives an HTTP request from the device and retrieves the corrected image. The server then converts the corrected image into a vector space using a machine learning library such as TensorFlow. The input is the corrected image, and the output is vector data.

[0156] Step 8:

[0157] The server uses the vector data to compare it with other image vectors in the product database and calculates the similarity. For example, it uses distance calculations such as cosine similarity. It extracts products with high similarity and lists them. The input is vector data, and the output is a list of similar products.

[0158] Step 9:

[0159] The server generates an HTTP response to send the list of similar products to the user terminal. The input is the list of similar products, and the output is the response to the user terminal.

[0160] Step 10:

[0161] The terminal receives the response from the server and displays similar products on the user interface. For example, three red sneakers are displayed. The user checks the displayed products and inputs feedback on their satisfaction into the terminal. The input is the response from the server, and the output is the user's feedback.

[0162] Step 11:

[0163] The terminal sends an HTTP request containing the user's feedback to the server. The input is the user's feedback data, and the output is the request to the server.

[0164] Step 12:

[0165] The server receives feedback from the device and analyzes the data to tune the product database, generative AI model, and search algorithm. Machine learning libraries such as scikit-learn are used for the analysis. The input is user feedback data, and the output is an updated and improved database and model.

[0166] (Application example 1)

[0167] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0168] Conventional keyword search systems have the problem that it is difficult for users to intuitively search for products. Also, there is no way to quickly search for an object found in the real world on an online shopping site. This makes it difficult for users to easily find the product they want, resulting in a decrease in satisfaction.

[0169] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0170] In this invention, the server includes means for extracting input search keywords, generating means for generating an image based on the extracted search keywords, display means for displaying the generated image on the user terminal, receiving means for receiving the image modified by the user, conversion means for converting the modified image into a vector space, search means for searching for similar products based on the converted vectors, transmission means for transmitting search results to the user terminal, and real-world image search means for receiving real-world image data captured by a camera of the user terminal and searching for products based on the real-world image data. This allows the user to quickly use objects found in the real world in product searches and easily find products with intuitive operations.

[0171] A "search keyword" is a word or phrase that a user enters to search for a product.

[0172] The "generation means" refers to a device or software that has the function of generating an image based on the input search keywords.

[0173] A "user terminal" is a device used by a user, including a smartphone, tablet, smart glasses, etc.

[0174] The "display means" is a device or software that has the function of displaying the generated image diagram on the screen of the user terminal.

[0175] The "receiving means" is a device or software that has the function of receiving an image modified by a user.

[0176] The "conversion means" is a device or software that has the function of converting the modified image into vector space.

[0177] The "search means" is a device or software that has the function of searching a product database for similar products based on the converted vector.

[0178] The "transmission means" is a device or software that has the function of transmitting search results to a user terminal.

[0179] The "real world image search means" is a device or software that has the function of receiving real world image data captured by the camera of the user terminal and searching for products based on that image data.

[0180] This invention provides a system that generates product images based on user search keywords and real-world images, and then searches for products based on the images. The details of the system are described below.

[0181] System configuration

[0182] The system consists of the following main components:

[0183] 1. Server

[0184] 2. User Device

[0185] 3. Generative AI Models

[0186] 4. Product Database

[0187] Hardware and Software Use

[0188] Server: A high-performance computer that processes data and runs AI models.

[0189] User terminal: A smartphone, tablet, smart glasses, or other mobile device with a camera and display.

[0190] Generative AI models: Examples include image generation models such as Stable Diffusion and DALL-E.

[0191] Product database: A database for registering and managing product information.

[0192] Data processing and calculation

[0193] Extract and submit search keywords:

[0194] A user enters keywords into the search bar using smart glasses or a smartphone, and the user device sends the entered keywords to the server.

[0195] Generate image diagrams:

[0196] The server extracts the received search keywords, passes them to the generative AI model, and generates an image. The generated image is then sent to the user's device.

[0197] View and modify images:

[0198] The user terminal displays the image received from the server, and the user can make any necessary corrections. The corrected image is then sent back to the server. This series of operations is performed through the user interface.

[0199] Transformation to vector space and product search:

[0200] The server converts the corrected image into a vector space and compares this vector with the image vectors in the product database to extract similar products. The extracted product list is sent to the user's terminal and displayed in a selectable format.

[0201] Search based on real-world objects:

[0202] Users use the smart glasses' camera to capture real-world objects, send the image data to a server, which uses a generative AI model to search for products that match the real-world image, and send the search results to the user's device for display.

[0203] Specific examples

[0204] For example, if a user likes a pair of red sneakers they see in the park, they can use the smart glasses to capture an image of the sneakers and then use the application-generated image to search for similar items on an online shopping site. An example prompt for this is:

[0205] Example prompt sentence:

[0206] "Search for products similar to the red sneakers found in the park."

[0207] The above system and processing enable users to quickly use objects they find in the real world to search for products, making it easy to find products through intuitive operations.

[0208] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0209] Step 1: Enter and submit your search keywords

[0210] The user enters a keyword describing the desired product into the search bar of the device. Specifically, the user enters a word such as "red sneakers." The device generates a request to send the entered keyword to the server and sends it to the server. This input data is sent to the server as is.

[0211] Step 2: Generate an image

[0212] The server extracts search keywords from the incoming request. The extracted keywords are passed to a generative AI model (e.g., an image generation model). The generative AI model generates an image of a product associated with the keyword and returns the image to the server. The server then sends the generated image to the user's device. The input data is the search keyword, and the generated image is obtained as the output.

[0213] Step 3: View and modify the image

[0214] The user terminal displays the image received from the server on its display. The user checks the displayed image and makes any necessary modifications, such as changing colors or adding / deleting items. Once modifications are complete, the terminal sends the modified image to the server. The input data is the generated image, and the image after modifications made by the user are output.

[0215] Step 4: Transform to vector space

[0216] The server converts the received modified image into a vector space, where the image features are represented as numerical vectors. The input data is the modified image, and the output is a vector representation.

[0217] Step 5: Product Search

[0218] The server compares the converted vector with the image vectors stored in the product database to search for similar products. The input data is a vector representation, and similar products are listed. The output data is a list of products with high matching scores.

[0219] Step 6: Presenting search results

[0220] The user terminal displays a list of search results received from the server. The user can check the search results and select the products they like. The input data is the search results sent from the server, and is presented to the user as output.

[0221] Step 7: Submit your feedback

[0222] The user checks the search results and inputs their satisfaction feedback into the terminal, which then sends this feedback information to the server. The input data is the user's feedback, which is sent to the server as output.

[0223] Step 8: Improve your product database

[0224] The server analyzes the received user feedback and considers updating the product database and improving the generative AI model. This process aims to improve the accuracy of the system and user satisfaction. The input data is user feedback, and the output is an update to the product database.

[0225] Step 9: Search based on real-world objects

[0226] The user captures a real-world object using the camera on the smart glasses. The device then sends the captured image data to the server. The server then passes the received image data to a generative AI model, which generates an image. Based on the generated image data, a product database is searched for similar products. The input data is the captured real-world image data, and the output is a list of similar products.

[0227] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0228] This invention is a system that generates product images associated with search keywords entered by a user and uses those images to search for products. This system further improves the user experience by incorporating an emotion engine that recognizes the user's emotions and adjusts search results accordingly.

[0229] System configuration

[0230] The system consists of the following main components:

[0231] 1. Server

[0232] 2. User Device

[0233] 3. Generative AI Models

[0234] 4. Product Database

[0235] 5. Emotion Engine

[0236] Program processing

[0237] (1) Enter keywords and submit

[0238] The user enters a keyword describing the desired product into the device's search bar and presses the search button (e.g., the user enters "red sneakers").

[0239] The terminal generates a request to send the input keyword to the server, and sends it to the server.

[0240] (2) Creating an image

[0241] The server extracts keywords from the incoming request and passes them as input to the generative AI model.

[0242] The generative AI model generates an image of a product associated with the search keyword and returns that image to the server.

[0243] The server transmits the generated image to the user's terminal.

[0244] (3) Display and edit the image

[0245] The terminal displays the image received from the server on the user interface.

[0246] The user checks the displayed image and makes corrections as necessary (e.g., the user changes the color of the sneakers to dark red).

[0247] The terminal generates new image data that reflects the corrections and sends it to the server.

[0248] (4) Product search

[0249] The server converts the corrected image into a vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN).

[0250] The server then uses the converted vector to calculate the similarity between the image vectors in the product database, using techniques such as cosine similarity and Euclidean distance.

[0251] The server lists similar products and sends the information to the user's terminal.

[0252] (5) Presentation of search results

[0253] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[0254] (6) Emotion recognition and search result adjustment

[0255] The emotion engine analyzes the user's facial expressions and voice to identify their emotions. For example, it uses a camera and microphone to analyze the user's facial expressions and tone of voice and recognize emotions such as "happiness," "confusion," and "anger."

[0256] The server uses the output of the emotion engine to adjust search results according to the user's emotions. For example, if the user has a dissatisfied expression, the server can change search results or add new suggestions.

[0257] (7) Feedback and database improvement

[0258] The user provides feedback on the search results (e.g., "As expected," "Disappointing," etc.).

[0259] The terminal transmits the user's feedback information to the server.

[0260] The server analyzes the feedback it receives and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[0261] Specific examples

[0262] If a user types in "red sneakers" and the emotion engine detects an excited expression on the user's face, the server will prioritize search results that display sneakers with bolder designs to attract the user's attention.

[0263] If the user frowns at the search results, the emotion engine detects the emotion of dissatisfaction and the server will either readjust the search results or display additional product suggestions.

[0264] In this way, the present invention realizes a system that provides higher user satisfaction and search accuracy by searching for and suggesting products while taking into account the user's emotions.

[0265] The processing flow will be explained below.

[0266] Step 1:

[0267] The user enters a keyword describing the desired product into the device's search bar and presses the search button (e.g., the user enters "red sneakers").

[0268] Step 2:

[0269] The terminal generates a request to send the input keyword to the server, and sends it to the server.

[0270] Step 3:

[0271] The server extracts search keywords from the incoming request and passes them as input to the generative AI model.

[0272] Step 4:

[0273] The generative AI model generates an image of a product associated with the search keyword and returns that image to the server.

[0274] Step 5:

[0275] The server transmits the generated image to the user's terminal.

[0276] Step 6:

[0277] The terminal displays the image received from the server on the user interface.

[0278] Step 7:

[0279] The user checks the displayed image, makes corrections as necessary (e.g., the user changes the color of the sneakers to dark red), and confirms the corrected image on the terminal.

[0280] Step 8:

[0281] The terminal sends the corrected image to the server.

[0282] Step 9:

[0283] The server converts the corrected image into a vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN).

[0284] Step 10:

[0285] The server then uses the converted vector to calculate the similarity between the image vectors in the product database, using techniques such as cosine similarity and Euclidean distance.

[0286] Step 11:

[0287] The server lists similar products and sends the information to the user's terminal.

[0288] Step 12:

[0289] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[0290] Step 13:

[0291] The emotion engine analyzes the user's facial expressions and voice to identify their emotions. For example, it uses a camera and microphone to analyze the user's facial expressions and tone of voice and recognize emotions such as "happiness," "confusion," and "anger."

[0292] Step 14:

[0293] The server uses the output of the emotion engine to adjust search results according to the user's emotions. For example, if the user has a dissatisfied expression, the server can change search results or add new suggestions.

[0294] Step 15:

[0295] The user provides feedback on the search results (e.g., "As expected," "Disappointing," etc.).

[0296] Step 16:

[0297] The terminal transmits the user's feedback information to the server.

[0298] Step 17:

[0299] The server analyzes the feedback it receives and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[0300] This allows the user to perform optimal product searches based on search keywords, intuitive images, and even their own emotions.

[0301] Example 2

[0302] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0303] Conventional search systems search for products based on keywords entered by the user, but the search results often do not match the user's expectations or preferences. In addition, there is no function to adjust the search results taking into account the user's emotions, which leads to a poor user experience.

[0304] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for extracting input search keywords, generating an image based on the extracted search keywords, displaying the generated image on the user terminal, receiving means for receiving an image modified by the user, converting means for converting the modified image into a vector space, searching means for searching for similar products based on the converted vectors, transmitting means for transmitting search results to the user terminal, and emotion recognition means for recognizing the user's emotion and adjusting the search results in accordance with the emotion. This makes it possible to present more personalized search results that take the user's emotion into consideration.

[0305] A "search keyword" is a word that a user inputs to specify a product or information to be searched for.

[0306] The "generation means" is a means for generating an image of a related product based on the input search keyword.

[0307] The "display means" is a means for displaying the generated image diagram and search results on the user terminal.

[0308] The "receiving means" is a means for transmitting the image diagram modified by the user and feedback information to the server.

[0309] The "transformation means" is a means for transforming the corrected image into a vector space.

[0310] The "search means" is a means for searching for similar products based on vectorized image diagrams.

[0311] The "transmission means" is a means for transmitting search results to the user terminal.

[0312] The "emotion recognition means" is a means for recognizing the user's emotions and adjusting search results according to the emotions.

[0313] The "editing means" is a means for the user to modify the displayed image diagram.

[0314] The "update means" is a means for updating the product database based on user feedback.

[0315] A "user terminal" is an electronic device that a user uses to conduct a search.

[0316] A "server" is a central computer device that manages the processing of the entire system and includes various means.

[0317] The "product database" is a database for storing information and image data about products.

[0318] A "vector space" is a multidimensional space that expresses an image in numerical form.

[0319] A "generative AI model" is an artificial intelligence model that generates image images of related products from input keywords.

[0320]

[0321] This invention is a system that generates product images associated with search keywords entered by a user and uses these images to search for products. The system incorporates an emotion recognition function that recognizes the user's emotions and adjusts search results accordingly. The system consists of the following main components: a server, a user terminal, a generative AI model, a product database, and an emotion engine.

[0322] (System Components)

[0323] Enter keywords and send

[0324] The user enters a keyword that describes the desired product into the search bar of the device and presses the search button. The device generates a request to send the entered keyword to the server and sends it to the server.

[0325] Generate image diagrams

[0326] The server extracts keywords from the incoming request and passes them as input to the generative AI model. The generative AI model generates product images associated with the search keywords and returns the images to the server. The server then sends the generated images to the user's device.

[0327] Display and edit images

[0328] The terminal displays the image received from the server on the user interface. The user checks the displayed image and makes any necessary corrections (e.g., changing the color). The terminal generates new image data that reflects the corrections and sends it to the server.

[0329] Product search

[0330] The server converts the corrected image it receives into vector space. Specifically, it vectorizes the image using a model such as a convolutional neural network (CNN). Based on the converted vector, the server calculates the similarity with the image vectors in the product database. This uses techniques such as cosine similarity and Euclidean distance. The server then lists similar products and sends this information to the user's device.

[0331] Presenting search results

[0332] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[0333] Emotion recognition and search result tailoring

[0334] The emotion engine analyzes the user's facial expressions and voice to identify their emotions. For example, it uses a camera and microphone to analyze the user's facial expressions and tone of voice and recognizes emotions such as "happiness," "confusion," and "anger." The server then adjusts the search results based on the output of the emotion engine according to the user's emotions. For example, if the user looks dissatisfied, the server will change the search results or add new suggestions.

[0335] Feedback and database improvements

[0336] The user enters feedback on the search results (e.g., "As expected," "Disappointing," etc.). The device sends the feedback information to the server. The server analyzes the received feedback and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[0337] (Example)

[0338] Example 1: Search for red sneakers

[0339] When a user enters the keyword "red sneakers" and the emotion engine detects the user's excited expression, the server prioritizes sneakers with bold designs in the search results list to attract the user's interest. Example prompt for the generative AI model: "Generate an image of red sneakers with an exciting design."

[0340] Example 2: Readjustment with a dissatisfied expression

[0341] If the user frowns at the search results, the emotion engine detects the emotion of dissatisfaction. The server then refines the search results or displays additional product suggestions. Example prompt for the generative AI model: "Generate an image of a classic red sneaker design."

[0342] In this way, the present invention takes user emotions into consideration when searching and suggesting products, thereby providing higher user satisfaction and search accuracy.

[0343] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0344] Step 1:

[0345] The user enters a search keyword into the search bar. The user enters a keyword such as "red sneakers" into the search bar of the device and presses the "Search" button. This causes the device to generate an HTTP request containing the keyword and send it to the server (input: search keyword, output: HTTP request).

[0346] Step 2:

[0347] The server receives the HTTP request and extracts keywords, such as "red sneakers," from the received request (input: HTTP request, output: search keyword).

[0348] Step 3:

[0349] The server starts generating an image by inputting keywords into the generative AI model. The server passes the extracted keywords to the generative AI model, generates a prompt, and inputs it into the AI ​​model. For example, a prompt such as "Please generate an image of red sneakers" is used (input: search keyword, output: prompt).

[0350] Step 4:

[0351] The generative AI model generates an image of a pair of red sneakers based on the input prompt and returns the data to the server (input: prompt, output: image).

[0352] Step 5:

[0353] The server sends the generated image to the user's terminal. It then generates an HTTP response including the image and sends it to the user's terminal (input: image, output: HTTP response).

[0354] Step 6:

[0355] The terminal receives the image and displays it on the user interface. The image is displayed on the terminal screen so that the user can check it (input: HTTP response, output: displayed image).

[0356] Step 7:

[0357] The user modifies the image. The user modifies each element (color, shape, etc.) of the image displayed in the user interface, and generates modified data (input: displayed image, output: modified image data).

[0358] Step 8:

[0359] The terminal sends the corrected image data to the server, which then generates new data reflecting the corrections and sends it back to the server (input: corrected image data, output: HTTP request).

[0360] Step 9:

[0361] The server receives the corrected image data and converts it into vector space. The server uses a convolutional neural network (CNN) to vectorize the corrected image data (input: corrected image data, output: vector data).

[0362] Step 10:

[0363] The server uses the vector data to search for similar products. Based on the converted vector, it calculates the cosine similarity and Euclidean distance between the image vector and the product database, and generates a list of similar products (input: vector data, output: list of similar products).

[0364] Step 11:

[0365] The server sends the search results to the user's device. It then generates an HTTP response containing a list of similar products and sends it to the user's device (input: list of similar products, output: HTTP response).

[0366] Step 12:

[0367] The terminal receives the search results and displays them on the user interface. The list of products obtained as search results is displayed on the screen (input: HTTP response, output: displayed search results).

[0368] Step 13:

[0369] The emotion engine recognizes the user's emotions. It uses a camera and microphone to analyze the user's facial expressions and tone of voice to identify emotions such as "joy," "confusion," and "anger" (input: user's facial expression and voice data, output: emotion data).

[0370] Step 14:

[0371] The server adjusts the search results based on the emotion data. It receives the emotion analysis results, adjusts the search results as needed, and regenerates the search results including new suggestions (input: emotion data, output: adjusted search results).

[0372] Step 15:

[0373] The user enters feedback on the search results. The user enters feedback on the search results (e.g., "As expected," "Disappointing," etc.) (Input: displayed search results, Output: feedback).

[0374] Step 16:

[0375] The terminal sends user feedback to the server. It generates a request including the feedback information and sends it to the server (input: feedback, output: HTTP request).

[0376] Step 17:

[0377] The server receives feedback and improves the product database and generative AI model. Based on the received feedback, the server updates the database, adjusts the model, and improves the search algorithm (input: feedback, output: updated database and improved model).

[0378] As a result, the present invention can perform searches and suggestions that take the user's emotions into consideration, providing higher user satisfaction and search accuracy.

[0379] (Application example 2)

[0380] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0381] Conventional search systems have the problem of being unable to provide search results that take users' emotions into account. As a result, users may not get the results they expect, resulting in a decrease in satisfaction. Another problem is that they are unable to suggest visually appealing products, making it difficult to stimulate purchasing desire.

[0382] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for extracting input search keywords, generating an image based on the extracted search keywords, displaying the generated image on the user terminal, receiving means for receiving an image modified by the user, conversion means for converting the modified image into a vector space, search means for searching for similar products based on the converted vectors, emotion recognition means for recognizing the user's emotion and adjusting the search results based on the emotion, and transmission means for transmitting the search results to the user terminal. This makes it possible to provide search results that reflect the user's emotion, improving user satisfaction and stimulating purchasing desire.

[0383] The "means for extracting input search keywords" refers to a device or software that has the function of detecting and acquiring search keywords input by a user.

[0384] The "means for generating an image based on the extracted search keywords" refers to a device or software that has the function of generating an image of a related product using the acquired search keywords.

[0385] The "display means for displaying the generated image diagram on the user terminal" refers to a device or software that displays the generated image diagram in a form that can be viewed by the user.

[0386] The "receiving means for receiving an image diagram modified by a user" refers to a device or software having a function for receiving data of an image diagram modified by a user.

[0387] The "conversion means for converting the corrected image into vector space" refers to a device or software that has the function of converting the image corrected by the user into vector data.

[0388] The "search means for searching for similar products based on the converted vectors" refers to a device or software that has the function of searching for similar products in a product database based on the vectorized image data.

[0389] "Emotion recognition means that recognizes a user's emotions and adjusts search results based on those emotions" refers to a device or software that has the function of analyzing emotions from a user's facial expressions, voice, etc., and changing or optimizing search results based on the results.

[0390] The "transmission means for transmitting search results to the user terminal" refers to a device or software that has the function of transmitting the search results to the user terminal.

[0391] System configuration

[0392] The system consists of the following main components:

[0393] 1. Server

[0394] 2. User Device

[0395] 3. Generative AI Models

[0396] 4. Product Database

[0397] 5. Emotion Recognition Engine

[0398] Program Overview

[0399] The system of the present invention allows users to input search keywords and search for products using images generated based on those keywords. It also includes a function to recognize user emotions and adjust search results accordingly.

[0400] Hardware and Software

[0401] Servers, user devices, generative AI models, etc. Specific examples include Python's PIL, Transformers (Huggingface), emotion_recognition library, and CLIP model.

[0402] Data processing and calculation

[0403] The server extracts the keywords entered by the user and generates an image using a generative AI model based on the keywords. The generated image is then sent to the user's device.

[0404] The user terminal displays the image received from the server, and the user can modify it as necessary. The modified image is then sent back to the server.

[0405] The server converts the corrected image into a vector space and searches for similar products in a product database based on the vectors. The search results are sent to the user's terminal.

[0406] The emotion recognition engine analyzes a user's facial expressions and voice to identify emotions and adjust search results accordingly.

[0407] Specific processing examples

[0408] For example, if a user enters the keyword "red sneakers," the system uses a generative AI model to generate a related image and display it to the user. The user can then modify the color of the sneakers to a darker red, and the modified image is sent to the server. The server then converts the modified image into a vector space and searches for similar products in the product database.

[0409] When search results are displayed on the user's device, an emotion recognition engine analyzes the user's facial expressions and voice, and the search results are adjusted based on their emotions. For example, if the user expresses dissatisfaction, the system can add other suggestions.

[0410] Recommended prompt examples

[0411] Keywords: "red sneakers"

[0412] Feedback: "User's Face Photo"

[0413] In this way, the present invention realizes higher user satisfaction and search accuracy by searching and proposing products taking into account the user's emotions.

[0414] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0415] Step 1:

[0416] The user enters a keyword describing a product into the search bar of the device and presses the search button. When "red sneakers" is entered as input, the device generates a request to send this keyword to the server and sends it to the server. The input is the keyword "red sneakers" and the output is the transmission of the request.

[0417] Step 2:

[0418] The server extracts keywords from the incoming request and passes them as input to the generative AI model. The model generates an image of the product from the keywords and returns the image data to the server. The input is the search keyword "red sneakers," and the output is the generated image data.

[0419] Step 3:

[0420] The server sends the generated image to the user's terminal. The input is the image data, and the output is the transmission of the image to the user's terminal.

[0421] Step 4:

[0422] The user terminal displays the image received from the server on the user interface. The user checks the displayed image and makes corrections, such as changing the color of the sneakers to dark red. The input is image data, and the output is a display containing the corrected image.

[0423] Step 5:

[0424] The user terminal generates new image data that reflects the modifications and sends it to the server. The input is the modified image, and the output is the transmission of the modified image data to the server.

[0425] Step 6:

[0426] The server converts the corrected image received into vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN). The input is the corrected image data, and the output is vector data.

[0427] Step 7:

[0428] The server uses the converted vector data to calculate the similarity with the image vectors in the product database. This uses techniques such as cosine similarity and Euclidean distance. The input is vector data, and the output is a list of similar products.

[0429] Step 8:

[0430] The server lists similar products and sends the information to the user terminal. The input is the list of similar products, and the output is the transmission of similar product information to the user terminal.

[0431] Step 9:

[0432] The user terminal displays the search results received from the server in a list on the user interface. The input is similar product information, and the output is the display to the user.

[0433] Step 10:

[0434] The emotion recognition engine analyzes the user's facial expressions and voice to identify emotions. Using a camera and microphone, the engine analyzes the user's facial expressions and tone of voice to recognize emotions such as "happiness," "confusion," and "anger." The input is the user's facial expression and voice data, and the output is identified emotional data.

[0435] Step 11:

[0436] The server adjusts search results based on the output of the emotion recognition engine according to the user's emotions. For example, if the user has a dissatisfied expression, the server changes the search results or adds new suggestions. The input is emotion data and search results, and the output is the adjusted search results.

[0437] Step 12:

[0438] The user inputs feedback on the search results, such as "as expected" or "disappointing," and the device sends this feedback information to the server. The input is user feedback, and the output is sending feedback information to the server.

[0439] Step 13:

[0440] The server analyzes the received feedback and updates the product database, adjusts the generative AI model, and improves the search algorithm. The input is the feedback information, and the output is improved system performance.

[0441] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0442] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0443] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0444] [Second embodiment]

[0445] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0446] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0447] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0448] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0449] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0450] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0451] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0452] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0453] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0454] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0455] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0456] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0457] This invention is a system that generates an image of a product associated with a user's search keywords and uses that image to search for products. This system is realized by linking a server and a user terminal.

[0458] System configuration

[0459] The system consists of the following main components:

[0460] 1. Server

[0461] 2. User Device

[0462] 3. Generative AI Models

[0463] 4. Product Database

[0464] Program processing

[0465] (1) Enter keywords and submit

[0466] User: Enters a keyword describing the desired product into the device's search bar (e.g., user enters "red sneakers").

[0467] Terminal: Generates a request to send the entered keyword to the server and sends it to the server.

[0468] (2) Creating an image

[0469] Server: Extract keywords from the incoming request.

[0470] Server: Passes the extracted keywords to a generative AI model (e.g., image generation model).

[0471] Generative AI model: Generates image images of products associated with keywords and returns them to the server.

[0472] Server: Sends the generated image to the user's device.

[0473] (3) Display and edit the image

[0474] Terminal: Displays the image received from the server on the user interface.

[0475] User: Check the displayed image and make any necessary modifications (e.g., change the color, add / delete items, etc.) (e.g., the user changes the color of the sneakers to dark red).

[0476] Terminal: Generate new image data that reflects the modifications and send it to the server.

[0477] (4) Product search

[0478] Server: Converts the received corrected image into vector space.

[0479] Server: Based on the converted vector, it compares it with the image vectors in the product database and extracts similar products.

[0480] Server: Lists products with high matching scores and sends this information to the user's device.

[0481] (5) Search result presentation and feedback

[0482] Terminal: The search results received from the server are displayed in a user interface (e.g., three red sneakers are displayed as options).

[0483] User: Review search results and enter satisfaction feedback (e.g., "As expected," "Disappointing," etc.) into the device.

[0484] Device: Sends feedback information to the server.

[0485] (6) Database Improvement

[0486] Server: Analyzes the received user feedback.

[0487] Server: Based on the feedback, consider ways to update and improve the product database (e.g., add new product data, modify existing data, or tune the generative AI model or search algorithm).

[0488] In this way, the present invention goes beyond the limitations of conventional keyword searches and provides a means for users to intuitively and effectively search for the products they want. Through specific steps, we achieve improved search quality and increased user satisfaction.

[0489] The processing flow will be explained below.

[0490] Step 1:

[0491] The user enters a keyword describing the desired product into the device's search bar and presses the search button (e.g., the user enters "red sneakers").

[0492] Step 2:

[0493] The terminal generates a request to send the input keyword to the server, and sends it to the server.

[0494] Step 3:

[0495] The server extracts search keywords from the incoming request and passes them as input to the generative AI model.

[0496] Step 4:

[0497] The generative AI model generates an image of a product associated with the search keyword and returns that image to the server.

[0498] Step 5:

[0499] The server transmits the generated image to the user's terminal.

[0500] Step 6:

[0501] The terminal displays the image received from the server on the user interface.

[0502] Step 7:

[0503] The user checks the displayed image, makes corrections as necessary (e.g., the user changes the color of the sneakers to dark red), and confirms the corrected image on the terminal.

[0504] Step 8:

[0505] The terminal sends the corrected image to the server.

[0506] Step 9:

[0507] The server converts the corrected image into a vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN).

[0508] Step 10:

[0509] The server then uses the converted vector to calculate the similarity between the image vectors in the product database, using techniques such as cosine similarity and Euclidean distance.

[0510] Step 11:

[0511] The server lists similar products and sends the information to the user's terminal.

[0512] Step 12:

[0513] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[0514] Step 13:

[0515] The user reviews the search results and enters their satisfaction feedback (e.g., "As expected," "Disappointed," etc.).

[0516] Step 14:

[0517] The terminal transmits the user's feedback information to the server.

[0518] Step 15:

[0519] The server analyzes the feedback it receives and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[0520] Example 1

[0521] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0522] Conventional keyword search systems have difficulty reflecting the specific image of the product a user is looking for, resulting in low search accuracy. Furthermore, they lack the ability for users to visually modify their image or provide feedback, which often results in search results that do not meet user expectations. Furthermore, there is a lack of a way to continuously improve the product database using user feedback.

[0523] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0524] In this invention, the server includes means for extracting input search keywords, generating means for generating an image based on the extracted search keywords, display means for displaying the generated image on a user terminal, receiving means for receiving an image modified by the user, conversion means for converting the modified image into a vector space, search means for searching for similar products based on the converted vectors, transmission means for transmitting search results to the user terminal, update means for receiving user feedback and updating the product database, and means for passing the input search keywords to the generative AI model as a prompt sentence and receiving an image generated based on the prompt sentence. This allows users to search for products based on a visual image, and enables search accuracy to be improved through corrections and feedback.

[0525] The "means for extracting input search keywords" refers to a method or device for extracting search keywords input by a user to a terminal.

[0526] The "means for generating an image diagram based on the extracted search keywords" refers to a method or device for generating a related image diagram based on the extracted keywords.

[0527] The "display means for displaying the generated image diagram on the user terminal" refers to a method or device for displaying the generated image diagram on the screen of the user's terminal.

[0528] The "receiving means for receiving an image diagram modified by a user" refers to a method or device for receiving image diagram data modified by a user.

[0529] The "conversion means for converting the corrected image into a vector space" is a method or device for converting the image corrected by the user into a numerical vector.

[0530] The "search means for searching for similar products based on the converted vectors" refers to a method or device for searching a database for similar products using the converted vector data.

[0531] The "transmission means for transmitting search results to the user terminal" refers to a method or device for transmitting information about the searched products to the user terminal.

[0532] The "means for receiving user feedback and updating the product database" refers to a method or device for receiving feedback information from users and updating the contents of the product database.

[0533] "Means for passing search keywords input to a generative AI model as a prompt sentence and receiving an image diagram generated based on the prompt sentence" refers to a method or device for passing search keywords as a prompt sentence to a generative AI model and receiving the resulting image diagram.

[0534] This invention is a system that generates product image images associated with a user's search keywords and uses the images to search for products. This system is implemented by a server and a user terminal, and utilizes a generative AI model and a product database.

[0535] Hardware and software used

[0536] 1. Server:

[0537] Use a high performance computer server.

[0538] The software uses DeepAI's generative AI model and TensorFlow for image generation and data processing.

[0539] The search algorithm uses Elasticsearch and other search engines.

[0540] 2. User Device:

[0541] A personal computer or smartphone is used for user operation.

[0542] The user interface is built using technologies such as HTML5, JavaScript, and CSS.

[0543] Specific processing of the program

[0544] Enter keywords and send

[0545] The user enters a keyword that describes the desired product into the search bar of the device. For example, the user enters "red sneakers." In response to this input, the device sends a request including the corresponding keyword to the server.

[0546] Generate image diagrams

[0547] The server receives the request sent from the device and extracts the search keywords. The extracted keywords are passed to the generative AI model as a prompt. The generative AI model generates an image based on the prompt, for example, "red sneakers," and returns it to the server. The server then sends the generated image to the user's device.

[0548] Display and edit images

[0549] The terminal displays the image received from the server on the user interface. The user can check the displayed image and modify it as necessary. For example, the user may change the color of the sneakers to dark red. Once the modification is made, the terminal transmits the modified image data to the server.

[0550] Product search

[0551] The server converts the corrected image into vector space using machine learning libraries such as TensorFlow. The converted vector is compared with other image vectors in a product database. For example, the server uses Elasticsearch to calculate similarity and extract similar products. The server then lists products with high similarity and sends the search results to the user's device.

[0552] Search results and feedback

[0553] The terminal displays the search result product list received from the server on a user interface. For example, three red sneaker options are displayed. The user can review the search results and enter feedback regarding satisfaction. The terminal then sends this feedback information to the server.

[0554] Database improvements

[0555] The server analyzes the feedback received from users. A machine learning library (e.g., scikit-learn) is used for the analysis. The server updates the product database based on the analysis results and also tunes the generative AI model and search algorithm. This makes it possible to continuously improve search accuracy.

[0556] Examples of concrete examples and prompts

[0557] As a concrete example, consider the case where a user enters the search keyword "red sneakers." The following prompt is passed to the generative AI model:

[0558] "Generate an image of red sneakers."

[0559] "Create an image of a pair of dark red sneakers."

[0560] Based on this prompt, the generative AI model generates an image, which the user can then review and modify, and then search for and present more similar products.

[0561] In this way, the present invention provides users with the ability to visually search for products and provides a means for continually improving the search system based on feedback.

[0562] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0563] Step 1:

[0564] A user enters a keyword that describes the desired product into the search bar of the device. For example, the user enters "red sneakers." The device captures the entered keyword and generates an HTTP request to send to the server. The input is the text data "red sneakers," and the output is the request data to the server.

[0565] Step 2:

[0566] The server receives the HTTP request sent from the device. Next, the server extracts the keyword "red sneakers" from the request. It generates a prompt based on this extracted keyword and sends it to the generative AI model. The input is the HTTP request data, and the output is the prompt to the generative AI model.

[0567] Step 3:

[0568] The generative AI model generates an image based on the received prompt "red sneakers." In this process, the generative AI model uses its internal algorithm to compose relevant image data. The generated image is then returned to the server. The input is the prompt, and the output is the generated image.

[0569] Step 4:

[0570] The server receives the image diagram returned from the generative AI model. It encodes the received image diagram in an appropriate format and generates an HTTP response to send to the user device. The input is the generated image diagram, and the output is an HTTP response containing the image diagram data to the user device.

[0571] Step 5:

[0572] The terminal receives the HTTP response from the server, decodes the image data, and displays it on the user interface. The user checks the displayed image and makes any necessary corrections. For example, the user changes the color of the sneakers to dark red. The input is the response from the server, and the output is the corrected image.

[0573] Step 6:

[0574] The terminal generates image data modified by the user and generates an HTTP request to send it to the server. The input is the modified image, and the output is an HTTP request including the modified image data to the server.

[0575] Step 7:

[0576] The server receives an HTTP request from the device and retrieves the corrected image. The server then converts the corrected image into a vector space using a machine learning library such as TensorFlow. The input is the corrected image, and the output is vector data.

[0577] Step 8:

[0578] The server uses the vector data to compare it with other image vectors in the product database and calculates the similarity. For example, it uses distance calculations such as cosine similarity. It extracts products with high similarity and lists them. The input is vector data, and the output is a list of similar products.

[0579] Step 9:

[0580] The server generates an HTTP response to send the list of similar products to the user terminal. The input is the list of similar products, and the output is the response to the user terminal.

[0581] Step 10:

[0582] The terminal receives the response from the server and displays similar products on the user interface. For example, three red sneakers are displayed. The user checks the displayed products and inputs feedback on their satisfaction into the terminal. The input is the response from the server, and the output is the user's feedback.

[0583] Step 11:

[0584] The terminal sends an HTTP request containing the user's feedback to the server. The input is the user's feedback data, and the output is the request to the server.

[0585] Step 12:

[0586] The server receives feedback from the device and analyzes the data to tune the product database, generative AI model, and search algorithm. Machine learning libraries such as scikit-learn are used for the analysis. The input is user feedback data, and the output is an updated and improved database and model.

[0587] (Application example 1)

[0588] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0589] Conventional keyword search systems have the problem that it is difficult for users to intuitively search for products. Also, there is no way to quickly search for an object found in the real world on an online shopping site. This makes it difficult for users to easily find the product they want, resulting in a decrease in satisfaction.

[0590] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0591] In this invention, the server includes means for extracting input search keywords, generating means for generating an image based on the extracted search keywords, display means for displaying the generated image on the user terminal, receiving means for receiving the image modified by the user, conversion means for converting the modified image into a vector space, search means for searching for similar products based on the converted vectors, transmission means for transmitting search results to the user terminal, and real-world image search means for receiving real-world image data captured by a camera of the user terminal and searching for products based on the real-world image data. This allows the user to quickly use objects found in the real world in product searches and easily find products with intuitive operations.

[0592] A "search keyword" is a word or phrase that a user enters to search for a product.

[0593] The "generation means" refers to a device or software that has the function of generating an image based on the input search keywords.

[0594] A "user terminal" is a device used by a user, including a smartphone, tablet, smart glasses, etc.

[0595] The "display means" is a device or software that has the function of displaying the generated image diagram on the screen of the user terminal.

[0596] The "receiving means" is a device or software that has the function of receiving an image modified by a user.

[0597] The "conversion means" is a device or software that has the function of converting the modified image into vector space.

[0598] The "search means" is a device or software that has the function of searching a product database for similar products based on the converted vector.

[0599] The "transmission means" is a device or software that has the function of transmitting search results to a user terminal.

[0600] The "real world image search means" is a device or software that has the function of receiving real world image data captured by the camera of the user terminal and searching for products based on that image data.

[0601] This invention provides a system that generates product images based on user search keywords and real-world images, and then searches for products based on the images. The details of the system are described below.

[0602] System configuration

[0603] The system consists of the following main components:

[0604] 1. Server

[0605] 2. User Device

[0606] 3. Generative AI Models

[0607] 4. Product Database

[0608] Hardware and Software Use

[0609] Server: A high-performance computer that processes data and runs AI models.

[0610] User terminal: A smartphone, tablet, smart glasses, or other mobile device with a camera and display.

[0611] Generative AI models: Examples include image generation models such as Stable Diffusion and DALL-E.

[0612] Product database: A database for registering and managing product information.

[0613] Data processing and calculation

[0614] Extract and submit search keywords:

[0615] A user enters keywords into the search bar using smart glasses or a smartphone, and the user device sends the entered keywords to the server.

[0616] Generate image diagrams:

[0617] The server extracts the received search keywords, passes them to the generative AI model, and generates an image. The generated image is then sent to the user's device.

[0618] View and modify images:

[0619] The user terminal displays the image received from the server, and the user can make any necessary corrections. The corrected image is then sent back to the server. This series of operations is performed through the user interface.

[0620] Transformation to vector space and product search:

[0621] The server converts the corrected image into a vector space and compares this vector with the image vectors in the product database to extract similar products. The extracted product list is sent to the user's terminal and displayed in a selectable format.

[0622] Search based on real-world objects:

[0623] Users use the smart glasses' camera to capture real-world objects, send the image data to a server, which uses a generative AI model to search for products that match the real-world image, and send the search results to the user's device for display.

[0624] Specific examples

[0625] For example, if a user likes a pair of red sneakers they see in the park, they can use the smart glasses to capture an image of the sneakers and then use the application-generated image to search for similar items on an online shopping site. An example prompt for this is:

[0626] Example prompt sentence:

[0627] "Search for products similar to the red sneakers found in the park."

[0628] The above system and processing enable users to quickly use objects they find in the real world to search for products, making it easy to find products through intuitive operations.

[0629] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0630] Step 1: Enter and submit your search keywords

[0631] The user enters a keyword describing the desired product into the search bar of the device. Specifically, the user enters a word such as "red sneakers." The device generates a request to send the entered keyword to the server and sends it to the server. This input data is sent to the server as is.

[0632] Step 2: Generate an image

[0633] The server extracts search keywords from the incoming request. The extracted keywords are passed to a generative AI model (e.g., an image generation model). The generative AI model generates an image of a product associated with the keyword and returns the image to the server. The server then sends the generated image to the user's device. The input data is the search keyword, and the generated image is obtained as the output.

[0634] Step 3: View and modify the image

[0635] The user terminal displays the image received from the server on its display. The user checks the displayed image and makes any necessary modifications, such as changing colors or adding / deleting items. Once modifications are complete, the terminal sends the modified image to the server. The input data is the generated image, and the image after modifications made by the user are output.

[0636] Step 4: Transform to vector space

[0637] The server converts the received modified image into a vector space, where the image features are represented as numerical vectors. The input data is the modified image, and the output is a vector representation.

[0638] Step 5: Product Search

[0639] The server compares the converted vector with the image vectors stored in the product database to search for similar products. The input data is a vector representation, and similar products are listed. The output data is a list of products with high matching scores.

[0640] Step 6: Presenting search results

[0641] The user terminal displays a list of search results received from the server. The user can check the search results and select the products they like. The input data is the search results sent from the server, and is presented to the user as output.

[0642] Step 7: Submit your feedback

[0643] The user checks the search results and inputs their satisfaction feedback into the terminal, which then sends this feedback information to the server. The input data is the user's feedback, which is sent to the server as output.

[0644] Step 8: Improve your product database

[0645] The server analyzes the received user feedback and considers updating the product database and improving the generative AI model. This process aims to improve the accuracy of the system and user satisfaction. The input data is user feedback, and the output is an update to the product database.

[0646] Step 9: Search based on real-world objects

[0647] The user captures a real-world object using the camera on the smart glasses. The device then sends the captured image data to the server. The server then passes the received image data to a generative AI model, which generates an image. Based on the generated image data, a product database is searched for similar products. The input data is the captured real-world image data, and the output is a list of similar products.

[0648] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0649] This invention is a system that generates product images associated with search keywords entered by a user and uses those images to search for products. This system further improves the user experience by incorporating an emotion engine that recognizes the user's emotions and adjusts search results accordingly.

[0650] System configuration

[0651] The system consists of the following main components:

[0652] 1. Server

[0653] 2. User Device

[0654] 3. Generative AI Models

[0655] 4. Product Database

[0656] 5. Emotion Engine

[0657] Program processing

[0658] (1) Enter keywords and submit

[0659] The user enters a keyword describing the desired product into the device's search bar and presses the search button (e.g., the user enters "red sneakers").

[0660] The terminal generates a request to send the input keyword to the server, and sends it to the server.

[0661] (2) Creating an image

[0662] The server extracts keywords from the incoming request and passes them as input to the generative AI model.

[0663] The generative AI model generates an image of a product associated with the search keyword and returns that image to the server.

[0664] The server transmits the generated image to the user's terminal.

[0665] (3) Display and edit the image

[0666] The terminal displays the image received from the server on the user interface.

[0667] The user checks the displayed image and makes corrections as necessary (e.g., the user changes the color of the sneakers to dark red).

[0668] The terminal generates new image data that reflects the corrections and sends it to the server.

[0669] (4) Product search

[0670] The server converts the corrected image into a vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN).

[0671] The server then uses the converted vector to calculate the similarity between the image vectors in the product database, using techniques such as cosine similarity and Euclidean distance.

[0672] The server lists similar products and sends the information to the user's terminal.

[0673] (5) Presentation of search results

[0674] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[0675] (6) Emotion recognition and search result adjustment

[0676] The emotion engine analyzes the user's facial expressions and voice to identify their emotions. For example, it uses a camera and microphone to analyze the user's facial expressions and tone of voice and recognize emotions such as "happiness," "confusion," and "anger."

[0677] The server uses the output of the emotion engine to adjust search results according to the user's emotions. For example, if the user has a dissatisfied expression, the server can change search results or add new suggestions.

[0678] (7) Feedback and database improvement

[0679] The user provides feedback on the search results (e.g., "As expected," "Disappointing," etc.).

[0680] The terminal transmits the user's feedback information to the server.

[0681] The server analyzes the feedback it receives and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[0682] Specific examples

[0683] If a user types in "red sneakers" and the emotion engine detects an excited expression on the user's face, the server will prioritize search results that display sneakers with bolder designs to attract the user's attention.

[0684] If the user frowns at the search results, the emotion engine detects the emotion of dissatisfaction and the server will either readjust the search results or display additional product suggestions.

[0685] In this way, the present invention realizes a system that provides higher user satisfaction and search accuracy by searching for and suggesting products while taking into account the user's emotions.

[0686] The processing flow will be explained below.

[0687] Step 1:

[0688] The user enters a keyword describing the desired product into the device's search bar and presses the search button (e.g., the user enters "red sneakers").

[0689] Step 2:

[0690] The terminal generates a request to send the input keyword to the server, and sends it to the server.

[0691] Step 3:

[0692] The server extracts search keywords from the incoming request and passes them as input to the generative AI model.

[0693] Step 4:

[0694] The generative AI model generates an image of a product associated with the search keyword and returns that image to the server.

[0695] Step 5:

[0696] The server transmits the generated image to the user's terminal.

[0697] Step 6:

[0698] The terminal displays the image received from the server on the user interface.

[0699] Step 7:

[0700] The user checks the displayed image, makes corrections as necessary (e.g., the user changes the color of the sneakers to dark red), and confirms the corrected image on the terminal.

[0701] Step 8:

[0702] The terminal sends the corrected image to the server.

[0703] Step 9:

[0704] The server converts the corrected image into a vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN).

[0705] Step 10:

[0706] The server then uses the converted vector to calculate the similarity between the image vectors in the product database, using techniques such as cosine similarity and Euclidean distance.

[0707] Step 11:

[0708] The server lists similar products and sends the information to the user's terminal.

[0709] Step 12:

[0710] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[0711] Step 13:

[0712] The emotion engine analyzes the user's facial expressions and voice to identify their emotions. For example, it uses a camera and microphone to analyze the user's facial expressions and tone of voice and recognize emotions such as "happiness," "confusion," and "anger."

[0713] Step 14:

[0714] The server adjusts search results based on the user's emotions based on the output of the emotion engine. For example, if the user has a dissatisfied expression, the server changes the search results or adds new suggestions.

[0715] Step 15:

[0716] The user provides feedback on the search results (e.g., "As expected," "Disappointing," etc.).

[0717] Step 16:

[0718] The terminal transmits the user's feedback information to the server.

[0719] Step 17:

[0720] The server analyzes the feedback it receives and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[0721] This allows the user to perform optimal product searches based on search keywords, intuitive images, and even their own emotions.

[0722] Example 2

[0723] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0724] Conventional search systems search for products based on keywords entered by the user, but the search results often do not match the user's expectations or preferences. In addition, there is no function to adjust the search results taking into account the user's emotions, which leads to a poor user experience.

[0725] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for extracting input search keywords, generating an image based on the extracted search keywords, displaying the generated image on the user terminal, receiving means for receiving an image modified by the user, converting means for converting the modified image into a vector space, searching means for searching for similar products based on the converted vectors, transmitting means for transmitting search results to the user terminal, and emotion recognition means for recognizing the user's emotion and adjusting the search results in accordance with the emotion. This makes it possible to present more personalized search results that take the user's emotion into consideration.

[0726] A "search keyword" is a word that a user inputs to specify a product or information to be searched for.

[0727] The "generation means" is a means for generating an image of a related product based on the input search keyword.

[0728] The "display means" is a means for displaying the generated image diagram and search results on the user terminal.

[0729] The "receiving means" is a means for transmitting the image diagram modified by the user and feedback information to the server.

[0730] The "transformation means" is a means for transforming the corrected image into a vector space.

[0731] The "search means" is a means for searching for similar products based on vectorized image diagrams.

[0732] The "transmission means" is a means for transmitting search results to the user terminal.

[0733] The "emotion recognition means" is a means for recognizing the user's emotions and adjusting search results according to the emotions.

[0734] The "editing means" is a means for the user to modify the displayed image diagram.

[0735] The "update means" is a means for updating the product database based on user feedback.

[0736] A "user terminal" is an electronic device that a user uses to conduct a search.

[0737] A "server" is a central computer device that manages the processing of the entire system and includes various means.

[0738] The "product database" is a database for storing information and image data about products.

[0739] A "vector space" is a multidimensional space that expresses an image in numerical form.

[0740] A "generative AI model" is an artificial intelligence model that generates image images of related products from input keywords.

[0741]

[0742] This invention is a system that generates product images associated with search keywords entered by a user and uses these images to search for products. The system incorporates an emotion recognition function that recognizes the user's emotions and adjusts search results accordingly. The system consists of the following main components: a server, a user terminal, a generative AI model, a product database, and an emotion engine.

[0743] (System Components)

[0744] Enter keywords and send

[0745] The user enters a keyword that describes the desired product into the search bar of the device and presses the search button. The device generates a request to send the entered keyword to the server and sends it to the server.

[0746] Generate image diagrams

[0747] The server extracts keywords from the incoming request and passes them as input to the generative AI model. The generative AI model generates product images associated with the search keywords and returns the images to the server. The server then sends the generated images to the user's device.

[0748] Display and edit images

[0749] The terminal displays the image received from the server on the user interface. The user checks the displayed image and makes any necessary corrections (e.g., changing the color). The terminal generates new image data that reflects the corrections and sends it to the server.

[0750] Product search

[0751] The server converts the corrected image it receives into vector space. Specifically, it vectorizes the image using a model such as a convolutional neural network (CNN). Based on the converted vector, the server calculates the similarity with the image vectors in the product database. This uses techniques such as cosine similarity and Euclidean distance. The server then lists similar products and sends this information to the user's device.

[0752] Presenting search results

[0753] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[0754] Emotion recognition and search result tailoring

[0755] The emotion engine analyzes the user's facial expressions and voice to identify their emotions. For example, it uses a camera and microphone to analyze the user's facial expressions and tone of voice and recognizes emotions such as "happiness," "confusion," and "anger." The server then adjusts the search results based on the output of the emotion engine according to the user's emotions. For example, if the user looks dissatisfied, the server will change the search results or add new suggestions.

[0756] Feedback and database improvements

[0757] The user enters feedback on the search results (e.g., "As expected," "Disappointing," etc.). The device sends the feedback information to the server. The server analyzes the received feedback and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[0758] (Example)

[0759] Example 1: Search for red sneakers

[0760] When a user enters the keyword "red sneakers" and the emotion engine detects the user's excited expression, the server prioritizes sneakers with bold designs in the search results list to attract the user's interest. Example prompt for the generative AI model: "Generate an image of red sneakers with an exciting design."

[0761] Example 2: Readjustment with a dissatisfied expression

[0762] If the user frowns at the search results, the emotion engine detects the emotion of dissatisfaction. The server then refines the search results or displays additional product suggestions. Example prompt for the generative AI model: "Generate an image of a classic red sneaker design."

[0763] In this way, the present invention takes user emotions into consideration when searching and suggesting products, thereby providing higher user satisfaction and search accuracy.

[0764] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0765] Step 1:

[0766] The user enters a search keyword into the search bar. The user enters a keyword such as "red sneakers" into the search bar of the device and presses the "Search" button. This causes the device to generate an HTTP request containing the keyword and send it to the server (input: search keyword, output: HTTP request).

[0767] Step 2:

[0768] The server receives the HTTP request and extracts keywords, such as "red sneakers," from the received request (input: HTTP request, output: search keyword).

[0769] Step 3:

[0770] The server starts generating an image by inputting keywords into the generative AI model. The server passes the extracted keywords to the generative AI model, generates a prompt, and inputs it into the AI ​​model. For example, a prompt such as "Please generate an image of red sneakers" is used (input: search keyword, output: prompt).

[0771] Step 4:

[0772] The generative AI model generates an image of a pair of red sneakers based on the input prompt and returns the data to the server (input: prompt, output: image).

[0773] Step 5:

[0774] The server sends the generated image to the user's terminal. It then generates an HTTP response including the image and sends it to the user's terminal (input: image, output: HTTP response).

[0775] Step 6:

[0776] The terminal receives the image and displays it on the user interface. The image is displayed on the terminal screen so that the user can check it (input: HTTP response, output: displayed image).

[0777] Step 7:

[0778] The user modifies the image. The user modifies each element (color, shape, etc.) of the image displayed in the user interface, and generates modified data (input: displayed image, output: modified image data).

[0779] Step 8:

[0780] The terminal sends the corrected image data to the server, which then generates new data reflecting the corrections and sends it back to the server (input: corrected image data, output: HTTP request).

[0781] Step 9:

[0782] The server receives the corrected image data and converts it into vector space. The server uses a convolutional neural network (CNN) to vectorize the corrected image data (input: corrected image data, output: vector data).

[0783] Step 10:

[0784] The server uses the vector data to search for similar products. Based on the converted vector, it calculates the cosine similarity and Euclidean distance between the image vector and the product database, and generates a list of similar products (input: vector data, output: list of similar products).

[0785] Step 11:

[0786] The server sends the search results to the user's device. It then generates an HTTP response containing a list of similar products and sends it to the user's device (input: list of similar products, output: HTTP response).

[0787] Step 12:

[0788] The terminal receives the search results and displays them on the user interface. The list of products obtained as search results is displayed on the screen (input: HTTP response, output: displayed search results).

[0789] Step 13:

[0790] The emotion engine recognizes the user's emotions. It uses a camera and microphone to analyze the user's facial expressions and tone of voice to identify emotions such as "joy," "confusion," and "anger" (input: user's facial expression and voice data, output: emotion data).

[0791] Step 14:

[0792] The server adjusts the search results based on the emotion data. It receives the emotion analysis results, adjusts the search results as needed, and regenerates the search results including new suggestions (input: emotion data, output: adjusted search results).

[0793] Step 15:

[0794] The user enters feedback on the search results. The user enters feedback on the search results (e.g., "As expected," "Disappointing," etc.) (Input: displayed search results, Output: feedback).

[0795] Step 16:

[0796] The terminal sends user feedback to the server. It generates a request including the feedback information and sends it to the server (input: feedback, output: HTTP request).

[0797] Step 17:

[0798] The server receives feedback and improves the product database and generative AI model. Based on the received feedback, the server updates the database, adjusts the model, and improves the search algorithm (input: feedback, output: updated database and improved model).

[0799] As a result, the present invention can perform searches and suggestions that take the user's emotions into consideration, providing higher user satisfaction and search accuracy.

[0800] (Application example 2)

[0801] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0802] Conventional search systems have the problem of being unable to provide search results that take users' emotions into account. As a result, users may not get the results they expect, resulting in a decrease in satisfaction. Another problem is that they are unable to suggest visually appealing products, making it difficult to stimulate purchasing desire.

[0803] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for extracting input search keywords, generating an image based on the extracted search keywords, displaying the generated image on the user terminal, receiving means for receiving an image modified by the user, conversion means for converting the modified image into a vector space, search means for searching for similar products based on the converted vectors, emotion recognition means for recognizing the user's emotion and adjusting the search results based on the emotion, and transmission means for transmitting the search results to the user terminal. This makes it possible to provide search results that reflect the user's emotion, improving user satisfaction and stimulating purchasing desire.

[0804] The "means for extracting input search keywords" refers to a device or software that has the function of detecting and acquiring search keywords input by a user.

[0805] The "means for generating an image based on the extracted search keywords" refers to a device or software that has the function of generating an image of a related product using the acquired search keywords.

[0806] The "display means for displaying the generated image diagram on the user terminal" refers to a device or software that displays the generated image diagram in a form that can be viewed by the user.

[0807] The "receiving means for receiving an image diagram modified by a user" refers to a device or software having a function for receiving data of an image diagram modified by a user.

[0808] The "conversion means for converting the corrected image into vector space" refers to a device or software that has the function of converting the image corrected by the user into vector data.

[0809] The "search means for searching for similar products based on the converted vectors" refers to a device or software that has the function of searching for similar products in a product database based on the vectorized image data.

[0810] "Emotion recognition means that recognizes a user's emotions and adjusts search results based on those emotions" refers to a device or software that has the function of analyzing emotions from a user's facial expressions, voice, etc., and changing or optimizing search results based on the results.

[0811] The "transmission means for transmitting search results to the user terminal" refers to a device or software that has the function of transmitting the search results to the user terminal.

[0812] System configuration

[0813] The system consists of the following main components:

[0814] 1. Server

[0815] 2. User Device

[0816] 3. Generative AI Models

[0817] 4. Product Database

[0818] 5. Emotion Recognition Engine

[0819] Program Overview

[0820] The system of the present invention allows users to input search keywords and search for products using images generated based on those keywords. It also includes a function to recognize user emotions and adjust search results accordingly.

[0821] Hardware and Software

[0822] Servers, user devices, generative AI models, etc. Specific examples include Python's PIL, Transformers (Huggingface), emotion_recognition library, and CLIP model.

[0823] Data processing and calculation

[0824] The server extracts the keywords entered by the user and generates an image using a generative AI model based on the keywords. The generated image is then sent to the user's device.

[0825] The user terminal displays the image received from the server, and the user can modify it as necessary. The modified image is then sent back to the server.

[0826] The server converts the corrected image into a vector space and searches for similar products in a product database based on the vectors. The search results are sent to the user's terminal.

[0827] The emotion recognition engine analyzes a user's facial expressions and voice to identify emotions and adjust search results accordingly.

[0828] Specific processing examples

[0829] For example, if a user enters the keyword "red sneakers," the system uses a generative AI model to generate a related image and display it to the user. The user can then modify the color of the sneakers to a darker red, and the modified image is sent to the server. The server then converts the modified image into a vector space and searches for similar products in the product database.

[0830] When search results are displayed on the user's device, an emotion recognition engine analyzes the user's facial expressions and voice, and the search results are adjusted based on their emotions. For example, if the user expresses dissatisfaction, the system can add other suggestions.

[0831] Recommended prompt examples

[0832] Keywords: "red sneakers"

[0833] Feedback: "User's Face Photo"

[0834] In this way, the present invention realizes higher user satisfaction and search accuracy by searching and proposing products taking into account the user's emotions.

[0835] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0836] Step 1:

[0837] The user enters a keyword describing a product into the search bar of the device and presses the search button. When "red sneakers" is entered as input, the device generates a request to send this keyword to the server and sends it to the server. The input is the keyword "red sneakers" and the output is the transmission of the request.

[0838] Step 2:

[0839] The server extracts keywords from the incoming request and passes them as input to the generative AI model. The model generates an image of the product from the keywords and returns the image data to the server. The input is the search keyword "red sneakers," and the output is the generated image data.

[0840] Step 3:

[0841] The server sends the generated image to the user's terminal. The input is the image data, and the output is the transmission of the image to the user's terminal.

[0842] Step 4:

[0843] The user terminal displays the image received from the server on the user interface. The user checks the displayed image and makes corrections, such as changing the color of the sneakers to dark red. The input is image data, and the output is a display containing the corrected image.

[0844] Step 5:

[0845] The user terminal generates new image data that reflects the modifications and sends it to the server. The input is the modified image, and the output is the transmission of the modified image data to the server.

[0846] Step 6:

[0847] The server converts the corrected image received into vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN). The input is the corrected image data, and the output is vector data.

[0848] Step 7:

[0849] The server uses the converted vector data to calculate the similarity with the image vectors in the product database. This uses techniques such as cosine similarity and Euclidean distance. The input is vector data, and the output is a list of similar products.

[0850] Step 8:

[0851] The server lists similar products and sends the information to the user terminal. The input is the list of similar products, and the output is the transmission of similar product information to the user terminal.

[0852] Step 9:

[0853] The user terminal displays the search results received from the server in a list on the user interface. The input is similar product information, and the output is the display to the user.

[0854] Step 10:

[0855] The emotion recognition engine analyzes the user's facial expressions and voice to identify emotions. Using a camera and microphone, the engine analyzes the user's facial expressions and tone of voice to recognize emotions such as "happiness," "confusion," and "anger." The input is the user's facial expression and voice data, and the output is identified emotional data.

[0856] Step 11:

[0857] The server adjusts search results based on the output of the emotion recognition engine according to the user's emotions. For example, if the user has a dissatisfied expression, the server changes the search results or adds new suggestions. The input is emotion data and search results, and the output is the adjusted search results.

[0858] Step 12:

[0859] The user inputs feedback on the search results, such as "as expected" or "disappointing," and the device sends this feedback information to the server. The input is user feedback, and the output is sending feedback information to the server.

[0860] Step 13:

[0861] The server analyzes the received feedback and updates the product database, adjusts the generative AI model, and improves the search algorithm. The input is the feedback information, and the output is improved system performance.

[0862] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0863] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0864] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0865] [Third embodiment]

[0866] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0867] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0868] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0869] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0870] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0871] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0872] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0873] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0874] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0875] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0876] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0877] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0878] This invention is a system that generates an image of a product associated with a user's search keywords and uses that image to search for products. This system is realized by linking a server and a user terminal.

[0879] System configuration

[0880] The system consists of the following main components:

[0881] 1. Server

[0882] 2. User Device

[0883] 3. Generative AI Models

[0884] 4. Product Database

[0885] Program processing

[0886] (1) Enter keywords and submit

[0887] User: Enters a keyword describing the desired product into the device's search bar (e.g., user enters "red sneakers").

[0888] Terminal: Generates a request to send the entered keyword to the server and sends it to the server.

[0889] (2) Creating an image

[0890] Server: Extract keywords from the incoming request.

[0891] Server: Passes the extracted keywords to a generative AI model (e.g., image generation model).

[0892] Generative AI model: Generates image images of products associated with keywords and returns them to the server.

[0893] Server: Sends the generated image to the user's device.

[0894] (3) Display and edit the image

[0895] Terminal: Displays the image received from the server on the user interface.

[0896] User: Check the displayed image and make any necessary modifications (e.g., change the color, add / delete items, etc.) (e.g., the user changes the color of the sneakers to dark red).

[0897] Terminal: Generate new image data that reflects the modifications and send it to the server.

[0898] (4) Product search

[0899] Server: Converts the received corrected image into vector space.

[0900] Server: Based on the converted vector, it compares it with the image vectors in the product database and extracts similar products.

[0901] Server: Lists products with high matching scores and sends this information to the user's device.

[0902] (5) Search result presentation and feedback

[0903] Terminal: The search results received from the server are displayed in a user interface (e.g., three red sneakers are displayed as options).

[0904] User: Review search results and enter satisfaction feedback (e.g., "As expected," "Disappointing," etc.) into the device.

[0905] Device: Sends feedback information to the server.

[0906] (6) Database Improvement

[0907] Server: Analyzes the received user feedback.

[0908] Server: Based on the feedback, consider ways to update and improve the product database (e.g., add new product data, modify existing data, or tune the generative AI model or search algorithm).

[0909] In this way, the present invention goes beyond the limitations of conventional keyword searches and provides a means for users to intuitively and effectively search for the products they want. Through specific steps, we achieve improved search quality and increased user satisfaction.

[0910] The processing flow will be explained below.

[0911] Step 1:

[0912] The user enters a keyword describing the desired product into the device's search bar and presses the search button (e.g., the user enters "red sneakers").

[0913] Step 2:

[0914] The terminal generates a request to send the input keyword to the server, and sends it to the server.

[0915] Step 3:

[0916] The server extracts search keywords from the incoming request and passes them as input to the generative AI model.

[0917] Step 4:

[0918] The generative AI model generates an image of a product associated with the search keyword and returns that image to the server.

[0919] Step 5:

[0920] The server transmits the generated image to the user's terminal.

[0921] Step 6:

[0922] The terminal displays the image received from the server on the user interface.

[0923] Step 7:

[0924] The user checks the displayed image, makes corrections as necessary (e.g., the user changes the color of the sneakers to dark red), and confirms the corrected image on the terminal.

[0925] Step 8:

[0926] The terminal sends the corrected image to the server.

[0927] Step 9:

[0928] The server converts the corrected image into a vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN).

[0929] Step 10:

[0930] The server then uses the converted vector to calculate the similarity between the image vectors in the product database, using techniques such as cosine similarity and Euclidean distance.

[0931] Step 11:

[0932] The server lists similar products and sends the information to the user's terminal.

[0933] Step 12:

[0934] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[0935] Step 13:

[0936] The user reviews the search results and enters their satisfaction feedback (e.g., "As expected," "Disappointed," etc.).

[0937] Step 14:

[0938] The terminal transmits the user's feedback information to the server.

[0939] Step 15:

[0940] The server analyzes the feedback it receives and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[0941] Example 1

[0942] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0943] Conventional keyword search systems have difficulty reflecting the specific image of the product a user is looking for, resulting in low search accuracy. Furthermore, they lack the ability for users to visually modify their image or provide feedback, which often results in search results that do not meet user expectations. Furthermore, there is a lack of a way to continuously improve the product database using user feedback.

[0944] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0945] In this invention, the server includes means for extracting input search keywords, generating means for generating an image based on the extracted search keywords, display means for displaying the generated image on a user terminal, receiving means for receiving an image modified by the user, conversion means for converting the modified image into a vector space, search means for searching for similar products based on the converted vectors, transmission means for transmitting search results to the user terminal, update means for receiving user feedback and updating the product database, and means for passing the input search keywords to the generative AI model as a prompt sentence and receiving an image generated based on the prompt sentence. This allows users to search for products based on a visual image, and enables search accuracy to be improved through corrections and feedback.

[0946] The "means for extracting input search keywords" refers to a method or device for extracting search keywords input by a user to a terminal.

[0947] The "means for generating an image diagram based on the extracted search keywords" refers to a method or device for generating a related image diagram based on the extracted keywords.

[0948] The "display means for displaying the generated image diagram on the user terminal" refers to a method or device for displaying the generated image diagram on the screen of the user's terminal.

[0949] The "receiving means for receiving an image diagram modified by a user" refers to a method or device for receiving image diagram data modified by a user.

[0950] The "conversion means for converting the corrected image into a vector space" is a method or device for converting the image corrected by the user into a numerical vector.

[0951] The "search means for searching for similar products based on the converted vectors" refers to a method or device for searching a database for similar products using the converted vector data.

[0952] The "transmission means for transmitting search results to the user terminal" refers to a method or device for transmitting information about the searched products to the user terminal.

[0953] The "means for receiving user feedback and updating the product database" refers to a method or device for receiving feedback information from users and updating the contents of the product database.

[0954] "Means for passing search keywords input to a generative AI model as a prompt sentence and receiving an image diagram generated based on the prompt sentence" refers to a method or device for passing search keywords as a prompt sentence to a generative AI model and receiving the resulting image diagram.

[0955] This invention is a system that generates product image images associated with a user's search keywords and uses the images to search for products. This system is implemented by a server and a user terminal, and utilizes a generative AI model and a product database.

[0956] Hardware and software used

[0957] 1. Server:

[0958] Use a high performance computer server.

[0959] The software uses DeepAI's generative AI model and TensorFlow for image generation and data processing.

[0960] The search algorithm uses Elasticsearch and other search engines.

[0961] 2. User Device:

[0962] A personal computer or smartphone is used for user operation.

[0963] The user interface is built using technologies such as HTML5, JavaScript, and CSS.

[0964] Specific processing of the program

[0965] Enter keywords and send

[0966] The user enters a keyword that describes the desired product into the search bar of the device. For example, the user enters "red sneakers." In response to this input, the device sends a request including the corresponding keyword to the server.

[0967] Generate image diagrams

[0968] The server receives the request sent from the device and extracts the search keywords. The extracted keywords are passed to the generative AI model as a prompt. The generative AI model generates an image based on the prompt, for example, "red sneakers," and returns it to the server. The server then sends the generated image to the user's device.

[0969] Display and edit images

[0970] The terminal displays the image received from the server on the user interface. The user can check the displayed image and modify it as necessary. For example, the user may change the color of the sneakers to dark red. Once the modification is made, the terminal transmits the modified image data to the server.

[0971] Product search

[0972] The server converts the corrected image into vector space using machine learning libraries such as TensorFlow. The converted vector is compared with other image vectors in a product database. For example, the server uses Elasticsearch to calculate similarity and extract similar products. The server then lists products with high similarity and sends the search results to the user's device.

[0973] Search results and feedback

[0974] The terminal displays the search result product list received from the server on a user interface. For example, three red sneaker options are displayed. The user can review the search results and enter feedback regarding satisfaction. The terminal then sends this feedback information to the server.

[0975] Database Improvements

[0976] The server analyzes the feedback received from users. A machine learning library (e.g., scikit-learn) is used for the analysis. The server updates the product database based on the analysis results and also tunes the generative AI model and search algorithm. This makes it possible to continuously improve search accuracy.

[0977] Examples of concrete examples and prompts

[0978] As a concrete example, consider the case where a user enters the search keyword "red sneakers." The following prompt is passed to the generative AI model:

[0979] "Generate an image of red sneakers."

[0980] "Create an image of a pair of dark red sneakers."

[0981] Based on this prompt, the generative AI model generates an image, which the user can then review and modify, and then search for and present more similar products.

[0982] In this way, the present invention provides users with the ability to visually search for products and provides a means for continually improving the search system based on feedback.

[0983] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0984] Step 1:

[0985] A user enters a keyword that describes the desired product into the search bar of the device. For example, the user enters "red sneakers." The device captures the entered keyword and generates an HTTP request to send to the server. The input is the text data "red sneakers," and the output is the request data to the server.

[0986] Step 2:

[0987] The server receives the HTTP request sent from the device. Next, the server extracts the keyword "red sneakers" from the request. It generates a prompt based on this extracted keyword and sends it to the generative AI model. The input is the HTTP request data, and the output is the prompt to the generative AI model.

[0988] Step 3:

[0989] The generative AI model generates an image based on the received prompt "red sneakers." In this process, the generative AI model uses its internal algorithm to compose relevant image data. The generated image is then returned to the server. The input is the prompt, and the output is the generated image.

[0990] Step 4:

[0991] The server receives the image diagram returned from the generative AI model. It encodes the received image diagram in an appropriate format and generates an HTTP response to send to the user device. The input is the generated image diagram, and the output is an HTTP response containing the image diagram data to the user device.

[0992] Step 5:

[0993] The terminal receives the HTTP response from the server, decodes the image data, and displays it on the user interface. The user checks the displayed image and makes any necessary corrections. For example, the user changes the color of the sneakers to dark red. The input is the response from the server, and the output is the corrected image.

[0994] Step 6:

[0995] The terminal generates image data modified by the user and generates an HTTP request to send it to the server. The input is the modified image, and the output is an HTTP request including the modified image data to the server.

[0996] Step 7:

[0997] The server receives an HTTP request from the device and retrieves the corrected image. The server then converts the corrected image into a vector space using a machine learning library such as TensorFlow. The input is the corrected image, and the output is vector data.

[0998] Step 8:

[0999] The server uses the vector data to compare it with other image vectors in the product database and calculates the similarity. For example, it uses distance calculations such as cosine similarity. It extracts products with high similarity and lists them. The input is vector data, and the output is a list of similar products.

[1000] Step 9:

[1001] The server generates an HTTP response to send the list of similar products to the user terminal. The input is the list of similar products, and the output is the response to the user terminal.

[1002] Step 10:

[1003] The terminal receives the response from the server and displays similar products on the user interface. For example, three red sneakers are displayed. The user checks the displayed products and inputs feedback on their satisfaction into the terminal. The input is the response from the server, and the output is the user's feedback.

[1004] Step 11:

[1005] The terminal sends an HTTP request containing the user's feedback to the server. The input is the user's feedback data, and the output is the request to the server.

[1006] Step 12:

[1007] The server receives feedback from the device and analyzes the data to tune the product database, generative AI model, and search algorithm. Machine learning libraries such as scikit-learn are used for the analysis. The input is user feedback data, and the output is an updated and improved database and model.

[1008] (Application example 1)

[1009] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1010] Conventional keyword search systems have the problem that it is difficult for users to intuitively search for products. Also, there is no way to quickly search for an object found in the real world on an online shopping site. This makes it difficult for users to easily find the product they want, resulting in a decrease in satisfaction.

[1011] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1012] In this invention, the server includes means for extracting input search keywords, generating means for generating an image based on the extracted search keywords, display means for displaying the generated image on the user terminal, receiving means for receiving the image modified by the user, conversion means for converting the modified image into a vector space, search means for searching for similar products based on the converted vectors, transmission means for transmitting search results to the user terminal, and real-world image search means for receiving real-world image data captured by a camera of the user terminal and searching for products based on the real-world image data. This allows the user to quickly use objects found in the real world in product searches and easily find products with intuitive operations.

[1013] A "search keyword" is a word or phrase that a user enters to search for a product.

[1014] The "generation means" refers to a device or software that has the function of generating an image based on the input search keywords.

[1015] A "user terminal" is a device used by a user, including a smartphone, tablet, smart glasses, etc.

[1016] The "display means" is a device or software that has the function of displaying the generated image diagram on the screen of the user terminal.

[1017] The "receiving means" is a device or software that has the function of receiving an image modified by a user.

[1018] The "conversion means" is a device or software that has the function of converting the modified image into vector space.

[1019] The "search means" is a device or software that has the function of searching a product database for similar products based on the converted vector.

[1020] The "transmission means" is a device or software that has the function of transmitting search results to a user terminal.

[1021] The "real world image search means" is a device or software that has the function of receiving real world image data captured by the camera of the user terminal and searching for products based on that image data.

[1022] This invention provides a system that generates product images based on user search keywords and real-world images, and then searches for products based on the images. The details of the system are described below.

[1023] System configuration

[1024] The system consists of the following main components:

[1025] 1. Server

[1026] 2. User Device

[1027] 3. Generative AI Models

[1028] 4. Product Database

[1029] Hardware and Software Use

[1030] Server: A high-performance computer that processes data and runs AI models.

[1031] User terminal: A smartphone, tablet, smart glasses, or other mobile device with a camera and display.

[1032] Generative AI models: Examples include image generation models such as Stable Diffusion and DALL-E.

[1033] Product database: A database for registering and managing product information.

[1034] Data processing and calculation

[1035] Extract and submit search keywords:

[1036] A user enters keywords into the search bar using smart glasses or a smartphone, and the user device sends the entered keywords to the server.

[1037] Generate image diagrams:

[1038] The server extracts the received search keywords, passes them to the generative AI model, and generates an image. The generated image is then sent to the user's device.

[1039] View and modify images:

[1040] The user terminal displays the image received from the server, and the user can make any necessary corrections. The corrected image is then sent back to the server. This series of operations is performed through the user interface.

[1041] Transformation to vector space and product search:

[1042] The server converts the corrected image into a vector space and compares this vector with the image vectors in the product database to extract similar products. The extracted product list is sent to the user's terminal and displayed in a selectable format.

[1043] Search based on real-world objects:

[1044] Users use the smart glasses' camera to capture real-world objects, send the image data to a server, which uses a generative AI model to search for products that match the real-world image, and send the search results to the user's device for display.

[1045] Specific examples

[1046] For example, if a user likes a pair of red sneakers they see in the park, they can use the smart glasses to capture an image of the sneakers and then use the application-generated image to search for similar items on an online shopping site. An example prompt for this is:

[1047] Example prompt sentence:

[1048] "Search for products similar to the red sneakers found in the park."

[1049] The above system and processing enable users to quickly use objects they find in the real world to search for products, making it easy to find products through intuitive operations.

[1050] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1051] Step 1: Enter and submit your search keywords

[1052] The user enters a keyword describing the desired product into the search bar of the device. Specifically, the user enters a word such as "red sneakers." The device generates a request to send the entered keyword to the server and sends it to the server. This input data is sent to the server as is.

[1053] Step 2: Generate an image

[1054] The server extracts search keywords from the incoming request. The extracted keywords are passed to a generative AI model (e.g., an image generation model). The generative AI model generates an image of a product associated with the keyword and returns the image to the server. The server then sends the generated image to the user's device. The input data is the search keyword, and the generated image is obtained as the output.

[1055] Step 3: View and modify the image

[1056] The user terminal displays the image received from the server on its display. The user checks the displayed image and makes any necessary modifications, such as changing colors or adding / deleting items. Once modifications are complete, the terminal sends the modified image to the server. The input data is the generated image, and the image after modifications made by the user are output.

[1057] Step 4: Transform to vector space

[1058] The server converts the received modified image into a vector space, where the image features are represented as numerical vectors. The input data is the modified image, and the output is a vector representation.

[1059] Step 5: Product Search

[1060] The server compares the converted vector with the image vectors stored in the product database to search for similar products. The input data is a vector representation, and similar products are listed. The output data is a list of products with high matching scores.

[1061] Step 6: Presenting search results

[1062] The user terminal displays a list of search results received from the server. The user can check the search results and select the products they like. The input data is the search results sent from the server, and is presented to the user as output.

[1063] Step 7: Submit your feedback

[1064] The user checks the search results and inputs their satisfaction feedback into the terminal, which then sends this feedback information to the server. The input data is the user's feedback, which is sent to the server as output.

[1065] Step 8: Improve your product database

[1066] The server analyzes the received user feedback and considers updating the product database and improving the generative AI model. This process aims to improve the accuracy of the system and user satisfaction. The input data is user feedback, and the output is an update to the product database.

[1067] Step 9: Search based on real-world objects

[1068] The user captures a real-world object using the camera on the smart glasses. The device then sends the captured image data to the server. The server then passes the received image data to a generative AI model, which generates an image. Based on the generated image data, a product database is searched for similar products. The input data is the captured real-world image data, and the output is a list of similar products.

[1069] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1070] This invention is a system that generates product images associated with search keywords entered by a user and uses those images to search for products. This system further improves the user experience by incorporating an emotion engine that recognizes the user's emotions and adjusts search results accordingly.

[1071] System configuration

[1072] The system consists of the following main components:

[1073] 1. Server

[1074] 2. User Device

[1075] 3. Generative AI Models

[1076] 4. Product Database

[1077] 5. Emotion Engine

[1078] Program processing

[1079] (1) Enter keywords and submit

[1080] The user enters a keyword describing the desired product into the device's search bar and presses the search button (e.g., the user enters "red sneakers").

[1081] The terminal generates a request to send the input keyword to the server, and sends it to the server.

[1082] (2) Creating an image

[1083] The server extracts keywords from the incoming request and passes them as input to the generative AI model.

[1084] The generative AI model generates an image of a product associated with the search keyword and returns that image to the server.

[1085] The server transmits the generated image to the user's terminal.

[1086] (3) Display and edit the image

[1087] The terminal displays the image received from the server on the user interface.

[1088] The user checks the displayed image and makes corrections as necessary (e.g., the user changes the color of the sneakers to dark red).

[1089] The terminal generates new image data that reflects the corrections and sends it to the server.

[1090] (4) Product search

[1091] The server converts the corrected image into a vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN).

[1092] The server then uses the converted vector to calculate the similarity between the image vectors in the product database, using techniques such as cosine similarity and Euclidean distance.

[1093] The server lists similar products and sends the information to the user's terminal.

[1094] (5) Presentation of search results

[1095] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[1096] (6) Emotion recognition and search result adjustment

[1097] The emotion engine analyzes the user's facial expressions and voice to identify their emotions. For example, it uses a camera and microphone to analyze the user's facial expressions and tone of voice and recognize emotions such as "happiness," "confusion," and "anger."

[1098] The server adjusts search results based on the user's emotions based on the output of the emotion engine. For example, if the user has a dissatisfied expression, the server changes the search results or adds new suggestions.

[1099] (7) Feedback and database improvement

[1100] The user provides feedback on the search results (e.g., "As expected," "Disappointing," etc.).

[1101] The terminal transmits the user's feedback information to the server.

[1102] The server analyzes the feedback it receives and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[1103] Specific examples

[1104] If a user types in "red sneakers" and the emotion engine detects an excited expression on the user's face, the server will prioritize search results that display sneakers with bolder designs to attract the user's attention.

[1105] If the user frowns at the search results, the emotion engine detects the emotion of dissatisfaction and the server will either readjust the search results or display additional product suggestions.

[1106] In this way, the present invention realizes a system that provides higher user satisfaction and search accuracy by searching for and suggesting products while taking into account the user's emotions.

[1107] The processing flow will be explained below.

[1108] Step 1:

[1109] The user enters a keyword describing the desired product into the device's search bar and presses the search button (e.g., the user enters "red sneakers").

[1110] Step 2:

[1111] The terminal generates a request to send the input keyword to the server, and sends it to the server.

[1112] Step 3:

[1113] The server extracts search keywords from the incoming request and passes them as input to the generative AI model.

[1114] Step 4:

[1115] The generative AI model generates an image of a product associated with the search keyword and returns that image to the server.

[1116] Step 5:

[1117] The server transmits the generated image to the user's terminal.

[1118] Step 6:

[1119] The terminal displays the image received from the server on the user interface.

[1120] Step 7:

[1121] The user checks the displayed image, makes corrections as necessary (e.g., the user changes the color of the sneakers to dark red), and confirms the corrected image on the terminal.

[1122] Step 8:

[1123] The terminal sends the corrected image to the server.

[1124] Step 9:

[1125] The server converts the corrected image into a vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN).

[1126] Step 10:

[1127] The server then uses the converted vector to calculate the similarity between the image vectors in the product database, using techniques such as cosine similarity and Euclidean distance.

[1128] Step 11:

[1129] The server lists similar products and sends the information to the user's terminal.

[1130] Step 12:

[1131] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[1132] Step 13:

[1133] The emotion engine analyzes the user's facial expressions and voice to identify their emotions. For example, it uses a camera and microphone to analyze the user's facial expressions and tone of voice and recognize emotions such as "happiness," "confusion," and "anger."

[1134] Step 14:

[1135] The server uses the output of the emotion engine to adjust search results according to the user's emotions. For example, if the user has a dissatisfied expression, the server can change search results or add new suggestions.

[1136] Step 15:

[1137] The user provides feedback on the search results (e.g., "As expected," "Disappointing," etc.).

[1138] Step 16:

[1139] The terminal transmits the user's feedback information to the server.

[1140] Step 17:

[1141] The server analyzes the feedback it receives and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[1142] This allows the user to perform optimal product searches based on search keywords, intuitive images, and even their own emotions.

[1143] Example 2

[1144] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1145] Conventional search systems search for products based on keywords entered by the user, but the search results often do not match the user's expectations or preferences. In addition, there is no function to adjust the search results taking into account the user's emotions, which leads to a poor user experience.

[1146] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for extracting input search keywords, generating an image based on the extracted search keywords, displaying the generated image on the user terminal, receiving means for receiving an image modified by the user, converting means for converting the modified image into a vector space, searching means for searching for similar products based on the converted vectors, transmitting means for transmitting search results to the user terminal, and emotion recognition means for recognizing the user's emotion and adjusting the search results in accordance with the emotion. This makes it possible to present more personalized search results that take the user's emotion into consideration.

[1147] A "search keyword" is a word that a user inputs to specify a product or information to be searched for.

[1148] The "generation means" is a means for generating an image of a related product based on the input search keyword.

[1149] The "display means" is a means for displaying the generated image diagram and search results on the user terminal.

[1150] The "receiving means" is a means for transmitting the image diagram modified by the user and feedback information to the server.

[1151] The "transformation means" is a means for transforming the corrected image into a vector space.

[1152] The "search means" is a means for searching for similar products based on vectorized image diagrams.

[1153] The "transmission means" is a means for transmitting search results to the user terminal.

[1154] The "emotion recognition means" is a means for recognizing the user's emotions and adjusting search results according to the emotions.

[1155] The "editing means" is a means for the user to modify the displayed image diagram.

[1156] The "update means" is a means for updating the product database based on user feedback.

[1157] A "user terminal" is an electronic device that a user uses to conduct a search.

[1158] A "server" is a central computer device that manages the processing of the entire system and includes various means.

[1159] The "product database" is a database for storing information and image data about products.

[1160] A "vector space" is a multidimensional space that expresses an image in numerical form.

[1161] A "generative AI model" is an artificial intelligence model that generates image images of related products from input keywords.

[1162]

[1163] This invention is a system that generates product images associated with search keywords entered by a user and uses these images to search for products. The system incorporates an emotion recognition function that recognizes the user's emotions and adjusts search results accordingly. The system consists of the following main components: a server, a user terminal, a generative AI model, a product database, and an emotion engine.

[1164] (System Components)

[1165] Enter keywords and send

[1166] The user enters a keyword that describes the desired product into the search bar of the device and presses the search button. The device generates a request to send the entered keyword to the server and sends it to the server.

[1167] Generate image diagrams

[1168] The server extracts keywords from the incoming request and passes them as input to the generative AI model. The generative AI model generates product images associated with the search keywords and returns the images to the server. The server then sends the generated images to the user's device.

[1169] Display and edit images

[1170] The terminal displays the image received from the server on the user interface. The user checks the displayed image and makes any necessary corrections (e.g., changing the color). The terminal generates new image data that reflects the corrections and sends it to the server.

[1171] Product search

[1172] The server converts the corrected image it receives into vector space. Specifically, it vectorizes the image using a model such as a convolutional neural network (CNN). Based on the converted vector, the server calculates the similarity with the image vectors in the product database. This uses techniques such as cosine similarity and Euclidean distance. The server then lists similar products and sends this information to the user's device.

[1173] Presenting search results

[1174] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[1175] Emotion recognition and search result tailoring

[1176] The emotion engine analyzes the user's facial expressions and voice to identify their emotions. For example, it uses a camera and microphone to analyze the user's facial expressions and tone of voice and recognizes emotions such as "happiness," "confusion," and "anger." The server then adjusts the search results based on the output of the emotion engine according to the user's emotions. For example, if the user looks dissatisfied, the server will change the search results or add new suggestions.

[1177] Feedback and database improvements

[1178] The user enters feedback on the search results (e.g., "As expected," "Disappointing," etc.). The device sends the feedback information to the server. The server analyzes the received feedback and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[1179] (Example)

[1180] Example 1: Search for red sneakers

[1181] When a user enters the keyword "red sneakers" and the emotion engine detects the user's excited expression, the server prioritizes sneakers with bold designs in the search results list to attract the user's interest. Example prompt for the generative AI model: "Generate an image of red sneakers with an exciting design."

[1182] Example 2: Readjustment with a dissatisfied expression

[1183] If the user frowns at the search results, the emotion engine detects the emotion of dissatisfaction. The server then refines the search results or displays additional product suggestions. Example prompt for the generative AI model: "Generate an image of a classic red sneaker design."

[1184] In this way, the present invention takes user emotions into consideration when searching and suggesting products, thereby providing higher user satisfaction and search accuracy.

[1185] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1186] Step 1:

[1187] The user enters a search keyword into the search bar. The user enters a keyword such as "red sneakers" into the search bar of the device and presses the "Search" button. This causes the device to generate an HTTP request containing the keyword and send it to the server (input: search keyword, output: HTTP request).

[1188] Step 2:

[1189] The server receives the HTTP request and extracts keywords, such as "red sneakers," from the received request (input: HTTP request, output: search keyword).

[1190] Step 3:

[1191] The server starts generating an image by inputting keywords into the generative AI model. The server passes the extracted keywords to the generative AI model, generates a prompt, and inputs it into the AI ​​model. For example, a prompt such as "Please generate an image of red sneakers" is used (input: search keyword, output: prompt).

[1192] Step 4:

[1193] The generative AI model generates an image of a pair of red sneakers based on the input prompt and returns the data to the server (input: prompt, output: image).

[1194] Step 5:

[1195] The server sends the generated image to the user's terminal. It then generates an HTTP response including the image and sends it to the user's terminal (input: image, output: HTTP response).

[1196] Step 6:

[1197] The terminal receives the image and displays it on the user interface. The image is displayed on the terminal screen so that the user can check it (input: HTTP response, output: displayed image).

[1198] Step 7:

[1199] The user modifies the image. The user modifies each element (color, shape, etc.) of the image displayed in the user interface, and generates modified data (input: displayed image, output: modified image data).

[1200] Step 8:

[1201] The terminal sends the corrected image data to the server, which then generates new data reflecting the corrections and sends it back to the server (input: corrected image data, output: HTTP request).

[1202] Step 9:

[1203] The server receives the corrected image data and converts it into vector space. The server uses a convolutional neural network (CNN) to vectorize the corrected image data (input: corrected image data, output: vector data).

[1204] Step 10:

[1205] The server uses the vector data to search for similar products. Based on the converted vector, it calculates the cosine similarity and Euclidean distance between the image vector and the product database, and generates a list of similar products (input: vector data, output: list of similar products).

[1206] Step 11:

[1207] The server sends the search results to the user's device. It then generates an HTTP response containing a list of similar products and sends it to the user's device (input: list of similar products, output: HTTP response).

[1208] Step 12:

[1209] The terminal receives the search results and displays them on the user interface. The list of products obtained as search results is displayed on the screen (input: HTTP response, output: displayed search results).

[1210] Step 13:

[1211] The emotion engine recognizes the user's emotions. It uses a camera and microphone to analyze the user's facial expressions and tone of voice to identify emotions such as "joy," "confusion," and "anger" (input: user's facial expression and voice data, output: emotion data).

[1212] Step 14:

[1213] The server adjusts the search results based on the emotion data. It receives the emotion analysis results, adjusts the search results as needed, and regenerates the search results including new suggestions (input: emotion data, output: adjusted search results).

[1214] Step 15:

[1215] The user enters feedback on the search results. The user enters feedback on the search results (e.g., "As expected," "Disappointing," etc.) (Input: displayed search results, Output: feedback).

[1216] Step 16:

[1217] The terminal sends user feedback to the server. It generates a request including the feedback information and sends it to the server (input: feedback, output: HTTP request).

[1218] Step 17:

[1219] The server receives feedback and improves the product database and generative AI model. Based on the received feedback, the server updates the database, adjusts the model, and improves the search algorithm (input: feedback, output: updated database and improved model).

[1220] As a result, the present invention can perform searches and suggestions that take the user's emotions into consideration, providing higher user satisfaction and search accuracy.

[1221] (Application example 2)

[1222] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1223] Conventional search systems have the problem of being unable to provide search results that take users' emotions into account. As a result, users may not get the results they expect, resulting in a decrease in satisfaction. Another problem is that they are unable to suggest visually appealing products, making it difficult to stimulate purchasing desire.

[1224] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for extracting input search keywords, generating an image based on the extracted search keywords, displaying the generated image on the user terminal, receiving means for receiving an image modified by the user, conversion means for converting the modified image into a vector space, search means for searching for similar products based on the converted vectors, emotion recognition means for recognizing the user's emotion and adjusting the search results based on the emotion, and transmission means for transmitting the search results to the user terminal. This makes it possible to provide search results that reflect the user's emotion, improving user satisfaction and stimulating purchasing desire.

[1225] The "means for extracting input search keywords" refers to a device or software that has the function of detecting and acquiring search keywords input by a user.

[1226] The "means for generating an image based on the extracted search keywords" refers to a device or software that has the function of generating an image of a related product using the acquired search keywords.

[1227] The "display means for displaying the generated image diagram on the user terminal" refers to a device or software that displays the generated image diagram in a form that can be viewed by the user.

[1228] The "receiving means for receiving an image diagram modified by a user" refers to a device or software having a function for receiving data of an image diagram modified by a user.

[1229] The "conversion means for converting the corrected image into vector space" refers to a device or software that has the function of converting the image corrected by the user into vector data.

[1230] The "search means for searching for similar products based on the converted vectors" refers to a device or software that has the function of searching for similar products in a product database based on the vectorized image data.

[1231] "Emotion recognition means that recognizes a user's emotions and adjusts search results based on those emotions" refers to a device or software that has the function of analyzing emotions from a user's facial expressions, voice, etc., and changing or optimizing search results based on the results.

[1232] The "transmission means for transmitting search results to the user terminal" refers to a device or software that has the function of transmitting the search results to the user terminal.

[1233] System configuration

[1234] The system consists of the following main components:

[1235] 1. Server

[1236] 2. User Device

[1237] 3. Generative AI Models

[1238] 4. Product Database

[1239] 5. Emotion Recognition Engine

[1240] Program Overview

[1241] The system of the present invention allows users to input search keywords and search for products using images generated based on those keywords. It also includes a function to recognize user emotions and adjust search results accordingly.

[1242] Hardware and Software

[1243] Servers, user devices, generative AI models, etc. Specific examples include Python's PIL, Transformers (Huggingface), emotion_recognition library, and CLIP model.

[1244] Data processing and calculation

[1245] The server extracts the keywords entered by the user and generates an image using a generative AI model based on the keywords. The generated image is then sent to the user's device.

[1246] The user terminal displays the image received from the server, and the user can modify it as necessary. The modified image is then sent back to the server.

[1247] The server converts the corrected image into a vector space and searches for similar products in a product database based on the vectors. The search results are sent to the user's terminal.

[1248] The emotion recognition engine analyzes a user's facial expressions and voice to identify emotions and adjust search results accordingly.

[1249] Specific processing examples

[1250] For example, if a user enters the keyword "red sneakers," the system uses a generative AI model to generate a related image and display it to the user. The user can then modify the color of the sneakers to a darker red, and the modified image is sent to the server. The server then converts the modified image into a vector space and searches for similar products in the product database.

[1251] When search results are displayed on the user's device, an emotion recognition engine analyzes the user's facial expressions and voice, and the search results are adjusted based on their emotions. For example, if the user expresses dissatisfaction, the system can add other suggestions.

[1252] Recommended prompt examples

[1253] Keywords: "red sneakers"

[1254] Feedback: "User's Face Photo"

[1255] In this way, the present invention realizes higher user satisfaction and search accuracy by searching and proposing products taking into account the user's emotions.

[1256] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1257] Step 1:

[1258] The user enters a keyword describing a product into the search bar of the device and presses the search button. When "red sneakers" is entered as input, the device generates a request to send this keyword to the server and sends it to the server. The input is the keyword "red sneakers" and the output is the transmission of the request.

[1259] Step 2:

[1260] The server extracts keywords from the incoming request and passes them as input to the generative AI model. The model generates an image of the product from the keywords and returns the image data to the server. The input is the search keyword "red sneakers," and the output is the generated image data.

[1261] Step 3:

[1262] The server sends the generated image to the user's terminal. The input is the image data, and the output is the transmission of the image to the user's terminal.

[1263] Step 4:

[1264] The user terminal displays the image received from the server on the user interface. The user checks the displayed image and makes corrections, such as changing the color of the sneakers to dark red. The input is image data, and the output is a display containing the corrected image.

[1265] Step 5:

[1266] The user terminal generates new image data that reflects the modifications and sends it to the server. The input is the modified image, and the output is the transmission of the modified image data to the server.

[1267] Step 6:

[1268] The server converts the corrected image received into vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN). The input is the corrected image data, and the output is vector data.

[1269] Step 7:

[1270] The server uses the converted vector data to calculate the similarity with the image vectors in the product database. This uses techniques such as cosine similarity and Euclidean distance. The input is vector data, and the output is a list of similar products.

[1271] Step 8:

[1272] The server lists similar products and sends the information to the user terminal. The input is the list of similar products, and the output is the transmission of similar product information to the user terminal.

[1273] Step 9:

[1274] The user terminal displays the search results received from the server in a list on the user interface. The input is similar product information, and the output is the display to the user.

[1275] Step 10:

[1276] The emotion recognition engine analyzes the user's facial expressions and voice to identify emotions. Using a camera and microphone, the engine analyzes the user's facial expressions and tone of voice to recognize emotions such as "happiness," "confusion," and "anger." The input is the user's facial expression and voice data, and the output is identified emotional data.

[1277] Step 11:

[1278] The server adjusts search results based on the output of the emotion recognition engine according to the user's emotions. For example, if the user has a dissatisfied expression, the server changes the search results or adds new suggestions. The input is emotion data and search results, and the output is the adjusted search results.

[1279] Step 12:

[1280] The user inputs feedback on the search results, such as "as expected" or "disappointing," and the device sends this feedback information to the server. The input is user feedback, and the output is sending feedback information to the server.

[1281] Step 13:

[1282] The server analyzes the received feedback and updates the product database, adjusts the generative AI model, and improves the search algorithm. The input is the feedback information, and the output is improved system performance.

[1283] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1284] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1285] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1286] [Fourth embodiment]

[1287] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1288] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1289] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1290] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1291] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1292] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1293] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1294] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1295] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1296] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1297] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1298] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1299] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1300] This invention is a system that generates an image of a product associated with a user's search keywords and uses that image to search for products. This system is realized by linking a server and a user terminal.

[1301] System configuration

[1302] The system consists of the following main components:

[1303] 1. Server

[1304] 2. User Device

[1305] 3. Generative AI Models

[1306] 4. Product Database

[1307] Program processing

[1308] (1) Enter keywords and submit

[1309] User: Enters a keyword describing the desired product into the device's search bar (e.g., user enters "red sneakers").

[1310] Terminal: Generates a request to send the entered keyword to the server and sends it to the server.

[1311] (2) Creating an image

[1312] Server: Extract keywords from the incoming request.

[1313] Server: Passes the extracted keywords to a generative AI model (e.g., image generation model).

[1314] Generative AI model: Generates image images of products associated with keywords and returns them to the server.

[1315] Server: Sends the generated image to the user's device.

[1316] (3) Display and edit the image

[1317] Terminal: Displays the image received from the server on the user interface.

[1318] User: Check the displayed image and make any necessary modifications (e.g., change the color, add / delete items, etc.) (e.g., the user changes the color of the sneakers to dark red).

[1319] Terminal: Generate new image data that reflects the modifications and send it to the server.

[1320] (4) Product search

[1321] Server: Converts the received corrected image into vector space.

[1322] Server: Based on the converted vector, it compares it with the image vectors in the product database and extracts similar products.

[1323] Server: Lists products with high matching scores and sends this information to the user's device.

[1324] (5) Search result presentation and feedback

[1325] Terminal: The search results received from the server are displayed in a user interface (e.g., three red sneakers are displayed as options).

[1326] User: Review search results and enter satisfaction feedback (e.g., "As expected," "Disappointing," etc.) into the device.

[1327] Device: Sends feedback information to the server.

[1328] (6) Database Improvement

[1329] Server: Analyzes the received user feedback.

[1330] Server: Based on the feedback, consider ways to update and improve the product database (e.g., add new product data, modify existing data, or tune the generative AI model or search algorithm).

[1331] In this way, the present invention goes beyond the limitations of conventional keyword searches and provides a means for users to intuitively and effectively search for the products they want. Through specific steps, we achieve improved search quality and increased user satisfaction.

[1332] The processing flow will be explained below.

[1333] Step 1:

[1334] The user enters a keyword describing the desired product into the device's search bar and presses the search button (e.g., the user enters "red sneakers").

[1335] Step 2:

[1336] The terminal generates a request to send the input keyword to the server, and sends it to the server.

[1337] Step 3:

[1338] The server extracts search keywords from the incoming request and passes them as input to the generative AI model.

[1339] Step 4:

[1340] The generative AI model generates an image of a product associated with the search keyword and returns that image to the server.

[1341] Step 5:

[1342] The server transmits the generated image to the user's terminal.

[1343] Step 6:

[1344] The terminal displays the image received from the server on the user interface.

[1345] Step 7:

[1346] The user checks the displayed image, makes corrections as necessary (e.g., the user changes the color of the sneakers to dark red), and confirms the corrected image on the terminal.

[1347] Step 8:

[1348] The terminal sends the corrected image to the server.

[1349] Step 9:

[1350] The server converts the corrected image into a vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN).

[1351] Step 10:

[1352] The server then uses the converted vector to calculate the similarity between the image vectors in the product database, using techniques such as cosine similarity and Euclidean distance.

[1353] Step 11:

[1354] The server lists similar products and sends the information to the user's terminal.

[1355] Step 12:

[1356] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[1357] Step 13:

[1358] The user reviews the search results and enters their satisfaction feedback (e.g., "As expected," "Disappointed," etc.).

[1359] Step 14:

[1360] The terminal transmits the user's feedback information to the server.

[1361] Step 15:

[1362] The server analyzes the feedback it receives and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[1363] Example 1

[1364] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1365] Conventional keyword search systems have difficulty reflecting the specific image of the product a user is looking for, resulting in low search accuracy. Furthermore, they lack the ability for users to visually modify their image or provide feedback, which often results in search results that do not meet user expectations. Furthermore, there is a lack of a way to continuously improve the product database using user feedback.

[1366] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1367] In this invention, the server includes means for extracting input search keywords, generating means for generating an image based on the extracted search keywords, display means for displaying the generated image on a user terminal, receiving means for receiving an image modified by the user, conversion means for converting the modified image into a vector space, search means for searching for similar products based on the converted vectors, transmission means for transmitting search results to the user terminal, update means for receiving user feedback and updating the product database, and means for passing the input search keywords to the generative AI model as a prompt sentence and receiving an image generated based on the prompt sentence. This allows users to search for products based on a visual image, and enables search accuracy to be improved through corrections and feedback.

[1368] The "means for extracting input search keywords" refers to a method or device for extracting search keywords input by a user to a terminal.

[1369] The "means for generating an image diagram based on the extracted search keywords" refers to a method or device for generating a related image diagram based on the extracted keywords.

[1370] The "display means for displaying the generated image diagram on the user terminal" refers to a method or device for displaying the generated image diagram on the screen of the user's terminal.

[1371] The "receiving means for receiving an image diagram modified by a user" refers to a method or device for receiving image diagram data modified by a user.

[1372] The "conversion means for converting the corrected image into a vector space" is a method or device for converting the image corrected by the user into a numerical vector.

[1373] The "search means for searching for similar products based on the converted vectors" refers to a method or device for searching a database for similar products using the converted vector data.

[1374] The "transmission means for transmitting search results to the user terminal" refers to a method or device for transmitting information about the searched products to the user terminal.

[1375] The "means for receiving user feedback and updating the product database" refers to a method or device for receiving feedback information from users and updating the contents of the product database.

[1376] "Means for passing search keywords input to a generative AI model as a prompt sentence and receiving an image diagram generated based on the prompt sentence" refers to a method or device for passing search keywords as a prompt sentence to a generative AI model and receiving the resulting image diagram.

[1377] This invention is a system that generates product image images associated with a user's search keywords and uses the images to search for products. This system is implemented by a server and a user terminal, and utilizes a generative AI model and a product database.

[1378] Hardware and software used

[1379] 1. Server:

[1380] Use a high performance computer server.

[1381] The software uses DeepAI's generative AI model and TensorFlow for image generation and data processing.

[1382] The search algorithm uses Elasticsearch and other search engines.

[1383] 2. User Device:

[1384] A personal computer or smartphone is used for user operation.

[1385] The user interface is built using technologies such as HTML5, JavaScript, and CSS.

[1386] Specific processing of the program

[1387] Enter keywords and send

[1388] The user enters a keyword that describes the desired product into the search bar of the device. For example, the user enters "red sneakers." In response to this input, the device sends a request including the corresponding keyword to the server.

[1389] Generate image diagrams

[1390] The server receives the request sent from the device and extracts the search keywords. The extracted keywords are passed to the generative AI model as a prompt. The generative AI model generates an image based on the prompt, for example, "red sneakers," and returns it to the server. The server then sends the generated image to the user's device.

[1391] Display and edit images

[1392] The terminal displays the image received from the server on the user interface. The user can check the displayed image and modify it as necessary. For example, the user may change the color of the sneakers to dark red. Once the modification is made, the terminal transmits the modified image data to the server.

[1393] Product search

[1394] The server converts the corrected image into vector space using machine learning libraries such as TensorFlow. The converted vector is compared with other image vectors in a product database. For example, the server uses Elasticsearch to calculate similarity and extract similar products. The server then lists products with high similarity and sends the search results to the user's device.

[1395] Search results and feedback

[1396] The terminal displays the search result product list received from the server on a user interface. For example, three red sneaker options are displayed. The user can review the search results and enter feedback regarding satisfaction. The terminal then sends this feedback information to the server.

[1397] Database improvements

[1398] The server analyzes the feedback received from users. A machine learning library (e.g., scikit-learn) is used for the analysis. The server updates the product database based on the analysis results and also tunes the generative AI model and search algorithm. This makes it possible to continuously improve search accuracy.

[1399] Examples of concrete examples and prompts

[1400] As a concrete example, consider the case where a user enters the search keyword "red sneakers." The following prompt is passed to the generative AI model:

[1401] "Generate an image of red sneakers."

[1402] "Create an image of a pair of dark red sneakers."

[1403] Based on this prompt, the generative AI model generates an image, which the user can then review and modify, and then search for and present more similar products.

[1404] In this way, the present invention provides users with the ability to visually search for products and provides a means for continually improving the search system based on feedback.

[1405] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1406] Step 1:

[1407] A user enters a keyword that describes the desired product into the search bar of the device. For example, the user enters "red sneakers." The device captures the entered keyword and generates an HTTP request to send to the server. The input is the text data "red sneakers," and the output is the request data to the server.

[1408] Step 2:

[1409] The server receives the HTTP request sent from the device. Next, the server extracts the keyword "red sneakers" from the request. It generates a prompt based on this extracted keyword and sends it to the generative AI model. The input is the HTTP request data, and the output is the prompt to the generative AI model.

[1410] Step 3:

[1411] The generative AI model generates an image based on the received prompt "red sneakers." In this process, the generative AI model uses its internal algorithm to compose relevant image data. The generated image is then returned to the server. The input is the prompt, and the output is the generated image.

[1412] Step 4:

[1413] The server receives the image diagram returned from the generative AI model. It encodes the received image diagram in an appropriate format and generates an HTTP response to send to the user device. The input is the generated image diagram, and the output is an HTTP response containing the image diagram data to the user device.

[1414] Step 5:

[1415] The terminal receives the HTTP response from the server, decodes the image data, and displays it on the user interface. The user checks the displayed image and makes any necessary corrections. For example, the user changes the color of the sneakers to dark red. The input is the response from the server, and the output is the corrected image.

[1416] Step 6:

[1417] The terminal generates image data modified by the user and generates an HTTP request to send it to the server. The input is the modified image, and the output is an HTTP request including the modified image data to the server.

[1418] Step 7:

[1419] The server receives an HTTP request from the device and retrieves the corrected image. The server then converts the corrected image into a vector space using a machine learning library such as TensorFlow. The input is the corrected image, and the output is vector data.

[1420] Step 8:

[1421] The server uses the vector data to compare it with other image vectors in the product database and calculates the similarity. For example, it uses distance calculations such as cosine similarity. It extracts products with high similarity and lists them. The input is vector data, and the output is a list of similar products.

[1422] Step 9:

[1423] The server generates an HTTP response to send the list of similar products to the user terminal. The input is the list of similar products, and the output is the response to the user terminal.

[1424] Step 10:

[1425] The terminal receives the response from the server and displays similar products on the user interface. For example, three red sneakers are displayed. The user checks the displayed products and inputs feedback on their satisfaction into the terminal. The input is the response from the server, and the output is the user's feedback.

[1426] Step 11:

[1427] The terminal sends an HTTP request containing the user's feedback to the server. The input is the user's feedback data, and the output is the request to the server.

[1428] Step 12:

[1429] The server receives feedback from the device and analyzes the data to tune the product database, generative AI model, and search algorithm. Machine learning libraries such as scikit-learn are used for the analysis. The input is user feedback data, and the output is an updated and improved database and model.

[1430] (Application example 1)

[1431] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1432] Conventional keyword search systems have the problem that it is difficult for users to intuitively search for products. Also, there is no way to quickly search for an object found in the real world on an online shopping site. This makes it difficult for users to easily find the product they want, resulting in a decrease in satisfaction.

[1433] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1434] In this invention, the server includes means for extracting input search keywords, generating means for generating an image based on the extracted search keywords, display means for displaying the generated image on the user terminal, receiving means for receiving the image modified by the user, conversion means for converting the modified image into a vector space, search means for searching for similar products based on the converted vectors, transmission means for transmitting search results to the user terminal, and real-world image search means for receiving real-world image data captured by a camera of the user terminal and searching for products based on the real-world image data. This allows the user to quickly use objects found in the real world in product searches and easily find products with intuitive operations.

[1435] A "search keyword" is a word or phrase that a user enters to search for a product.

[1436] The "generation means" refers to a device or software that has the function of generating an image based on the input search keywords.

[1437] A "user terminal" is a device used by a user, including a smartphone, tablet, smart glasses, etc.

[1438] The "display means" is a device or software that has the function of displaying the generated image diagram on the screen of the user terminal.

[1439] The "receiving means" is a device or software that has the function of receiving an image modified by a user.

[1440] The "conversion means" is a device or software that has the function of converting the modified image into vector space.

[1441] The "search means" is a device or software that has the function of searching a product database for similar products based on the converted vector.

[1442] The "transmission means" is a device or software that has the function of transmitting search results to a user terminal.

[1443] The "real world image search means" is a device or software that has the function of receiving real world image data captured by the camera of the user terminal and searching for products based on that image data.

[1444] This invention provides a system that generates product images based on user search keywords and real-world images, and then searches for products based on the images. The details of the system are described below.

[1445] System configuration

[1446] The system consists of the following main components:

[1447] 1. Server

[1448] 2. User Device

[1449] 3. Generative AI Models

[1450] 4. Product Database

[1451] Hardware and Software Use

[1452] Server: A high-performance computer that processes data and runs AI models.

[1453] User terminal: A smartphone, tablet, smart glasses, or other mobile device with a camera and display.

[1454] Generative AI models: Examples include image generation models such as Stable Diffusion and DALL-E.

[1455] Product database: A database for registering and managing product information.

[1456] Data processing and calculation

[1457] Extract and submit search keywords:

[1458] A user enters keywords into the search bar using smart glasses or a smartphone, and the user device sends the entered keywords to the server.

[1459] Generate image diagrams:

[1460] The server extracts the received search keywords, passes them to the generative AI model, and generates an image. The generated image is then sent to the user's device.

[1461] View and modify images:

[1462] The user terminal displays the image received from the server, and the user can make any necessary corrections. The corrected image is then sent back to the server. This series of operations is performed through the user interface.

[1463] Transformation to vector space and product search:

[1464] The server converts the corrected image into a vector space and compares this vector with the image vectors in the product database to extract similar products. The extracted product list is sent to the user's terminal and displayed in a selectable format.

[1465] Search based on real-world objects:

[1466] Users use the smart glasses' camera to capture real-world objects, send the image data to a server, which uses a generative AI model to search for products that match the real-world image, and send the search results to the user's device for display.

[1467] Specific examples

[1468] For example, if a user likes a pair of red sneakers they see in the park, they can use the smart glasses to capture an image of the sneakers and then use the application-generated image to search for similar items on an online shopping site. An example prompt for this is:

[1469] Example prompt sentence:

[1470] "Search for products similar to the red sneakers found in the park."

[1471] The above system and processing enable users to quickly use objects they find in the real world to search for products, making it easy to find products through intuitive operations.

[1472] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1473] Step 1: Enter and submit your search keywords

[1474] The user enters a keyword describing the desired product into the search bar of the device. Specifically, the user enters a word such as "red sneakers." The device generates a request to send the entered keyword to the server and sends it to the server. This input data is sent to the server as is.

[1475] Step 2: Generate an image

[1476] The server extracts search keywords from the incoming request. The extracted keywords are passed to a generative AI model (e.g., an image generation model). The generative AI model generates an image of a product associated with the keyword and returns the image to the server. The server then sends the generated image to the user's device. The input data is the search keyword, and the generated image is obtained as the output.

[1477] Step 3: View and modify the image

[1478] The user terminal displays the image received from the server on its display. The user checks the displayed image and makes any necessary modifications, such as changing colors or adding / deleting items. Once modifications are complete, the terminal sends the modified image to the server. The input data is the generated image, and the image after modifications made by the user are output.

[1479] Step 4: Transform to vector space

[1480] The server converts the received modified image into a vector space, where the image features are represented as numerical vectors. The input data is the modified image, and the output is a vector representation.

[1481] Step 5: Product Search

[1482] The server compares the converted vector with the image vectors stored in the product database to search for similar products. The input data is a vector representation, and similar products are listed. The output data is a list of products with high matching scores.

[1483] Step 6: Presenting search results

[1484] The user terminal displays a list of search results received from the server. The user can check the search results and select the products they like. The input data is the search results sent from the server, and is presented to the user as output.

[1485] Step 7: Submit your feedback

[1486] The user checks the search results and inputs their satisfaction feedback into the terminal, which then sends this feedback information to the server. The input data is the user's feedback, which is sent to the server as output.

[1487] Step 8: Improve your product database

[1488] The server analyzes the received user feedback and considers updating the product database and improving the generative AI model. This process aims to improve the accuracy of the system and user satisfaction. The input data is user feedback, and the output is an update to the product database.

[1489] Step 9: Search based on real-world objects

[1490] The user captures a real-world object using the camera on the smart glasses. The device then sends the captured image data to the server. The server then passes the received image data to a generative AI model, which generates an image. Based on the generated image data, a product database is searched for similar products. The input data is the captured real-world image data, and the output is a list of similar products.

[1491] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1492] This invention is a system that generates product images associated with search keywords entered by a user and uses those images to search for products. This system further improves the user experience by incorporating an emotion engine that recognizes the user's emotions and adjusts search results accordingly.

[1493] System configuration

[1494] The system consists of the following main components:

[1495] 1. Server

[1496] 2. User Device

[1497] 3. Generative AI Models

[1498] 4. Product Database

[1499] 5. Emotion Engine

[1500] Program processing

[1501] (1) Enter keywords and submit

[1502] The user enters a keyword describing the desired product into the device's search bar and presses the search button (e.g., the user enters "red sneakers").

[1503] The terminal generates a request to send the input keyword to the server, and sends it to the server.

[1504] (2) Creating an image

[1505] The server extracts keywords from the incoming request and passes them as input to the generative AI model.

[1506] The generative AI model generates an image of a product associated with the search keyword and returns that image to the server.

[1507] The server transmits the generated image to the user's terminal.

[1508] (3) Display and edit the image

[1509] The terminal displays the image received from the server on the user interface.

[1510] The user checks the displayed image and makes corrections as necessary (e.g., the user changes the color of the sneakers to dark red).

[1511] The terminal generates new image data that reflects the corrections and sends it to the server.

[1512] (4) Product search

[1513] The server converts the corrected image into a vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN).

[1514] The server then uses the converted vector to calculate the similarity between the image vectors in the product database, using techniques such as cosine similarity and Euclidean distance.

[1515] The server lists similar products and sends the information to the user's terminal.

[1516] (5) Presentation of search results

[1517] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[1518] (6) Emotion recognition and search result adjustment

[1519] The emotion engine analyzes the user's facial expressions and voice to identify their emotions. For example, it uses a camera and microphone to analyze the user's facial expressions and tone of voice and recognize emotions such as "happiness," "confusion," and "anger."

[1520] The server uses the output of the emotion engine to adjust search results according to the user's emotions. For example, if the user has a dissatisfied expression, the server can change search results or add new suggestions.

[1521] (7) Feedback and database improvement

[1522] The user provides feedback on the search results (e.g., "As expected," "Disappointing," etc.).

[1523] The terminal transmits the user's feedback information to the server.

[1524] The server analyzes the feedback it receives and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[1525] Specific examples

[1526] If a user types in "red sneakers" and the emotion engine detects an excited expression on the user's face, the server will prioritize search results that display sneakers with bolder designs to attract the user's attention.

[1527] If the user frowns at the search results, the emotion engine detects the emotion of dissatisfaction and the server will either readjust the search results or display additional product suggestions.

[1528] In this way, the present invention realizes a system that provides higher user satisfaction and search accuracy by searching for and suggesting products while taking into account the user's emotions.

[1529] The processing flow will be explained below.

[1530] Step 1:

[1531] The user enters a keyword describing the desired product into the device's search bar and presses the search button (e.g., the user enters "red sneakers").

[1532] Step 2:

[1533] The terminal generates a request to send the input keyword to the server, and sends it to the server.

[1534] Step 3:

[1535] The server extracts search keywords from the incoming request and passes them as input to the generative AI model.

[1536] Step 4:

[1537] The generative AI model generates an image of a product associated with the search keyword and returns that image to the server.

[1538] Step 5:

[1539] The server transmits the generated image to the user's terminal.

[1540] Step 6:

[1541] The terminal displays the image received from the server on the user interface.

[1542] Step 7:

[1543] The user checks the displayed image, makes corrections as necessary (e.g., the user changes the color of the sneakers to dark red), and confirms the corrected image on the terminal.

[1544] Step 8:

[1545] The terminal sends the corrected image to the server.

[1546] Step 9:

[1547] The server converts the corrected image into a vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN).

[1548] Step 10:

[1549] The server then uses the converted vector to calculate the similarity between the image vectors in the product database, using techniques such as cosine similarity and Euclidean distance.

[1550] Step 11:

[1551] The server lists similar products and sends the information to the user's terminal.

[1552] Step 12:

[1553] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[1554] Step 13:

[1555] The emotion engine analyzes the user's facial expressions and voice to identify their emotions. For example, it uses a camera and microphone to analyze the user's facial expressions and tone of voice and recognize emotions such as "happiness," "confusion," and "anger."

[1556] Step 14:

[1557] The server uses the output of the emotion engine to adjust search results according to the user's emotions. For example, if the user has a dissatisfied expression, the server can change search results or add new suggestions.

[1558] Step 15:

[1559] The user provides feedback on the search results (e.g., "As expected," "Disappointing," etc.).

[1560] Step 16:

[1561] The terminal transmits the user's feedback information to the server.

[1562] Step 17:

[1563] The server analyzes the feedback it receives and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[1564] This allows the user to perform optimal product searches based on search keywords, intuitive images, and even their own emotions.

[1565] Example 2

[1566] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1567] Conventional search systems search for products based on keywords entered by the user, but the search results often do not match the user's expectations or preferences. In addition, there is no function to adjust the search results taking into account the user's emotions, which leads to a poor user experience.

[1568] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for extracting input search keywords, generating an image based on the extracted search keywords, displaying the generated image on the user terminal, receiving means for receiving an image modified by the user, converting means for converting the modified image into a vector space, searching means for searching for similar products based on the converted vectors, transmitting means for transmitting search results to the user terminal, and emotion recognition means for recognizing the user's emotion and adjusting the search results in accordance with the emotion. This makes it possible to present more personalized search results that take the user's emotion into consideration.

[1569] A "search keyword" is a word that a user inputs to specify a product or information to be searched for.

[1570] The "generation means" is a means for generating an image of a related product based on the input search keyword.

[1571] The "display means" is a means for displaying the generated image diagram and search results on the user terminal.

[1572] The "receiving means" is a means for transmitting the image diagram modified by the user and feedback information to the server.

[1573] The "transformation means" is a means for transforming the corrected image into a vector space.

[1574] The "search means" is a means for searching for similar products based on vectorized image diagrams.

[1575] The "transmission means" is a means for transmitting search results to the user terminal.

[1576] The "emotion recognition means" is a means for recognizing the user's emotions and adjusting search results according to the emotions.

[1577] The "editing means" is a means for the user to modify the displayed image diagram.

[1578] The "update means" is a means for updating the product database based on user feedback.

[1579] A "user terminal" is an electronic device that a user uses to conduct a search.

[1580] A "server" is a central computer device that manages the processing of the entire system and includes various means.

[1581] The "product database" is a database for storing information and image data about products.

[1582] A "vector space" is a multidimensional space that expresses an image in numerical form.

[1583] A "generative AI model" is an artificial intelligence model that generates image images of related products from input keywords.

[1584]

[1585] This invention is a system that generates product images associated with search keywords entered by a user and uses these images to search for products. The system incorporates an emotion recognition function that recognizes the user's emotions and adjusts search results accordingly. The system consists of the following main components: a server, a user terminal, a generative AI model, a product database, and an emotion engine.

[1586] (System Components)

[1587] Enter keywords and send

[1588] The user enters a keyword that describes the desired product into the search bar of the device and presses the search button. The device generates a request to send the entered keyword to the server and sends it to the server.

[1589] Generate image diagrams

[1590] The server extracts keywords from the incoming request and passes them as input to the generative AI model. The generative AI model generates product images associated with the search keywords and returns the images to the server. The server then sends the generated images to the user's device.

[1591] Display and edit images

[1592] The terminal displays the image received from the server on the user interface. The user checks the displayed image and makes any necessary corrections (e.g., changing the color). The terminal generates new image data that reflects the corrections and sends it to the server.

[1593] Product search

[1594] The server converts the corrected image it receives into vector space. Specifically, it vectorizes the image using a model such as a convolutional neural network (CNN). Based on the converted vector, the server calculates the similarity with the image vectors in the product database. This uses techniques such as cosine similarity and Euclidean distance. The server then lists similar products and sends this information to the user's device.

[1595] Presenting search results

[1596] The terminal displays a list of search results received from the server in a user interface (e.g., multiple options for red sneakers are displayed).

[1597] Emotion recognition and search result tailoring

[1598] The emotion engine analyzes the user's facial expressions and voice to identify their emotions. For example, it uses a camera and microphone to analyze the user's facial expressions and tone of voice and recognizes emotions such as "happiness," "confusion," and "anger." The server then adjusts the search results based on the output of the emotion engine according to the user's emotions. For example, if the user looks dissatisfied, the server will change the search results or add new suggestions.

[1599] Feedback and database improvements

[1600] The user enters feedback on the search results (e.g., "As expected," "Disappointing," etc.). The device sends the feedback information to the server. The server analyzes the received feedback and updates the product database, adjusts the generative AI model, and improves the search algorithm.

[1601] (Example)

[1602] Example 1: Search for red sneakers

[1603] When a user enters the keyword "red sneakers" and the emotion engine detects the user's excited expression, the server prioritizes sneakers with bold designs in the search results list to attract the user's interest. Example prompt for the generative AI model: "Generate an image of red sneakers with an exciting design."

[1604] Example 2: Readjustment with a dissatisfied expression

[1605] If the user frowns at the search results, the emotion engine detects the emotion of dissatisfaction. The server can then refine the search results or display additional product suggestions. Example prompt for the generative AI model: "Generate an image of a classic red sneaker design."

[1606] In this way, the present invention takes user emotions into consideration when searching and suggesting products, thereby providing higher user satisfaction and search accuracy.

[1607] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1608] Step 1:

[1609] The user enters a search keyword into the search bar. The user enters a keyword such as "red sneakers" into the search bar of the device and presses the "Search" button. This causes the device to generate an HTTP request containing the keyword and send it to the server (input: search keyword, output: HTTP request).

[1610] Step 2:

[1611] The server receives the HTTP request and extracts keywords, such as "red sneakers," from the received request (input: HTTP request, output: search keyword).

[1612] Step 3:

[1613] The server starts generating an image by inputting keywords into the generative AI model. The server passes the extracted keywords to the generative AI model, generates a prompt, and inputs it into the AI ​​model. For example, a prompt such as "Please generate an image of red sneakers" is used (input: search keyword, output: prompt).

[1614] Step 4:

[1615] The generative AI model generates an image of a pair of red sneakers based on the input prompt and returns the data to the server (input: prompt, output: image).

[1616] Step 5:

[1617] The server sends the generated image to the user's terminal. It then generates an HTTP response including the image and sends it to the user's terminal (input: image, output: HTTP response).

[1618] Step 6:

[1619] The terminal receives the image and displays it on the user interface. The image is displayed on the terminal screen so that the user can check it (input: HTTP response, output: displayed image).

[1620] Step 7:

[1621] The user modifies the image. The user modifies each element (color, shape, etc.) of the image displayed in the user interface, and generates modified data (input: displayed image, output: modified image data).

[1622] Step 8:

[1623] The terminal sends the corrected image data to the server, which then generates new data reflecting the corrections and sends it back to the server (input: corrected image data, output: HTTP request).

[1624] Step 9:

[1625] The server receives the corrected image data and converts it into vector space. The server uses a convolutional neural network (CNN) to vectorize the corrected image data (input: corrected image data, output: vector data).

[1626] Step 10:

[1627] The server uses the vector data to search for similar products. Based on the converted vector, it calculates the cosine similarity and Euclidean distance between the image vector and the product database, and generates a list of similar products (input: vector data, output: list of similar products).

[1628] Step 11:

[1629] The server sends the search results to the user's device. It then generates an HTTP response containing a list of similar products and sends it to the user's device (input: list of similar products, output: HTTP response).

[1630] Step 12:

[1631] The terminal receives the search results and displays them on the user interface. The list of products obtained as search results is displayed on the screen (input: HTTP response, output: displayed search results).

[1632] Step 13:

[1633] The emotion engine recognizes the user's emotions. It uses a camera and microphone to analyze the user's facial expressions and tone of voice to identify emotions such as "joy," "confusion," and "anger" (input: user's facial expression and voice data, output: emotion data).

[1634] Step 14:

[1635] The server adjusts the search results based on the emotion data. It receives the emotion analysis results, adjusts the search results as needed, and regenerates the search results including new suggestions (input: emotion data, output: adjusted search results).

[1636] Step 15:

[1637] The user enters feedback on the search results. The user enters feedback on the search results (e.g., "As expected," "Disappointing," etc.) (Input: displayed search results, Output: feedback).

[1638] Step 16:

[1639] The terminal sends user feedback to the server. It generates a request including the feedback information and sends it to the server (input: feedback, output: HTTP request).

[1640] Step 17:

[1641] The server receives feedback and improves the product database and generative AI model. Based on the received feedback, the server updates the database, adjusts the model, and improves the search algorithm (input: feedback, output: updated database and improved model).

[1642] As a result, the present invention can perform searches and suggestions that take the user's emotions into consideration, providing higher user satisfaction and search accuracy.

[1643] (Application example 2)

[1644] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1645] Conventional search systems have the problem of being unable to provide search results that take users' emotions into account. As a result, users may not get the results they expect, resulting in a decrease in satisfaction. Another problem is that they are unable to suggest visually appealing products, making it difficult to stimulate purchasing desire.

[1646] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for extracting input search keywords, generating an image based on the extracted search keywords, displaying the generated image on the user terminal, receiving means for receiving an image modified by the user, conversion means for converting the modified image into a vector space, search means for searching for similar products based on the converted vectors, emotion recognition means for recognizing the user's emotion and adjusting the search results based on the emotion, and transmission means for transmitting the search results to the user terminal. This makes it possible to provide search results that reflect the user's emotion, improving user satisfaction and stimulating purchasing desire.

[1647] The "means for extracting input search keywords" refers to a device or software that has the function of detecting and acquiring search keywords input by a user.

[1648] The "means for generating an image based on the extracted search keywords" refers to a device or software that has the function of generating an image of a related product using the acquired search keywords.

[1649] The "display means for displaying the generated image diagram on the user terminal" refers to a device or software that displays the generated image diagram in a form that can be viewed by the user.

[1650] The "receiving means for receiving an image diagram modified by a user" refers to a device or software having a function for receiving data of an image diagram modified by a user.

[1651] The "conversion means for converting the corrected image into vector space" refers to a device or software that has the function of converting the image corrected by the user into vector data.

[1652] The "search means for searching for similar products based on the converted vectors" refers to a device or software that has the function of searching for similar products in a product database based on the vectorized image data.

[1653] "Emotion recognition means that recognizes a user's emotions and adjusts search results based on those emotions" refers to a device or software that has the function of analyzing emotions from a user's facial expressions, voice, etc., and changing or optimizing search results based on the results.

[1654] The "transmission means for transmitting search results to the user terminal" refers to a device or software that has the function of transmitting the search results to the user terminal.

[1655] System configuration

[1656] The system consists of the following main components:

[1657] 1. Server

[1658] 2. User Device

[1659] 3. Generative AI Models

[1660] 4. Product Database

[1661] 5. Emotion Recognition Engine

[1662] Program Overview

[1663] The system of the present invention allows users to input search keywords and search for products using images generated based on those keywords. It also includes a function to recognize user emotions and adjust search results accordingly.

[1664] Hardware and Software

[1665] Servers, user devices, generative AI models, etc. Specific examples include Python's PIL, Transformers (Huggingface), emotion_recognition library, and CLIP model.

[1666] Data processing and calculation

[1667] The server extracts the keywords entered by the user and generates an image using a generative AI model based on the keywords. The generated image is then sent to the user's device.

[1668] The user terminal displays the image received from the server, and the user can modify it as necessary. The modified image is then sent back to the server.

[1669] The server converts the corrected image into a vector space and searches for similar products in a product database based on the vectors. The search results are sent to the user's terminal.

[1670] The emotion recognition engine analyzes a user's facial expressions and voice to identify emotions and adjust search results accordingly.

[1671] Specific processing examples

[1672] For example, if a user enters the keyword "red sneakers," the system uses a generative AI model to generate a related image and display it to the user. The user can then modify the color of the sneakers to a darker red, and the modified image is sent to the server. The server then converts the modified image into a vector space and searches for similar products in the product database.

[1673] When search results are displayed on the user's device, an emotion recognition engine analyzes the user's facial expressions and voice, and the search results are adjusted based on their emotions. For example, if the user expresses dissatisfaction, the system can add other suggestions.

[1674] Recommended prompt examples

[1675] Keywords: "red sneakers"

[1676] Feedback: "User's Face Photo"

[1677] In this way, the present invention realizes higher user satisfaction and search accuracy by searching and proposing products taking into account the user's emotions.

[1678] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1679] Step 1:

[1680] The user enters a keyword describing a product into the search bar of the device and presses the search button. When "red sneakers" is entered as input, the device generates a request to send this keyword to the server and sends it to the server. The input is the keyword "red sneakers" and the output is the transmission of the request.

[1681] Step 2:

[1682] The server extracts keywords from the incoming request and passes them as input to the generative AI model. The model generates an image of the product from the keywords and returns the image data to the server. The input is the search keyword "red sneakers," and the output is the generated image data.

[1683] Step 3:

[1684] The server sends the generated image to the user's terminal. The input is the image data, and the output is the transmission of the image to the user's terminal.

[1685] Step 4:

[1686] The user terminal displays the image received from the server on the user interface. The user checks the displayed image and makes corrections, such as changing the color of the sneakers to dark red. The input is image data, and the output is a display containing the corrected image.

[1687] Step 5:

[1688] The user terminal generates new image data that reflects the modifications and sends it to the server. The input is the modified image, and the output is the transmission of the modified image data to the server.

[1689] Step 6:

[1690] The server converts the corrected image received into vector space. Specifically, the image is vectorized using a model such as a convolutional neural network (CNN). The input is the corrected image data, and the output is vector data.

[1691] Step 7:

[1692] The server uses the converted vector data to calculate the similarity with the image vectors in the product database. This uses techniques such as cosine similarity and Euclidean distance. The input is vector data, and the output is a list of similar products.

[1693] Step 8:

[1694] The server lists similar products and sends the information to the user terminal. The input is the list of similar products, and the output is the transmission of similar product information to the user terminal.

[1695] Step 9:

[1696] The user terminal displays the search results received from the server in a list on the user interface. The input is similar product information, and the output is the display to the user.

[1697] Step 10:

[1698] The emotion recognition engine analyzes the user's facial expressions and voice to identify emotions. Using a camera and microphone, the engine analyzes the user's facial expressions and tone of voice to recognize emotions such as "happiness," "confusion," and "anger." The input is the user's facial expression and voice data, and the output is identified emotional data.

[1699] Step 11:

[1700] The server adjusts search results based on the output of the emotion recognition engine according to the user's emotions. For example, if the user has a dissatisfied expression, the server changes the search results or adds new suggestions. The input is emotion data and search results, and the output is the adjusted search results.

[1701] Step 12:

[1702] The user inputs feedback on the search results, such as "as expected" or "disappointing," and the device sends this feedback information to the server. The input is user feedback, and the output is sending feedback information to the server.

[1703] Step 13:

[1704] The server analyzes the received feedback and updates the product database, adjusts the generative AI model, and improves the search algorithm. The input is the feedback information, and the output is improved system performance.

[1705] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1706] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1707] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1708] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1709] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1710] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1711] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1712] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1713] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1714] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1715] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1716] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1717] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1718] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1719] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1720] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1721] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1722] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1723] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1724] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1725] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1726] The following is further disclosed regarding the above embodiment.

[1727] (Claim 1)

[1728] A means for extracting input search keywords;

[1729] a generating means for generating an image diagram based on the extracted search keywords;

[1730] a display means for displaying the generated image diagram on a user terminal;

[1731] a receiving means for receiving the image diagram modified by the user;

[1732] a transforming means for transforming the corrected image into a vector space;

[1733] A search means for searching for similar products based on the transformed vector;

[1734] a transmitting means for transmitting the search results to a user terminal;

[1735] A system including:

[1736] (Claim 2)

[1737] 10. The system of claim 1, further comprising editing means for modifying the image displayed on the user terminal.

[1738] (Claim 3)

[1739] 10. The system of claim 1, further comprising an update means for receiving user feedback and updating the product database.

[1740] "Example 1"

[1741] (Claim 1)

[1742] A means for extracting input search keywords;

[1743] a generating means for generating an image diagram based on the extracted search keywords;

[1744] a display means for displaying the generated image diagram on a user terminal;

[1745] a receiving means for receiving the image diagram modified by the user;

[1746] a transforming means for transforming the corrected image into a vector space;

[1747] A search means for searching for similar products based on the transformed vector;

[1748] a transmitting means for transmitting the search results to a user terminal;

[1749] an updating means for receiving user feedback and updating the product database;

[1750] A system including:

[1751] (Claim 2)

[1752] 10. The system of claim 1, further comprising editing means for modifying the image displayed on the user terminal.

[1753] (Claim 3)

[1754] 2. The system according to claim 1, further comprising means for passing the input search keyword to the generative AI model as a prompt sentence and receiving an image diagram generated based on the prompt sentence.

[1755] "Application Example 1"

[1756] (Claim 1)

[1757] A means for extracting input search keywords;

[1758] a generating means for generating an image diagram based on the extracted search keywords;

[1759] a display means for displaying the generated image diagram on a user terminal;

[1760] a receiving means for receiving the image diagram modified by the user;

[1761] a transforming means for transforming the corrected image into a vector space;

[1762] A search means for searching for similar products based on the transformed vector;

[1763] a transmitting means for transmitting the search results to a user terminal;

[1764] a real-world image search means for receiving real-world image data captured by a camera of a user terminal and searching for products based on the real-world image data;

[1765] A system including:

[1766] (Claim 2)

[1767] 10. The system of claim 1, further comprising editing means for modifying the image displayed on the user terminal.

[1768] (Claim 3)

[1769] 10. The system of claim 1, further comprising an update means for receiving user feedback and updating the product database.

[1770] "Example 2: Combining Emotion Engines"

[1771] (Claim 1)

[1772] A means for extracting input search keywords;

[1773] a generating means for generating an image diagram based on the extracted search keywords;

[1774] a display means for displaying the generated image diagram on a user terminal;

[1775] a receiving means for receiving the image diagram modified by the user;

[1776] a transforming means for transforming the corrected image into a vector space;

[1777] A search means for searching for similar products based on the transformed vector;

[1778] a transmitting means for transmitting the search results to a user terminal;

[1779] emotion recognition means for recognizing a user's emotion and adjusting search results in response to the emotion;

[1780] A system including:

[1781] (Claim 2)

[1782] 10. The system of claim 1, further comprising editing means for modifying the image displayed on the user terminal.

[1783] (Claim 3)

[1784] 10. The system of claim 1, further comprising an update means for receiving user feedback and updating the product database.

[1785] "Application example 2 when combining emotion engines"

[1786] (Claim 1)

[1787] A means for extracting input search keywords;

[1788] a generating means for generating an image diagram based on the extracted search keywords;

[1789] a display means for displaying the generated image diagram on a user terminal;

[1790] a receiving means for receiving the image diagram modified by the user;

[1791] a transforming means for transforming the corrected image into a vector space;

[1792] A search means for searching for similar products based on the transformed vector;

[1793] emotion recognition means for recognizing an emotion of a user and adjusting search results based on the emotion;

[1794] a transmitting means for transmitting the search results to a user terminal;

[1795] A system including:

[1796] (Claim 2)

[1797] 10. The system of claim 1, further comprising editing means for modifying the image displayed on the user terminal.

[1798] (Claim 3)

[1799] 10. The system of claim 1, further comprising an update means for receiving user feedback and updating the product database. [Explanation of symbols]

[1800] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for extracting input search keywords; a generating means for generating an image diagram based on the extracted search keywords; a display means for displaying the generated image diagram on a user terminal; a receiving means for receiving the image diagram modified by the user; a transforming means for transforming the corrected image into a vector space; A search means for searching for similar products based on the transformed vector; a transmitting means for transmitting the search results to a user terminal; A system including:

2. 2. The system of claim 1, further comprising editing means for modifying the image displayed on the user terminal.

3. 10. The system of claim 1, further comprising an update means for receiving user feedback and updating the product database.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A