System
The system addresses the challenge of manual categorization in search services by using image and natural language processing to automate the categorization of search results, reducing costs and ensuring consistent quality.
Patent Information
- Application Number
- JP2024116507
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Modern search services face challenges in providing relevant products and information based on user-entered queries due to the significant manual effort required for proper categorization, which is prone to human error and increases operational costs.
A system that utilizes image recognition algorithms to extract features from images and natural language processing algorithms to extract keywords and topics from text data, automatically categorizing search results and allowing for visual review to ensure accuracy.
Reduces manual labor, lowers operational costs, and maintains consistent search result quality by efficiently categorizing search results.
Smart Images

Figure 2026015033000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Modern search services are required to provide relevant products and information based on user-entered queries, but proper categorization of search results requires significant manual effort. This places a heavy burden on operators and increases costs. Furthermore, manual categorization can be prone to human error, making it difficult to maintain consistent result quality. This invention aims to solve these problems, improve the accuracy of search results, and reduce the amount of manual work required, thereby reducing operational costs and maintaining quality. [Means for solving the problem]
[0005] The present invention provides a system for receiving a search query entered by a user, searching a database based on the search query, and retrieving relevant image and text data. The system includes means for applying an image recognition algorithm to the retrieved images to extract features, and means for applying a natural language processing algorithm to the retrieved text data to extract keywords and topics. The system further includes means for classifying search results into specific categories based on the features, keywords, and topics, and means for transmitting the classified search results to an interface for visual review and displaying the reviewed search results to the user.
[0006] A "search query" refers to a character string or image data that a user enters to search for specific information.
[0007] A "database" refers to a system for systematically storing and managing information such as images and text data.
[0008] An "image recognition algorithm" refers to a set of steps in a computer program for extracting and analyzing specific features from image data.
[0009] A "natural language processing (NLP) algorithm" refers to a set of computer program steps that analyzes text data and makes sense of it.
[0010] "Features" refer to identifiable information such as color, shape, and pattern extracted from image data.
[0011] "Keywords" refer to important words or phrases extracted from text data.
[0012] A "topic" is a concept that summarizes the content of text data and represents a major theme or topic.
[0013] "Category" refers to a group or classification for classifying search results.
[0014] "Visual review interface" refers to a display screen that allows a human to review and modify the classified search results. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] This invention relates to a system that automatically categorizes search results based on a user-entered search query. By using an algorithm to efficiently categorize search results, it is possible to significantly reduce manual labor, lower operational costs, and maintain consistent search result quality. The invention operates as follows.
[0037] When a user enters a search query from their device, the data is sent to the server. The server searches a database based on the received search query to retrieve relevant image and text data. It then applies image recognition algorithms to the retrieved image data to extract features such as color, shape, and patterns from each image. In parallel, it applies natural language processing (NLP) algorithms to the text data to extract important keywords and topics.
[0038] The server classifies search results into specific categories based on image and text features, keywords, and topics. The classified search results are sent to an interface for visual review before being displayed to the user. Visual review involves the user or administrator, who corrects incorrectly classified results. After final review is complete, the server sends the organized search results to the user's device for display.
[0039] Specific examples are shown below.
[0040] Example: A user searches for "red shirt"
[0041] 1. Enter your query:
[0042] A user enters the search query "red shirt" into a device.
[0043] The device sends a search query to the server.
[0044] 2. Obtaining and processing search results:
[0045] The server receives the search query and searches the database to retrieve image and text data related to "red shirt."
[0046] The server applies an image recognition algorithm to the images it acquires and extracts features such as "red," "short sleeves," and "design."
[0047] The server applies NLP algorithms to the text data it acquires to extract important keywords and topics such as "red shirt," "cotton," and "fashion."
[0048] 3. Automatic Category Assignment:
[0049] The server assigns categories such as "clothes," "shirt," and "red" based on the extracted features, keywords, and topics.
[0050] Group search results by category and store them in a new data structure.
[0051] 4. Final check and display:
[0052] The server then sends the compiled list of search results by category to an interface for visual review.
[0053] The user or administrator visually checks and corrects as necessary.
[0054] After final confirmation, the server sends the final search results to the user terminal and displays them on the user's screen.
[0055] This system allows users to easily browse products and information related to "red shirts." It also reduces the burden of manual categorization, reducing operational costs. Furthermore, it is expected to improve the user experience by ensuring consistent quality of search results.
[0056] The processing flow will be explained below.
[0057] Step 1:
[0058] The user enters a search query into the device. The user types in "red shirt" and presses the search button.
[0059] Step 2:
[0060] The device sends the search query entered by the user to the server. The query sent is "red shirt."
[0061] Step 3:
[0062] The server analyzes the received search query. The query analysis module extracts the keyword "red shirt."
[0063] Step 4:
[0064] The server executes a search query against the database to retrieve relevant image and text data, such as product information and images related to "red shirt."
[0065] Step 5:
[0066] The server sends the acquired image data to an image recognition module, which extracts features such as color, shape, and pattern from each image.
[0067] Step 6:
[0068] The image recognition module returns the extracted features to the server. Specifically, it recognizes that each image has features such as "red," "shirt," and "short sleeves."
[0069] Step 7:
[0070] The server sends the acquired text data to a natural language processing (NLP) module, which extracts important keywords and topics from the text data.
[0071] Step 8:
[0072] The NLP module returns the extracted keywords and topics to the server. Specifically, keywords such as "red shirt," "cotton," and "fashion" are extracted.
[0073] Step 9:
[0074] The server integrates the features, keywords, and topics returned by the image recognition and NLP modules.
[0075] Step 10:
[0076] The server uses the combined information to categorize the search results into specific categories, such as "clothes," "shirts," and "red."
[0077] Step 11:
[0078] The server sends the search results, organized by category, to a visual interface for review by the user or administrator.
[0079] Step 12:
[0080] A user or administrator visually inspects the results and corrects any incorrectly classified results.
[0081] Step 13:
[0082] After the final confirmation is completed, the server sends the final organized search results to the user terminal.
[0083] Step 14:
[0084] The device displays the final result on the user's screen, where the user can view accurate information about the "red shirt."
[0085] This series of steps allows for efficient delivery of search results to users and also reduces the burden of manual categorization.
[0086] Example 1
[0087] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0088] Existing search systems have had difficulty efficiently classifying and organizing information related to user-entered search queries. In particular, the integrated classification of image and text data relies on manual work, which is time-consuming and results in high operational costs. Furthermore, manual categorization often results in inconsistent quality of search results. The purpose of this invention is to solve the above problems and achieve efficient classification of search results while maintaining high quality.
[0089] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0090] In this invention, the server includes means for receiving a search query entered by a user, means for searching a database based on the search query to obtain related visual data and text data, and means for applying an image recognition algorithm to the obtained visual data to extract features, thereby enabling automatic acquisition of visual and text data related to the search query and efficient extraction of its features.
[0091] The server includes means for applying a natural language processing algorithm to the acquired text data to extract keywords and topics, means for classifying search results into specific categories based on the features, keywords, and topics, means for sending the classified search results to an interface for visual confirmation, and means for displaying the confirmed search results to the user. This enables automatic classification of search results and final confirmation by the user or administrator, and provides accurate, high-quality search results to the user.
[0092] A "user" is an entity that enters a search query into the system and receives the results.
[0093] A "search query" is an instruction for searching information that a user inputs to a system.
[0094] "Database" means a data storage that stores and retrieves relevant data based on a search query.
[0095] "Visual data" refers to visual information such as images and illustrations.
[0096] "Text data" refers to information expressed in text format.
[0097] An "image recognition algorithm" is a software process for extracting features such as color, shape, and pattern from image data.
[0098] A "natural language processing algorithm" is a software process for extracting keywords and topics from text data.
[0099] "Features" refer to visual attributes extracted by image recognition algorithms.
[0100] "Keywords" refer to important words or terms extracted by natural language processing algorithms.
[0101] "Topics" refer to themes or topics extracted by natural language processing algorithms.
[0102] A "category" is a group for classifying search results.
[0103] A "visual review interface" is a screen or tool that allows a user or administrator to review and modify search results.
[0104] "Means for displaying to the user" refers to the mechanism for displaying the verified search results on the user's device.
[0105] This invention relates to a system that automatically classifies and organizes related visual and textual data based on a user-entered search query. The system utilizes image recognition and natural language processing algorithms to achieve high-quality classification of search results.
[0106] Hardware and software used
[0107] The system hardware includes a device (such as a PC or smartphone) where users enter search queries, and a server that searches the database and processes the results. The server is a powerful machine that provides high-speed data processing and search functions.
[0108] The system's software includes the following components:
[0109] 1. Database search engine: An engine is needed to quickly retrieve relevant data based on a search query, such as various commercial databases.
[0110] 2. Image recognition algorithm: To extract features from the acquired visual data, we use image recognition libraries such as TensorFlow and OpenCV.
[0111] 3. Natural Language Processing (NLP) algorithms: We use libraries such as spaCy and hugface Transformers to extract keywords and topics from the acquired text data.
[0112] Specific examples
[0113] Take the example of a user entering the search query "red shirt" on a terminal.
[0114] 1. Enter your query:
[0115] A user enters the search query "red shirt" into a device.
[0116] The device sends a search query to the server.
[0117] 2. Obtaining and processing search results:
[0118] The server receives the search query and searches the database to retrieve image and text data related to "red shirt."
[0119] The server uses TensorFlow and OpenCV to perform image recognition on the images it acquires, extracting features such as color, shape, and pattern.
[0120] The server performs natural language processing on the text data it acquires using spaCy and hugface Transformers to extract important keywords and topics such as "red shirt," "cotton," and "fashion."
[0121] 3. Automatic Category Assignment:
[0122] The server classifies the search results into categories such as "clothes," "shirts," and "red color" based on the extracted features, keywords, and topics.
[0123] The server groups the search results by category and stores them in a new data structure.
[0124] 4. Final check and display:
[0125] The server then sends the compiled list of search results by category to an interface for visual review.
[0126] The user or administrator visually checks the results and makes any necessary corrections. Once the final check is complete, the server sends the organized search results to the user's device and displays them on the user's screen.
[0127] Prompt Sentence Examples
[0128] "Category search results based on the query: red shirt"
[0129] In this way, users can easily browse products and information related to "red shirts," significantly reducing the burden of manual categorization, and the quality of search results is maintained at a consistent level, improving the user experience.
[0130] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0131] Step 1:
[0132] A user enters a search query on a device.
[0133] Input: A search query, such as "red shirt," that a user types into the search bar.
[0134] What happens: The user opens a browser or application on their device and enters the desired information into the search bar.
[0135] Output: The entered search query is stored in the message to be sent by the terminal.
[0136] Step 2:
[0137] The device sends a search query to the server.
[0138] Input: The search query entered by the user in step 1.
[0139] Specific operation: The terminal generates an HTTP request and sends request data including a search query to the server.
[0140] Output: The server receives the search query.
[0141] Step 3:
[0142] The server receives the search query and searches the database.
[0143] Input: The search query received from the device.
[0144] Specific operation: The server generates a database query to search for relevant data.
[0145] Output: Visual data (images) and textual data (text) related to the search query are obtained.
[0146] Step 4:
[0147] The server retrieves the visual and textual data from the search results.
[0148] Input: Database search results containing visual and textual data that match the search query.
[0149] Specific operation: The server extracts relevant image and text data from the database.
[0150] Output: The extracted visual and textual data is stored in memory.
[0151] Step 5:
[0152] The server applies image recognition algorithms to the visual data to extract features.
[0153] Input: Acquired visual data.
[0154] Specific operation: The server uses image recognition libraries such as TensorFlow and OpenCV to analyze visual data.
[0155] Output: Features extracted from visual data (e.g., color, shape, pattern).
[0156] Step 6:
[0157] The server applies natural language processing algorithms to the text data to extract keywords and topics.
[0158] Input: Acquired sentence data.
[0159] Specific operation: The server analyzes the text data using natural language processing algorithms such as spaCy and HugFace Transformers.
[0160] Output: Keywords and topics extracted from the text data.
[0161] Step 7:
[0162] The server categorizes search results into specific categories based on features, keywords, and topics.
[0163] Input: Features extracted from visual data, and keywords and topics extracted from text data.
[0164] What it does: The server automatically categorizes the results based on a predefined list of categories.
[0165] Output: Search results organized by category.
[0166] Step 8:
[0167] The server sends the categorized search results to an interface for visual review.
[0168] Input: Search results organized by category.
[0169] Specific operation: The server performs data conversion to send the search results to the visual confirmation interface in an appropriate format.
[0170] Output: Search results displayed in a visual interface.
[0171] Step 9:
[0172] The user or administrator visually checks and corrects as necessary.
[0173] Input: Search results displayed on the visual interface.
[0174] Specific Action: A user or administrator reviews search results through the interface and corrects any incorrectly classified items.
[0175] Output: The final confirmed search results.
[0176] Step 10:
[0177] The server sends the final confirmed search results to the user's device.
[0178] Input: Last confirmed search results.
[0179] Specific operation: The server converts the search results into a data format suitable for transmission to the user terminal, and transmits the results.
[0180] Output: The final search results that are displayed on the user's device.
[0181] Step 11:
[0182] The user views the search results.
[0183] Input: The final search results displayed on the user's device.
[0184] Specific operation: The user views the search results displayed on the device and obtains the required information.
[0185] Output: Information that leads to a user's decision or next action.
[0186] (Application example 1)
[0187] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0188] Existing search systems face the challenge of effectively categorizing and displaying relevant image and text data in response to user queries. Manual categorization of search results is also labor-intensive and expensive. Smartphone shopping applications, in particular, need to efficiently manage large amounts of search results and provide a user-friendly experience.
[0189] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0190] In this invention, the server includes: means for receiving a search query entered by a user; means for searching a database based on the search query to acquire related images and text data; means for applying an image recognition algorithm to the acquired images to extract features; means for applying a natural language processing algorithm to the acquired text data to extract keywords and topics; means for categorizing search results into specific categories based on the features, keywords, and topics; means for sending the categorized search results to an interface for visual confirmation; means for displaying the confirmed search results to the user; and means for categorizing the search results and returning them to the user in JSON format. This automates the categorization of search results by category, reducing manual work and operational costs. Furthermore, users can quickly and efficiently browse related product information through a smartphone mail-order application.
[0191] "User" refers to a person who enters a search query using a device, including a smartphone.
[0192] A "search query" refers to a search term or phrase that a user enters into a search box.
[0193] "Database" refers to an information management system that stores related image and text data and provides data based on a search query.
[0194] An "image recognition algorithm" refers to a computational method for analyzing and extracting features such as color, shape, and pattern from acquired image data.
[0195] "Features" refer to specific data attributes such as color, shape, and pattern extracted using image recognition algorithms.
[0196] A "natural language processing algorithm" refers to a computational method for analyzing and extracting keywords and topics from acquired text data.
[0197] "Keywords" refer to important words and phrases extracted by natural language processing algorithms.
[0198] "Topic" refers to the main subject or theme of text data analyzed by natural language processing algorithms.
[0199] "Category" refers to various groups that categorize search results based on extracted features, keywords, and topics.
[0200] "Search Results" refers to a collection of relevant image and text data retrieved from a database based on a user's search query.
[0201] "Visual confirmation interface" refers to the display means by which a user or administrator can confirm and modify classified search results.
[0202] "JSON format" refers to a method of structuring and representing data in JavaScript Object Notation format.
[0203] "Smartphone online shopping application" refers to application software that runs on a smartphone and allows users to search and view related product information by entering a search query.
[0204] The present invention relates to a system that provides a mail-order application for smartphones that allows a user to input a search query, efficiently categorize related product information based on the query, and make the information available for browsing.
[0205] This system works as follows: When a user enters a search query using a smartphone, the data is sent to the server. Based on the entered query, the server searches a database to retrieve relevant image and text data. An image recognition algorithm is applied to the retrieved image data to extract features such as color, shape, and pattern. In parallel, a natural language processing algorithm is applied to the retrieved text data to extract important keywords and topics. Based on this, the server classifies the search results into specific categories. The classified search results are sent to a visual confirmation interface where the user or administrator can review and modify them. The confirmed search results are finally sent to the user's smartphone and displayed in JSON format.
[0206] The hardware used includes the following: a smartphone, which acts as a mobile device for entering search queries and displaying results; a server, which processes and stores the data; and finally, a database, which stores the images and text data to be searched.
[0207] The software used includes the following: Flask (Python) is used as the web application development framework; the OpenCV library is used for image recognition algorithms, and TfidfVectorizer from the Scikit-learn library is used for natural language processing; and the JSON format is used to organize and transmit data structures.
[0208] Examples:
[0209] A user enters the query "red shirt" into an online shopping application on their smartphone. The server retrieves image and text data related to "red shirt" from the database and analyzes the data. The retrieved image data is subjected to an image recognition algorithm to extract features such as "red," "shirt," and "design." The text data is analyzed using an NLP algorithm to extract keywords such as "red shirt," "cotton," and "fashion." Based on these features and keywords, the server classifies search results into categories such as "clothing," "shirt," and "red." The categorized search results are then sent to a visual review interface before being displayed to the user, where the user or administrator can make any necessary corrections. Finally, the reviewed and corrected search results are sent in JSON format to the user's smartphone and displayed.
[0210] Example prompt sentence:
[0211] When a user searches for "red shirt," related products are automatically categorized into categories such as "red," "shirt," and "clothing" based on the query and displayed to the user.
[0212] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0213] Step 1:
[0214] A user enters a search query into a shopping application on their smartphone and sends it to a server. The input data includes search terms such as "red shirt." The server receives the query and stores it for further data processing.
[0215] Step 2:
[0216] Based on the received search query, the server searches a database containing relevant image and text data. In this step, the server executes the database query to retrieve product data related to "red shirt." The output is a set of image and text data.
[0217] Step 3:
[0218] The server applies an image recognition algorithm to the acquired image data. Specifically, it uses the OpenCV library to extract features such as color, shape, and pattern from the image. The input to this process is the image data, and the output is the extracted feature data.
[0219] Step 4:
[0220] The server applies natural language processing (NLP) algorithms to the acquired text data. It uses the TfidfVectorizer from the Scikit-learn library to extract important keywords and topics from the text data. The input of this step is the text data, and the output is the extracted keywords and topics.
[0221] Step 5:
[0222] The server classifies search results into specific categories based on the extracted image features and text keywords and topics. Based on features such as "red," "shirt," and "cotton," products are automatically sorted into categories such as "clothing," "shirt," and "red." The input for this step is feature data, keywords, and topics, and the output is search results classified by category.
[0223] Step 6:
[0224] The categorized search results are sent to a visual review interface, where administrators or users can review the results and make corrections as needed. The input is the categorized search results, and the output is the reviewed and corrected search results.
[0225] Step 7:
[0226] The verified and corrected search results are sent from the server to the smartphone. The data is sent in JSON format, and the smartphone application parses it and displays it to the user. The input is the verified search results, and the output is the final search results displayed on the user's smartphone.
[0227] The above is the specific processing flow in the system of this application example. The specific data processing and calculations performed at each step enable the user to efficiently obtain product information.
[0228] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0229] This invention relates to a system that automatically categorizes search results based on a user-entered search query. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions, helping to filter and customize the search results.
[0230] When a user enters a search query from their device, the data is sent to the server. The server searches a database based on the received search query to retrieve relevant image and text data. It then applies image recognition algorithms to the retrieved image data to extract features such as color, shape, and patterns from each image. It also applies natural language processing (NLP) algorithms to the text data to extract important keywords and topics.
[0231] At the same time, the emotion engine analyzes the user's input behavior and facial expressions to recognize their emotions. This emotion information is used by the server to filter search results. For example, if the user is feeling stressed, such as "urgent," information requiring a quick response will be displayed first.
[0232] The server categorizes search results into specific categories based on image and text features, keywords, and topics, and then sends the categorized results to an interface for visual review, where users or administrators can participate and check the accuracy of the classification. After the final review is complete, the server sends the organized search results to the user's device for display.
[0233] Example: A user searches for "red shirt"
[0234] 1. Enter your query:
[0235] A user enters the search query "red shirt" into a device.
[0236] The device sends a search query to the server.
[0237] 2. Obtaining and processing search results:
[0238] The server receives the search query and searches the database to retrieve image and text data related to "red shirt."
[0239] The server applies an image recognition algorithm to the images it acquires and extracts features such as "red," "short sleeves," and "design."
[0240] The server applies NLP algorithms to the text data it acquires to extract important keywords and topics such as "red shirt," "cotton," and "fashion."
[0241] 3. Use of Emotion Engine:
[0242] When a user enters a search query, the emotion engine analyzes the user's facial expressions and typing speed to recognize their emotion.
[0243] The emotion engine provides emotion information such as whether the user is "excited" or "stressed" to the server.
[0244] 4. Automatic category assignment and filtering:
[0245] The server assigns categories such as "clothes," "shirt," and "red" based on the extracted features, keywords, and topics.
[0246] Based on the emotional information provided by the emotion engine, filtering is performed to prioritize and display search results that best suit the user's emotions.
[0247] 5. Final check and display:
[0248] The server then sends the compiled list of search results by category to an interface for visual review.
[0249] The user or administrator visually checks and corrects as necessary.
[0250] After final confirmation, the server sends the final search results to the user terminal and displays them on the user's screen.
[0251] This system allows users to easily browse products and information related to "red shirts" that are customized to their own emotions. It also reduces the burden of manual categorization, reducing operational costs. Furthermore, it is expected to improve the user experience by ensuring consistent quality of search results.
[0252] The processing flow will be explained below.
[0253] Step 1:
[0254] The user enters a search query into the device. The user types in "red shirt" and presses the search button.
[0255] Step 2:
[0256] The device sends the search query entered by the user to the server. The query sent is "red shirt."
[0257] Step 3:
[0258] The server analyzes the received search query. The query analysis module extracts the keyword "red shirt."
[0259] Step 4:
[0260] The server executes a search query against the database to retrieve relevant image and text data, such as product information and images related to "red shirt."
[0261] Step 5:
[0262] The server sends the acquired image data to an image recognition module, which extracts features such as color, shape, and pattern from each image.
[0263] Step 6:
[0264] The image recognition module returns the extracted features to the server. Specifically, it recognizes that each image has features such as "red," "shirt," and "short sleeves."
[0265] Step 7:
[0266] The server sends the acquired text data to a natural language processing (NLP) module, which extracts important keywords and topics from the text data.
[0267] Step 8:
[0268] The NLP module returns the extracted keywords and topics to the server. Specifically, keywords such as "red shirt," "cotton," and "fashion" are extracted.
[0269] Step 9:
[0270] The server integrates the features, keywords, and topics returned by the image recognition and NLP modules.
[0271] Step 10:
[0272] The server uses the combined information to categorize the search results into specific categories, such as "clothes," "shirts," and "red."
[0273] Step 11:
[0274] The server instructs the emotion engine to acquire emotion data when the user enters a search query. The emotion engine analyzes the user's facial expressions and typing speed to identify the emotion.
[0275] Step 12:
[0276] The emotion engine returns the analyzed emotion data to the server and recognizes whether the user is feeling "excited" or "stressed."
[0277] Step 13:
[0278] The server uses the emotion data to filter search results that match the recognized emotion. For example, if the user is excited, information that matches fashion trends will be displayed preferentially.
[0279] Step 14:
[0280] The server sends the categorized and filtered search results to a visual interface for review by the user or administrator.
[0281] Step 15:
[0282] A user or administrator visually inspects the results and corrects any incorrectly classified results.
[0283] Step 16:
[0284] After the final confirmation is completed, the server sends the final organized search results to the user terminal.
[0285] Step 17:
[0286] The device displays the final results on the user's screen, allowing them to view accurate and emotionally relevant information about the "red shirt."
[0287] This series of steps results in efficient search results for users, reduces the burden of manual categorization, and leverages the emotion engine to improve the quality of the user experience.
[0288] Example 2
[0289] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0290] Conventional search systems only retrieve information related to the query entered by the user, and do not adequately provide search results that reflect the user's emotions or circumstances. Furthermore, proper categorization of the retrieved search results requires manual work, resulting in high operational costs and inconsistent search accuracy. Furthermore, insufficient customization based on user emotions makes it difficult to improve the user experience.
[0291] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0292] In this invention, the server includes means for receiving a search query entered by a user, means for searching a database based on the search query to obtain related data, means for applying an image recognition algorithm to the obtained images to extract features, means for applying a natural language processing algorithm to the obtained text data to extract keywords and topics, means for classifying search results into specific categories based on the features, keywords, and topics, means for recognizing a user's emotion, means for filtering the search results based on the recognized user's emotion, means for sending the classified search results to an interface for visual confirmation, and means for displaying the confirmed search results to the user. This makes it possible to provide search results customized according to the user's emotion, reduce the burden of manual categorization, reduce operating costs, and improve the quality of search results.
[0293] A "search query" is a string of characters that a user enters to obtain information.
[0294] A "database" is a collection of information that is accessed based on a search query.
[0295] An "image recognition algorithm" is a technique for extracting features from image data.
[0296] "Features" are unique information extracted from image or text data.
[0297] "Natural language processing algorithms" are techniques for extracting keywords and topics from text data.
[0298] A "keyword" is an important word in text data.
[0299] A "topic" is the main theme or content that the text data deals with.
[0300] A "category" is a group for classifying search results.
[0301] "Emotion engine" is a technology that analyzes and recognizes the user's emotions.
[0302] "Filtering" is the process of selecting data based on specific criteria.
[0303] The "visual confirmation interface" is a display screen that allows users and administrators to check and modify search results.
[0304] "Means for displaying to the user" refers to the technology for displaying the final search results on the user's terminal.
[0305] This invention relates to a system that automatically categorizes search results based on a search query entered by a user, and then recognizes the user's emotions to provide customized search results. Specific processing for this includes the following steps:
[0306] When a user inputs a search query from a device, the device sends the data to the server. The server searches the database based on the received search query to retrieve related image and text data. For example, if the search query is "red shirt," the server extracts image and text data related to "red shirt" from the database.
[0307] The server applies image recognition algorithms to the acquired image data and extracts features such as color, shape, and pattern from each image. Libraries such as OpenCV and TensorFlow can be used for image recognition. For example, features such as "red," "short sleeves," and "design" can be extracted from an image.
[0308] The server then applies natural language processing (NLP) algorithms to the text data to extract important keywords and topics. NLP algorithms such as spaCy and NLTK are available. For example, keywords such as "red shirt," "cotton," and "fashion" are extracted from the text data.
[0309] At the same time, the emotion engine analyzes the user's input behavior and facial expressions to recognize their emotions. This emotion recognition is performed using software such as Emotion AI. As the user types a search query, the device's camera captures the user's facial expressions and identifies emotions such as "excitement" or "stress."
[0310] The server categorizes search results into specific categories based on image and text features, keywords, and topics, and also filters search results based on the user's emotional state. For example, if a user is feeling stressed, it prioritizes simple, easy-to-understand information.
[0311] The classified search results are sent to a visual review interface where the user or administrator can review the results and make any necessary corrections. The final review is then sent to the user's device and displayed on the user's screen.
[0312] Specific examples
[0313] The specific flow when a user searches for "red shirt" is as follows.
[0314] 1. Enter your query:
[0315] A user types "red shirt" into the device's search bar, and the device generates an HTTP request and sends it to the server.
[0316] 2. Obtaining and processing search results:
[0317] The server analyzes the received query and retrieves product data related to the "red shirt" from the database using an SQL query. The retrieved data includes images and text, and can also be searched using a search engine such as Apache Solr.
[0318] 3. Image Recognition and Text Analysis:
[0319] The server uses OpenCV to apply image recognition algorithms to the acquired images to extract features such as color, shape, and pattern.The acquired text data is then analyzed using spaCy and NLTK to extract important keywords.
[0320] 4. Use of Emotion Engine:
[0321] While the user is entering a query, the device's camera captures the user's facial expressions, which Emotion AI analyzes and sends emotional information to the server.
[0322] 5. Automatic category assignment and filtering:
[0323] The server assigns categories such as "clothing," "shirt," and "red" based on features obtained from image and text analysis, keywords, and emotional information provided by the emotion engine, and then prioritizes displaying search results that best match the user's emotions.
[0324] 6. Visual inspection:
[0325] The server generates search results and sends them to a visual review interface where they can be reviewed by the user or administrator, and corrections can be made if necessary.
[0326] 7. Final check and display:
[0327] The server sends the final confirmed result to the user terminal, which displays it.
[0328] This allows users to easily browse products and information related to "red shirts" that are customized to their own emotions. Furthermore, it reduces the burden of manual categorization, which is expected to reduce operational costs and improve the quality of search results.
[0329] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0330] The flow of this system's program processing
[0331] Step 1:
[0332] The user enters a search query into the search bar on the device. After completing the input, the entered search query is sent from the device to the server as an HTTP request. The request contains the query content and the user ID. Input data: The search query entered by the user in the search bar (e.g., "red shirt"). Output data: The HTTP request sent to the server.
[0333] Step 2:
[0334] The server extracts the search query from the HTTP request received and searches the database based on the query. Apache Solr or other search engines may be used, and related images and text data are obtained as search results. Input data: HTTP request (search query) sent to the server. Output data: Related images and text data.
[0335] Step 3:
[0336] The server applies an image recognition algorithm (e.g., OpenCV) to the image data it acquires to extract features such as color, shape, and pattern. This analyzes the image data and obtains each feature (e.g., "red," "short sleeves," "design"). Input data: acquired image data. Output data: extracted image feature information.
[0337] Step 4:
[0338] A natural language processing algorithm (e.g., spaCy or NLTK) is applied to the acquired text data to extract important keywords and topics. This allows related keywords (e.g., "red shirt," "cotton," "fashion") to be obtained from the text data. Input data: The acquired text data. Output data: The keywords and topics of the extracted text.
[0339] Step 5:
[0340] The emotion engine analyzes the user's input behavior and facial expression data to recognize the user's emotions. While the user is entering a query, the device's camera captures their facial expressions, which are then analyzed by emotion analysis software such as Emotion AI. Input data: User's input behavior and facial expression data. Output data: Recognized user emotional information (e.g., "excitement" or "stress").
[0341] Step 6:
[0342] The server classifies search results into specific categories based on image and text features, keywords and topics, and user sentiment information. For example, categories such as "clothing," "shirt," and "red" are assigned. Input data: Image and text feature information, and user sentiment information. Output data: Search results classified by category.
[0343] Step 7:
[0344] The server filters search results based on the user's emotional information and prioritizes displaying the information most appropriate for the user. Information is prioritized, taking into account emotions such as stress and excitement in particular. Input data: Search results categorized. Output data: Filtered search results.
[0345] Step 8:
[0346] The server sends the categorized search results to an interface for visual confirmation. Here, a web console or similar is used, allowing the user or administrator to check the search results and make corrections as necessary. Input data: Filtered search results. Output data: Search results displayed in the visual confirmation interface.
[0347] Step 9:
[0348] The user or administrator checks the search results and makes any necessary corrections. Based on this, the server generates the final confirmed search results. Input data: Search results displayed in the visual confirmation interface. Output data: Final corrected search results.
[0349] Step 10:
[0350] The server generates the final search results as an HTTP response and sends it to the user's device. The device displays the received search results on the screen, allowing the user to view customized search results. Input data: Final confirmed search results. Output data: Final search results displayed on the user's device.
[0351] This allows users to efficiently obtain and view search results that are customized to their own emotions. It also reduces the burden of manual categorization, which is expected to reduce operational costs and improve the quality of search results.
[0352] (Application example 2)
[0353] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0354] Conventional search systems can appropriately categorize related data in response to a user's search query, but they are unable to provide optimal search results that take the user's emotions into account. In particular, when a user is in a specific emotional state, such as stress or excitement, it is difficult to provide search results that are optimal for that emotion. This leads to problems such as reduced user satisfaction and an inability to obtain effective search results.
[0355] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0356] In this invention, the server includes means for receiving a search query entered by a user, means for searching a database based on the search query to obtain related images and text data, means for applying an image recognition algorithm to the obtained images to extract features, means for applying a natural language processing algorithm to the obtained text data to extract keywords and topics, means for recognizing the user's emotions, means for filtering search results based on the recognized emotions, means for classifying search results into specific categories based on the features, keywords, and topics, means for sending the classified search results to an interface for visual confirmation, and means for displaying the confirmed search results to the user, thereby enabling the provision of optimal search results that take the user's emotional state into consideration.
[0357] A "user-entered search query" is a word or phrase that a user enters into a system to request information.
[0358] A "database" is a collection of information that organizes and stores related information and provides data in response to search queries.
[0359] An "image recognition algorithm" is a computational procedure for automatically analyzing and extracting specific features or patterns from image data.
[0360] A "natural language processing algorithm" is a computational procedure for analyzing and extracting meaning, keywords, and topics from text data.
[0361] "Emotion recognition means" is a technology for determining and analyzing a user's emotional state from their facial expressions and input actions.
[0362] A "search result filtering method" is a technique for sorting and ordering search results based on specific criteria or conditions.
[0363] "Means for categorizing" refers to technology for automatically classifying search result images and text data into specific categories.
[0364] A "visual review interface" is a screen or tool that allows a user or administrator to review and correct the classification and accuracy of search results.
[0365] "Means of displaying to the user" refers to the technology used to display the organized and filtered search results on the user's device.
[0366] This invention relates to a system that automatically categorizes related information based on a search query entered by a user and provides search results that take the user's emotions into consideration. In this system, a server processes and analyzes data using various algorithms.
[0367] The server first receives a search query from the user's device. Based on the received search query, the server searches the database to retrieve relevant image and text data. Image recognition algorithms are applied to the retrieved image data to extract features such as color, shape, and pattern. Natural language processing (NLP) algorithms are applied to the retrieved text data to extract important keywords and topics.
[0368] To recognize the user's emotions, the server uses an emotion recognition means that captures the user's facial expressions with a camera and analyzes them using a machine learning model to determine the user's emotional state. It also analyzes the user's input behavior, such as typing speed on a keyboard or touchscreen.
[0369] Next, the server filters the search results based on the emotion information provided by the emotion recognition means. For example, if the user is in an "urgent" state, it prioritizes displaying information that requires a quick response, and if the user is "excited," it prioritizes displaying information about popular or new products.
[0370] The server classifies the search results into specific categories based on the extracted features and keywords. The classified search results are sent to a visual review interface for review by the user or administrator. Once review is complete, the final search results are sent to the user's terminal and displayed to the user.
[0371] The system uses the following hardware and software:
[0372] Hardware:
[0373] 1. Smartphone camera - used to capture the user's facial expressions.
[0374] 2. Server - Used to process and manage data.
[0375] software:
[0376] 1. OpenCV - Used for camera capture and face detection to recognize the user's facial expressions.
[0377] 2. TensorFlow / Keras - Used for machine learning models for emotion recognition.
[0378] 3. Transformers - Used for natural language processing algorithms.
[0379] 4. Requests - Used to communicate with the server.
[0380] As a specific example, if a user searches for "new smartphone cases" and it is recognized that the user is excited at that time, the search results will prioritize smartphone cases with high popularity rankings and the latest designs. An example of an input prompt for a generative AI model is, "Write the code for a Python program that will output search results appropriate for when the user is searching for new smartphone cases and is in an excited state."
[0381] This makes it possible to provide personalized search results that reflect the user's emotional state, improving the user experience.
[0382] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0383] Step 1:
[0384] A search query entered by a user is received at the terminal.
[0385] A user types "new phone case" into the search bar of their smartphone and presses the search button. This entered search query ("new phone case") is sent from the device to the server. The input data is a search query, and the server processes this query as received data.
[0386] Step 2:
[0387] The server searches a database based on the received search query to retrieve relevant image and text data.
[0388] Based on the search query "new smartphone case," the server searches for related information in the database. It retrieves multiple image and text data (product names, descriptions, etc.) from the database. The input data is the search query, and the output data is the related image and text data.
[0389] Step 3:
[0390] The server applies an image recognition algorithm to the image data it acquires and extracts features.
[0391] The server uses OpenCV to perform image recognition on the acquired image data, extracting features such as color, shape, and pattern from the image data. The input data is the image data, and the output data is the extracted feature data.
[0392] Step 4:
[0393] The server applies natural language processing algorithms to the text data it acquires to extract keywords and topics.
[0394] The server uses Transformers to perform natural language analysis on the acquired text data, extracting important keywords and topics from the text data. The input data is the text data, and the output data is the extracted keywords and topics.
[0395] Step 5:
[0396] The server uses an emotion recognition means to recognize the user's emotion.
[0397] When a user enters a search query, the device's camera captures their facial expression. The server uses TensorFlow / Keras to analyze the captured facial expression data and determine the user's emotional state. It also uses behavioral data such as input speed for analysis. The input data are facial expression capture data and input behavior data, and the output data is the user's emotional state.
[0398] Step 6:
[0399] The server filters the search results based on the emotion information provided by the emotion recognition means.
[0400] Based on the recognized emotion information, the server filters the extracted search results. For example, if the user is "excited," popular products and the latest designs are displayed preferentially. The input data are emotion information and search result data, and the output data are the filtered search results.
[0401] Step 7:
[0402] The server categorizes the search results into specific categories based on features, keywords and topics.
[0403] Based on the image and text data of the filtered search results, the server classifies each search result into an appropriate category (e.g., "smartphone cases," "new products," etc.). The input data is the search result data, and the output data is the search results classified by category.
[0404] Step 8:
[0405] The server sends the categorized search results to an interface for visual review.
[0406] The server sends the categorized search results to an interface where the user or administrator can check them. The input data is the categorized search results, and the output data is an interface for visual confirmation.
[0407] Step 9:
[0408] The server displays the verified search results to the user.
[0409] After the visual confirmation is completed, the server sends the final search results to the user's terminal and displays them on the screen. The input data are the search results after visual confirmation, and the output data are the search results displayed on the user's screen.
[0410] The above are the specific processing steps of the system that realizes the application example.
[0411] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0412] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0413] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0414] [Second embodiment]
[0415] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0416] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0417] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0418] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0419] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0420] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0421] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0422] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0423] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0424] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0425] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0426] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0427] This invention relates to a system that automatically categorizes search results based on a user-entered search query. By using an algorithm to efficiently categorize search results, it is possible to significantly reduce manual labor, lower operational costs, and maintain consistent search result quality. The invention operates as follows.
[0428] When a user enters a search query from their device, the data is sent to the server. The server searches a database based on the received search query to retrieve relevant image and text data. It then applies image recognition algorithms to the retrieved image data to extract features such as color, shape, and patterns from each image. In parallel, it applies natural language processing (NLP) algorithms to the text data to extract important keywords and topics.
[0429] The server classifies search results into specific categories based on image and text features, keywords, and topics. The classified search results are sent to an interface for visual review before being displayed to the user. Visual review involves the user or administrator, who corrects incorrectly classified results. After final review is complete, the server sends the organized search results to the user's device for display.
[0430] Specific examples are shown below.
[0431] Example: A user searches for "red shirt"
[0432] 1. Enter your query:
[0433] A user enters the search query "red shirt" into a device.
[0434] The device sends a search query to the server.
[0435] 2. Obtaining and processing search results:
[0436] The server receives the search query and searches the database to retrieve image and text data related to "red shirt."
[0437] The server applies an image recognition algorithm to the images it acquires and extracts features such as "red," "short sleeves," and "design."
[0438] The server applies NLP algorithms to the text data it acquires to extract important keywords and topics such as "red shirt," "cotton," and "fashion."
[0439] 3. Automatic Category Assignment:
[0440] The server assigns categories such as "clothes," "shirt," and "red" based on the extracted features, keywords, and topics.
[0441] Group search results by category and store them in a new data structure.
[0442] 4. Final check and display:
[0443] The server then sends the compiled list of search results by category to an interface for visual review.
[0444] The user or administrator visually checks and corrects as necessary.
[0445] After final confirmation, the server sends the final search results to the user terminal and displays them on the user's screen.
[0446] This system allows users to easily browse products and information related to "red shirts." It also reduces the burden of manual categorization, reducing operational costs. Furthermore, it is expected to improve the user experience by ensuring consistent quality of search results.
[0447] The processing flow will be explained below.
[0448] Step 1:
[0449] The user enters a search query into the device. The user types in "red shirt" and presses the search button.
[0450] Step 2:
[0451] The device sends the search query entered by the user to the server. The query sent is "red shirt."
[0452] Step 3:
[0453] The server analyzes the received search query. The query analysis module extracts the keyword "red shirt."
[0454] Step 4:
[0455] The server executes a search query against the database to retrieve relevant image and text data, such as product information and images related to "red shirt."
[0456] Step 5:
[0457] The server sends the acquired image data to an image recognition module, which extracts features such as color, shape, and pattern from each image.
[0458] Step 6:
[0459] The image recognition module returns the extracted features to the server. Specifically, it recognizes that each image has features such as "red," "shirt," and "short sleeves."
[0460] Step 7:
[0461] The server sends the acquired text data to a natural language processing (NLP) module, which extracts important keywords and topics from the text data.
[0462] Step 8:
[0463] The NLP module returns the extracted keywords and topics to the server. Specifically, keywords such as "red shirt," "cotton," and "fashion" are extracted.
[0464] Step 9:
[0465] The server integrates the features, keywords, and topics returned by the image recognition and NLP modules.
[0466] Step 10:
[0467] The server uses the combined information to categorize the search results into specific categories, such as "clothes," "shirts," and "red."
[0468] Step 11:
[0469] The server sends the search results, organized by category, to a visual interface for review by the user or administrator.
[0470] Step 12:
[0471] A user or administrator visually inspects the results and corrects any incorrectly classified results.
[0472] Step 13:
[0473] After the final confirmation is completed, the server sends the final organized search results to the user terminal.
[0474] Step 14:
[0475] The device displays the final result on the user's screen, where the user can view accurate information about the "red shirt."
[0476] This series of steps allows for efficient delivery of search results to users and also reduces the burden of manual categorization.
[0477] Example 1
[0478] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0479] Existing search systems have had difficulty efficiently classifying and organizing information related to user-entered search queries. In particular, the integrated classification of image and text data relies on manual work, which is time-consuming and results in high operational costs. Furthermore, manual categorization often results in inconsistent quality of search results. The purpose of this invention is to solve the above problems and achieve efficient classification of search results while maintaining high quality.
[0480] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0481] In this invention, the server includes means for receiving a search query entered by a user, means for searching a database based on the search query to obtain related visual data and text data, and means for applying an image recognition algorithm to the obtained visual data to extract features, thereby enabling automatic acquisition of visual and text data related to the search query and efficient extraction of its features.
[0482] The server includes means for applying a natural language processing algorithm to the acquired text data to extract keywords and topics, means for classifying search results into specific categories based on the features, keywords, and topics, means for sending the classified search results to an interface for visual confirmation, and means for displaying the confirmed search results to the user. This enables automatic classification of search results and final confirmation by the user or administrator, and provides accurate, high-quality search results to the user.
[0483] A "user" is an entity that enters a search query into the system and receives the results.
[0484] A "search query" is an instruction for searching information that a user inputs to a system.
[0485] "Database" means a data storage that stores and retrieves relevant data based on a search query.
[0486] "Visual data" refers to visual information such as images and illustrations.
[0487] "Text data" refers to information expressed in text format.
[0488] An "image recognition algorithm" is a software process for extracting features such as color, shape, and pattern from image data.
[0489] A "natural language processing algorithm" is a software process for extracting keywords and topics from text data.
[0490] "Features" refer to visual attributes extracted by image recognition algorithms.
[0491] "Keywords" refer to important words or terms extracted by natural language processing algorithms.
[0492] "Topics" refer to themes or topics extracted by natural language processing algorithms.
[0493] A "category" is a group for classifying search results.
[0494] A "visual review interface" is a screen or tool that allows a user or administrator to review and modify search results.
[0495] "Means for displaying to the user" refers to the mechanism for displaying the verified search results on the user's device.
[0496] This invention relates to a system that automatically classifies and organizes related visual and textual data based on a user-entered search query. The system utilizes image recognition and natural language processing algorithms to achieve high-quality classification of search results.
[0497] Hardware and software used
[0498] The system hardware includes a device (such as a PC or smartphone) where users enter search queries, and a server that searches the database and processes the results. The server is a powerful machine that provides high-speed data processing and search functions.
[0499] The system's software includes the following components:
[0500] 1. Database search engine: An engine is needed to quickly retrieve relevant data based on a search query, such as various commercial databases.
[0501] 2. Image recognition algorithm: To extract features from the acquired visual data, we use image recognition libraries such as TensorFlow and OpenCV.
[0502] 3. Natural Language Processing (NLP) algorithms: We use libraries such as spaCy and hugface Transformers to extract keywords and topics from the acquired text data.
[0503] Specific examples
[0504] Take the example of a user entering the search query "red shirt" on a terminal.
[0505] 1. Enter your query:
[0506] A user enters the search query "red shirt" into a device.
[0507] The device sends a search query to the server.
[0508] 2. Obtaining and processing search results:
[0509] The server receives the search query and searches the database to retrieve image and text data related to "red shirt."
[0510] The server uses TensorFlow and OpenCV to perform image recognition on the images it acquires, extracting features such as color, shape, and pattern.
[0511] The server performs natural language processing on the text data it acquires using spaCy and hugface Transformers to extract important keywords and topics such as "red shirt," "cotton," and "fashion."
[0512] 3. Automatic Category Assignment:
[0513] The server classifies the search results into categories such as "clothes," "shirts," and "red color" based on the extracted features, keywords, and topics.
[0514] The server groups the search results by category and stores them in a new data structure.
[0515] 4. Final check and display:
[0516] The server then sends the compiled list of search results by category to an interface for visual review.
[0517] The user or administrator visually checks the results and makes any necessary corrections. Once the final check is complete, the server sends the organized search results to the user's device and displays them on the user's screen.
[0518] Prompt Sentence Examples
[0519] "Category search results based on the query: red shirt"
[0520] In this way, users can easily browse products and information related to "red shirts," significantly reducing the burden of manual categorization, and the quality of search results is maintained at a consistent level, improving the user experience.
[0521] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0522] Step 1:
[0523] A user enters a search query on a device.
[0524] Input: A search query, such as "red shirt," that a user types into the search bar.
[0525] What happens: The user opens a browser or application on their device and enters the desired information into the search bar.
[0526] Output: The entered search query is stored in the message to be sent by the terminal.
[0527] Step 2:
[0528] The device sends a search query to the server.
[0529] Input: The search query entered by the user in step 1.
[0530] Specific operation: The terminal generates an HTTP request and sends request data including a search query to the server.
[0531] Output: The server receives the search query.
[0532] Step 3:
[0533] The server receives the search query and searches the database.
[0534] Input: The search query received from the device.
[0535] Specific operation: The server generates a database query to search for relevant data.
[0536] Output: Visual data (images) and textual data (text) related to the search query are obtained.
[0537] Step 4:
[0538] The server retrieves the visual and textual data from the search results.
[0539] Input: Database search results containing visual and textual data that match the search query.
[0540] Specific operation: The server extracts relevant image and text data from the database.
[0541] Output: The extracted visual and textual data is stored in memory.
[0542] Step 5:
[0543] The server applies image recognition algorithms to the visual data to extract features.
[0544] Input: Acquired visual data.
[0545] Specific operation: The server uses image recognition libraries such as TensorFlow and OpenCV to analyze visual data.
[0546] Output: Features extracted from visual data (e.g., color, shape, pattern).
[0547] Step 6:
[0548] The server applies natural language processing algorithms to the text data to extract keywords and topics.
[0549] Input: Acquired sentence data.
[0550] Specific operation: The server analyzes the text data using natural language processing algorithms such as spaCy and HugFace Transformers.
[0551] Output: Keywords and topics extracted from the text data.
[0552] Step 7:
[0553] The server categorizes search results into specific categories based on features, keywords, and topics.
[0554] Input: Features extracted from visual data, and keywords and topics extracted from text data.
[0555] What it does: The server automatically categorizes the results based on a predefined list of categories.
[0556] Output: Search results organized by category.
[0557] Step 8:
[0558] The server sends the categorized search results to an interface for visual review.
[0559] Input: Search results organized by category.
[0560] Specific operation: The server performs data conversion to send the search results to the visual confirmation interface in an appropriate format.
[0561] Output: Search results displayed in a visual interface.
[0562] Step 9:
[0563] The user or administrator visually checks and corrects as necessary.
[0564] Input: Search results displayed on the visual interface.
[0565] Specific Action: A user or administrator reviews search results through the interface and corrects any incorrectly classified items.
[0566] Output: The final confirmed search results.
[0567] Step 10:
[0568] The server sends the final confirmed search results to the user's device.
[0569] Input: Last confirmed search results.
[0570] Specific operation: The server converts the search results into a data format suitable for transmission to the user terminal, and transmits the results.
[0571] Output: The final search results that are displayed on the user's device.
[0572] Step 11:
[0573] The user views the search results.
[0574] Input: The final search results displayed on the user's device.
[0575] Specific operation: The user views the search results displayed on the device and obtains the required information.
[0576] Output: Information that leads to a user's decision or next action.
[0577] (Application example 1)
[0578] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0579] Existing search systems face the challenge of effectively categorizing and displaying relevant image and text data in response to user queries. Manual categorization of search results is also labor-intensive and expensive. Smartphone shopping applications, in particular, need to efficiently manage large amounts of search results and provide a user-friendly experience.
[0580] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0581] In this invention, the server includes: means for receiving a search query entered by a user; means for searching a database based on the search query to acquire related images and text data; means for applying an image recognition algorithm to the acquired images to extract features; means for applying a natural language processing algorithm to the acquired text data to extract keywords and topics; means for categorizing search results into specific categories based on the features, keywords, and topics; means for sending the categorized search results to an interface for visual confirmation; means for displaying the confirmed search results to the user; and means for categorizing the search results and returning them to the user in JSON format. This automates the categorization of search results by category, reducing manual work and operational costs. Furthermore, users can quickly and efficiently browse related product information through a smartphone mail-order application.
[0582] "User" refers to a person who enters a search query using a device, including a smartphone.
[0583] A "search query" refers to a search term or phrase that a user enters into a search box.
[0584] "Database" refers to an information management system that stores related image and text data and provides data based on a search query.
[0585] An "image recognition algorithm" refers to a computational method for analyzing and extracting features such as color, shape, and pattern from acquired image data.
[0586] "Features" refer to specific data attributes such as color, shape, and pattern extracted using image recognition algorithms.
[0587] A "natural language processing algorithm" refers to a computational method for analyzing and extracting keywords and topics from acquired text data.
[0588] "Keywords" refer to important words and phrases extracted by natural language processing algorithms.
[0589] "Topic" refers to the main subject or theme of text data analyzed by natural language processing algorithms.
[0590] "Category" refers to various groups that categorize search results based on extracted features, keywords, and topics.
[0591] "Search Results" refers to a collection of relevant image and text data retrieved from a database based on a user's search query.
[0592] "Visual confirmation interface" refers to the display means by which a user or administrator can confirm and modify classified search results.
[0593] "JSON format" refers to a method of structuring and representing data in JavaScript Object Notation format.
[0594] "Smartphone online shopping application" refers to application software that runs on a smartphone and allows users to search and view related product information by entering a search query.
[0595] The present invention relates to a system that provides a mail-order application for smartphones that allows a user to input a search query, efficiently categorize related product information based on the query, and make the information available for browsing.
[0596] This system works as follows: When a user enters a search query using a smartphone, the data is sent to the server. Based on the entered query, the server searches a database to retrieve relevant image and text data. An image recognition algorithm is applied to the retrieved image data to extract features such as color, shape, and pattern. In parallel, a natural language processing algorithm is applied to the retrieved text data to extract important keywords and topics. Based on this, the server classifies the search results into specific categories. The classified search results are sent to a visual confirmation interface where the user or administrator can review and modify them. The confirmed search results are finally sent to the user's smartphone and displayed in JSON format.
[0597] The hardware used includes the following: a smartphone, which acts as a mobile device for entering search queries and displaying results; a server, which processes and stores the data; and finally, a database, which stores the images and text data to be searched.
[0598] The software used includes the following: Flask (Python) is used as the web application development framework; the OpenCV library is used for image recognition algorithms, and TfidfVectorizer from the Scikit-learn library is used for natural language processing; and the JSON format is used to organize and transmit data structures.
[0599] Examples:
[0600] A user enters the query "red shirt" into an online shopping application on their smartphone. The server retrieves image and text data related to "red shirt" from the database and analyzes the data. The retrieved image data is subjected to an image recognition algorithm to extract features such as "red," "shirt," and "design." The text data is analyzed using an NLP algorithm to extract keywords such as "red shirt," "cotton," and "fashion." Based on these features and keywords, the server classifies search results into categories such as "clothing," "shirt," and "red." The categorized search results are then sent to a visual review interface before being displayed to the user, where the user or administrator can make any necessary corrections. Finally, the reviewed and corrected search results are sent in JSON format to the user's smartphone and displayed.
[0601] Example prompt sentence:
[0602] When a user searches for "red shirt," related products are automatically categorized into categories such as "red," "shirt," and "clothing" based on the query and displayed to the user.
[0603] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0604] Step 1:
[0605] A user enters a search query into a shopping application on their smartphone and sends it to a server. The input data includes search terms such as "red shirt." The server receives the query and stores it for further data processing.
[0606] Step 2:
[0607] Based on the received search query, the server searches a database containing relevant image and text data. In this step, the server executes the database query to retrieve product data related to "red shirt." The output is a set of image and text data.
[0608] Step 3:
[0609] The server applies an image recognition algorithm to the acquired image data. Specifically, it uses the OpenCV library to extract features such as color, shape, and pattern from the image. The input to this process is the image data, and the output is the extracted feature data.
[0610] Step 4:
[0611] The server applies natural language processing (NLP) algorithms to the acquired text data. It uses the TfidfVectorizer from the Scikit-learn library to extract important keywords and topics from the text data. The input of this step is the text data, and the output is the extracted keywords and topics.
[0612] Step 5:
[0613] The server classifies search results into specific categories based on the extracted image features and text keywords and topics. Based on features such as "red," "shirt," and "cotton," products are automatically sorted into categories such as "clothing," "shirt," and "red." The input for this step is feature data, keywords, and topics, and the output is search results classified by category.
[0614] Step 6:
[0615] The categorized search results are sent to a visual review interface, where administrators or users can review the results and make corrections as needed. The input is the categorized search results, and the output is the reviewed and corrected search results.
[0616] Step 7:
[0617] The verified and corrected search results are sent from the server to the smartphone. The data is sent in JSON format, and the smartphone application parses it and displays it to the user. The input is the verified search results, and the output is the final search results displayed on the user's smartphone.
[0618] The above is the specific processing flow in the system of this application example. The specific data processing and calculations performed at each step enable the user to efficiently obtain product information.
[0619] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0620] This invention relates to a system that automatically categorizes search results based on a user-entered search query. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions, helping to filter and customize the search results.
[0621] When a user enters a search query from their device, the data is sent to the server. The server searches a database based on the received search query to retrieve relevant image and text data. It then applies image recognition algorithms to the retrieved image data to extract features such as color, shape, and patterns from each image. It also applies natural language processing (NLP) algorithms to the text data to extract important keywords and topics.
[0622] At the same time, the emotion engine analyzes the user's input behavior and facial expressions to recognize their emotions. This emotion information is used by the server to filter search results. For example, if the user is feeling stressed, such as "urgent," information requiring a quick response will be displayed first.
[0623] The server categorizes search results into specific categories based on image and text features, keywords, and topics, and then sends the categorized results to an interface for visual review, where users or administrators can participate and check the accuracy of the classification. After the final review is complete, the server sends the organized search results to the user's device for display.
[0624] Example: A user searches for "red shirt"
[0625] 1. Enter your query:
[0626] A user enters the search query "red shirt" into a device.
[0627] The device sends a search query to the server.
[0628] 2. Obtaining and processing search results:
[0629] The server receives the search query and searches the database to retrieve image and text data related to "red shirt."
[0630] The server applies an image recognition algorithm to the images it acquires and extracts features such as "red," "short sleeves," and "design."
[0631] The server applies NLP algorithms to the text data it acquires to extract important keywords and topics such as "red shirt," "cotton," and "fashion."
[0632] 3. Use of Emotion Engine:
[0633] When a user enters a search query, the emotion engine analyzes the user's facial expressions and typing speed to recognize their emotion.
[0634] The emotion engine provides emotion information such as whether the user is "excited" or "stressed" to the server.
[0635] 4. Automatic category assignment and filtering:
[0636] The server assigns categories such as "clothes," "shirt," and "red" based on the extracted features, keywords, and topics.
[0637] Based on the emotional information provided by the emotion engine, filtering is performed to prioritize and display search results that best suit the user's emotions.
[0638] 5. Final check and display:
[0639] The server then sends the compiled list of search results by category to an interface for visual review.
[0640] The user or administrator visually checks and corrects as necessary.
[0641] After final confirmation, the server sends the final search results to the user terminal and displays them on the user's screen.
[0642] This system allows users to easily browse products and information related to "red shirts" that are customized to their own emotions. It also reduces the burden of manual categorization, reducing operational costs. Furthermore, it is expected to improve the user experience by ensuring consistent quality of search results.
[0643] The processing flow will be explained below.
[0644] Step 1:
[0645] The user enters a search query into the device. The user types in "red shirt" and presses the search button.
[0646] Step 2:
[0647] The device sends the search query entered by the user to the server. The query sent is "red shirt."
[0648] Step 3:
[0649] The server analyzes the received search query. The query analysis module extracts the keyword "red shirt."
[0650] Step 4:
[0651] The server executes a search query against the database to retrieve relevant image and text data, such as product information and images related to "red shirt."
[0652] Step 5:
[0653] The server sends the acquired image data to an image recognition module, which extracts features such as color, shape, and pattern from each image.
[0654] Step 6:
[0655] The image recognition module returns the extracted features to the server. Specifically, it recognizes that each image has features such as "red," "shirt," and "short sleeves."
[0656] Step 7:
[0657] The server sends the acquired text data to a natural language processing (NLP) module, which extracts important keywords and topics from the text data.
[0658] Step 8:
[0659] The NLP module returns the extracted keywords and topics to the server. Specifically, keywords such as "red shirt," "cotton," and "fashion" are extracted.
[0660] Step 9:
[0661] The server integrates the features, keywords, and topics returned by the image recognition and NLP modules.
[0662] Step 10:
[0663] The server uses the combined information to categorize the search results into specific categories, such as "clothes," "shirts," and "red."
[0664] Step 11:
[0665] The server instructs the emotion engine to acquire emotion data when the user enters a search query. The emotion engine analyzes the user's facial expressions and typing speed to identify the emotion.
[0666] Step 12:
[0667] The emotion engine returns the analyzed emotion data to the server and recognizes whether the user is feeling "excited" or "stressed."
[0668] Step 13:
[0669] The server uses the emotion data to filter search results that match the recognized emotion. For example, if the user is excited, information that matches fashion trends will be displayed preferentially.
[0670] Step 14:
[0671] The server sends the categorized and filtered search results to a visual interface for review by the user or administrator.
[0672] Step 15:
[0673] A user or administrator visually inspects the results and corrects any incorrectly classified results.
[0674] Step 16:
[0675] After the final confirmation is completed, the server sends the final organized search results to the user terminal.
[0676] Step 17:
[0677] The device displays the final results on the user's screen, allowing them to view accurate and emotionally relevant information about the "red shirt."
[0678] This series of steps results in efficient search results for users, reduces the burden of manual categorization, and leverages the emotion engine to improve the quality of the user experience.
[0679] Example 2
[0680] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0681] Conventional search systems only retrieve information related to the query entered by the user, and do not adequately provide search results that reflect the user's emotions or circumstances. Furthermore, proper categorization of the retrieved search results requires manual work, resulting in high operational costs and inconsistent search accuracy. Furthermore, insufficient customization based on user emotions makes it difficult to improve the user experience.
[0682] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0683] In this invention, the server includes means for receiving a search query entered by a user, means for searching a database based on the search query to obtain related data, means for applying an image recognition algorithm to the obtained images to extract features, means for applying a natural language processing algorithm to the obtained text data to extract keywords and topics, means for classifying search results into specific categories based on the features, keywords, and topics, means for recognizing a user's emotion, means for filtering the search results based on the recognized user's emotion, means for sending the classified search results to an interface for visual confirmation, and means for displaying the confirmed search results to the user. This makes it possible to provide search results customized according to the user's emotion, reduce the burden of manual categorization, reduce operating costs, and improve the quality of search results.
[0684] A "search query" is a string of characters that a user enters to obtain information.
[0685] A "database" is a collection of information that is accessed based on a search query.
[0686] An "image recognition algorithm" is a technique for extracting features from image data.
[0687] "Features" are unique information extracted from image or text data.
[0688] "Natural language processing algorithms" are techniques for extracting keywords and topics from text data.
[0689] A "keyword" is an important word in text data.
[0690] A "topic" is the main theme or content that the text data deals with.
[0691] A "category" is a group for classifying search results.
[0692] "Emotion engine" is a technology that analyzes and recognizes the user's emotions.
[0693] "Filtering" is the process of selecting data based on specific criteria.
[0694] The "visual confirmation interface" is a display screen that allows users and administrators to check and modify search results.
[0695] "Means for displaying to the user" refers to the technology for displaying the final search results on the user's terminal.
[0696] This invention relates to a system that automatically categorizes search results based on a search query entered by a user, and then recognizes the user's emotions to provide customized search results. Specific processing for this includes the following steps:
[0697] When a user inputs a search query from a device, the device sends the data to the server. The server searches the database based on the received search query to retrieve related image and text data. For example, if the search query is "red shirt," the server extracts image and text data related to "red shirt" from the database.
[0698] The server applies image recognition algorithms to the acquired image data and extracts features such as color, shape, and pattern from each image. Libraries such as OpenCV and TensorFlow can be used for image recognition. For example, features such as "red," "short sleeves," and "design" can be extracted from an image.
[0699] The server then applies natural language processing (NLP) algorithms to the text data to extract important keywords and topics. NLP algorithms such as spaCy and NLTK are available. For example, keywords such as "red shirt," "cotton," and "fashion" are extracted from the text data.
[0700] At the same time, the emotion engine analyzes the user's input behavior and facial expressions to recognize their emotions. This emotion recognition is performed using software such as Emotion AI. As the user types a search query, the device's camera captures the user's facial expressions and identifies emotions such as "excitement" or "stress."
[0701] The server categorizes search results into specific categories based on image and text features, keywords, and topics, and also filters search results based on the user's emotional state. For example, if a user is feeling stressed, it prioritizes simple, easy-to-understand information.
[0702] The classified search results are sent to a visual review interface where the user or administrator can review the results and make any necessary corrections. The final review is then sent to the user's device and displayed on the user's screen.
[0703] Specific examples
[0704] The specific flow when a user searches for "red shirt" is as follows.
[0705] 1. Enter your query:
[0706] A user types "red shirt" into the device's search bar, and the device generates an HTTP request and sends it to the server.
[0707] 2. Obtaining and processing search results:
[0708] The server analyzes the received query and retrieves product data related to the "red shirt" from the database using an SQL query. The retrieved data includes images and text, and can also be searched using a search engine such as Apache Solr.
[0709] 3. Image Recognition and Text Analysis:
[0710] The server uses OpenCV to apply image recognition algorithms to the acquired images to extract features such as color, shape, and pattern.The acquired text data is then analyzed using spaCy and NLTK to extract important keywords.
[0711] 4. Use of Emotion Engine:
[0712] While the user is entering a query, the device's camera captures the user's facial expressions, which Emotion AI analyzes and sends emotional information to the server.
[0713] 5. Automatic category assignment and filtering:
[0714] The server assigns categories such as "clothing," "shirt," and "red" based on features obtained from image and text analysis, keywords, and emotional information provided by the emotion engine, and then prioritizes displaying search results that best match the user's emotions.
[0715] 6. Visual inspection:
[0716] The server generates search results and sends them to a visual review interface where they can be reviewed by the user or administrator, and corrections can be made if necessary.
[0717] 7. Final check and display:
[0718] The server sends the final confirmed result to the user terminal, which displays it.
[0719] This allows users to easily browse products and information related to "red shirts" that are customized to their own emotions. Furthermore, it reduces the burden of manual categorization, which is expected to reduce operational costs and improve the quality of search results.
[0720] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0721] The flow of this system's program processing
[0722] Step 1:
[0723] The user enters a search query into the search bar on the device. After completing the input, the entered search query is sent from the device to the server as an HTTP request. The request contains the query content and the user ID. Input data: The search query entered by the user in the search bar (e.g., "red shirt"). Output data: The HTTP request sent to the server.
[0724] Step 2:
[0725] The server extracts the search query from the HTTP request received and searches the database based on the query. Apache Solr or other search engines may be used, and related images and text data are obtained as search results. Input data: HTTP request (search query) sent to the server. Output data: Related images and text data.
[0726] Step 3:
[0727] The server applies an image recognition algorithm (e.g., OpenCV) to the image data it acquires to extract features such as color, shape, and pattern. This analyzes the image data and obtains each feature (e.g., "red," "short sleeves," "design"). Input data: acquired image data. Output data: extracted image feature information.
[0728] Step 4:
[0729] A natural language processing algorithm (e.g., spaCy or NLTK) is applied to the acquired text data to extract important keywords and topics. This allows related keywords (e.g., "red shirt," "cotton," "fashion") to be obtained from the text data. Input data: The acquired text data. Output data: The keywords and topics of the extracted text.
[0730] Step 5:
[0731] The emotion engine analyzes the user's input behavior and facial expression data to recognize the user's emotions. While the user is entering a query, the device's camera captures their facial expressions, which are then analyzed by emotion analysis software such as Emotion AI. Input data: User's input behavior and facial expression data. Output data: Recognized user emotional information (e.g., "excitement" or "stress").
[0732] Step 6:
[0733] The server classifies search results into specific categories based on image and text features, keywords and topics, and user sentiment information. For example, categories such as "clothing," "shirt," and "red" are assigned. Input data: Image and text feature information, and user sentiment information. Output data: Search results classified by category.
[0734] Step 7:
[0735] The server filters search results based on the user's emotional information and prioritizes displaying the information most appropriate for the user. Information is prioritized, taking into account emotions such as stress and excitement in particular. Input data: Search results categorized. Output data: Filtered search results.
[0736] Step 8:
[0737] The server sends the categorized search results to an interface for visual confirmation. Here, a web console or similar is used, allowing the user or administrator to check the search results and make corrections as necessary. Input data: Filtered search results. Output data: Search results displayed in the visual confirmation interface.
[0738] Step 9:
[0739] The user or administrator checks the search results and makes any necessary corrections. Based on this, the server generates the final confirmed search results. Input data: Search results displayed in the visual confirmation interface. Output data: Final corrected search results.
[0740] Step 10:
[0741] The server generates the final search results as an HTTP response and sends it to the user's device. The device displays the received search results on the screen, allowing the user to view customized search results. Input data: Final confirmed search results. Output data: Final search results displayed on the user's device.
[0742] This allows users to efficiently obtain and view search results that are customized to their own emotions. It also reduces the burden of manual categorization, which is expected to reduce operational costs and improve the quality of search results.
[0743] (Application example 2)
[0744] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0745] Conventional search systems can appropriately categorize related data in response to a user's search query, but they are unable to provide optimal search results that take the user's emotions into account. In particular, when a user is in a specific emotional state, such as stress or excitement, it is difficult to provide search results that are optimal for that emotion. This leads to problems such as reduced user satisfaction and an inability to obtain effective search results.
[0746] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0747] In this invention, the server includes means for receiving a search query entered by a user, means for searching a database based on the search query to obtain related images and text data, means for applying an image recognition algorithm to the obtained images to extract features, means for applying a natural language processing algorithm to the obtained text data to extract keywords and topics, means for recognizing the user's emotions, means for filtering search results based on the recognized emotions, means for classifying search results into specific categories based on the features, keywords, and topics, means for sending the classified search results to an interface for visual confirmation, and means for displaying the confirmed search results to the user, thereby enabling the provision of optimal search results that take the user's emotional state into consideration.
[0748] A "user-entered search query" is a word or phrase that a user enters into a system to request information.
[0749] A "database" is a collection of information that organizes and stores related information and provides data in response to search queries.
[0750] An "image recognition algorithm" is a computational procedure for automatically analyzing and extracting specific features or patterns from image data.
[0751] A "natural language processing algorithm" is a computational procedure for analyzing and extracting meaning, keywords, and topics from text data.
[0752] "Emotion recognition means" is a technology for determining and analyzing a user's emotional state from their facial expressions and input actions.
[0753] A "search result filtering method" is a technique for sorting and ordering search results based on specific criteria or conditions.
[0754] "Means for categorizing" refers to technology for automatically classifying search result images and text data into specific categories.
[0755] A "visual review interface" is a screen or tool that allows a user or administrator to review and correct the classification and accuracy of search results.
[0756] "Means of displaying to the user" refers to the technology used to display the organized and filtered search results on the user's device.
[0757] This invention relates to a system that automatically categorizes related information based on a search query entered by a user and provides search results that take the user's emotions into consideration. In this system, a server processes and analyzes data using various algorithms.
[0758] The server first receives a search query from the user's device. Based on the received search query, the server searches the database to retrieve relevant image and text data. Image recognition algorithms are applied to the retrieved image data to extract features such as color, shape, and pattern. Natural language processing (NLP) algorithms are applied to the retrieved text data to extract important keywords and topics.
[0759] To recognize the user's emotions, the server uses an emotion recognition means that captures the user's facial expressions with a camera and analyzes them using a machine learning model to determine the user's emotional state. It also analyzes the user's input behavior, such as typing speed on a keyboard or touchscreen.
[0760] Next, the server filters the search results based on the emotion information provided by the emotion recognition means. For example, if the user is in an "urgent" state, it prioritizes displaying information that requires a quick response, and if the user is "excited," it prioritizes displaying information about popular or new products.
[0761] The server classifies the search results into specific categories based on the extracted features and keywords. The classified search results are sent to a visual review interface for review by the user or administrator. Once review is complete, the final search results are sent to the user's terminal and displayed to the user.
[0762] The system uses the following hardware and software:
[0763] Hardware:
[0764] 1. Smartphone camera - used to capture the user's facial expressions.
[0765] 2. Server - Used to process and manage data.
[0766] software:
[0767] 1. OpenCV - Used for camera capture and face detection to recognize the user's facial expressions.
[0768] 2. TensorFlow / Keras - Used for machine learning models for emotion recognition.
[0769] 3. Transformers - Used for natural language processing algorithms.
[0770] 4. Requests - Used to communicate with the server.
[0771] As a specific example, if a user searches for "new smartphone cases" and it is recognized that the user is excited at that time, the search results will prioritize smartphone cases with high popularity rankings and the latest designs. An example of an input prompt for a generative AI model is, "Write the code for a Python program that will output search results appropriate for when the user is searching for new smartphone cases and is in an excited state."
[0772] This makes it possible to provide personalized search results that reflect the user's emotional state, improving the user experience.
[0773] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0774] Step 1:
[0775] A search query entered by a user is received at the terminal.
[0776] A user types "new phone case" into the search bar of their smartphone and presses the search button. This entered search query ("new phone case") is sent from the device to the server. The input data is a search query, and the server processes this query as received data.
[0777] Step 2:
[0778] The server searches a database based on the received search query to retrieve relevant image and text data.
[0779] Based on the search query "new smartphone case," the server searches for related information in the database. It retrieves multiple image and text data (product names, descriptions, etc.) from the database. The input data is the search query, and the output data is the related image and text data.
[0780] Step 3:
[0781] The server applies an image recognition algorithm to the image data it acquires and extracts features.
[0782] The server uses OpenCV to perform image recognition on the acquired image data, extracting features such as color, shape, and pattern from the image data. The input data is the image data, and the output data is the extracted feature data.
[0783] Step 4:
[0784] The server applies natural language processing algorithms to the text data it acquires to extract keywords and topics.
[0785] The server uses Transformers to perform natural language analysis on the acquired text data, extracting important keywords and topics from the text data. The input data is the text data, and the output data is the extracted keywords and topics.
[0786] Step 5:
[0787] The server uses an emotion recognition means to recognize the user's emotion.
[0788] When a user enters a search query, the device's camera captures their facial expression. The server uses TensorFlow / Keras to analyze the captured facial expression data and determine the user's emotional state. It also uses behavioral data such as input speed for analysis. The input data are facial expression capture data and input behavior data, and the output data is the user's emotional state.
[0789] Step 6:
[0790] The server filters the search results based on the emotion information provided by the emotion recognition means.
[0791] Based on the recognized emotion information, the server filters the extracted search results. For example, if the user is "excited," popular products and the latest designs are displayed preferentially. The input data are emotion information and search result data, and the output data are the filtered search results.
[0792] Step 7:
[0793] The server categorizes the search results into specific categories based on features, keywords and topics.
[0794] Based on the image and text data of the filtered search results, the server classifies each search result into an appropriate category (e.g., "smartphone cases," "new products," etc.). The input data is the search result data, and the output data is the search results classified by category.
[0795] Step 8:
[0796] The server sends the categorized search results to an interface for visual review.
[0797] The server sends the categorized search results to an interface where the user or administrator can check them. The input data is the categorized search results, and the output data is an interface for visual confirmation.
[0798] Step 9:
[0799] The server displays the verified search results to the user.
[0800] After the visual confirmation is completed, the server sends the final search results to the user's terminal and displays them on the screen. The input data are the search results after visual confirmation, and the output data are the search results displayed on the user's screen.
[0801] The above are the specific processing steps of the system that realizes the application example.
[0802] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0803] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0804] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0805] [Third embodiment]
[0806] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0807] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0808] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0809] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0810] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0811] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0812] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0813] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0814] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0815] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0816] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0817] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0818] This invention relates to a system that automatically categorizes search results based on a user-entered search query. By using an algorithm to efficiently categorize search results, it is possible to significantly reduce manual labor, lower operational costs, and maintain consistent search result quality. The invention operates as follows.
[0819] When a user enters a search query from their device, the data is sent to the server. The server searches a database based on the received search query to retrieve relevant image and text data. It then applies image recognition algorithms to the retrieved image data to extract features such as color, shape, and patterns from each image. In parallel, it applies natural language processing (NLP) algorithms to the text data to extract important keywords and topics.
[0820] The server classifies search results into specific categories based on image and text features, keywords, and topics. The classified search results are sent to an interface for visual review before being displayed to the user. Visual review involves the user or administrator, who corrects incorrectly classified results. After final review is complete, the server sends the organized search results to the user's device for display.
[0821] Specific examples are shown below.
[0822] Example: A user searches for "red shirt"
[0823] 1. Enter your query:
[0824] A user enters the search query "red shirt" into a device.
[0825] The device sends a search query to the server.
[0826] 2. Obtaining and processing search results:
[0827] The server receives the search query and searches the database to retrieve image and text data related to "red shirt."
[0828] The server applies an image recognition algorithm to the images it acquires and extracts features such as "red," "short sleeves," and "design."
[0829] The server applies NLP algorithms to the text data it acquires to extract important keywords and topics such as "red shirt," "cotton," and "fashion."
[0830] 3. Automatic Category Assignment:
[0831] The server assigns categories such as "clothes," "shirt," and "red" based on the extracted features, keywords, and topics.
[0832] Group search results by category and store them in a new data structure.
[0833] 4. Final check and display:
[0834] The server then sends the compiled list of search results by category to an interface for visual review.
[0835] The user or administrator visually checks and corrects as necessary.
[0836] After final confirmation, the server sends the final search results to the user terminal and displays them on the user's screen.
[0837] This system allows users to easily browse products and information related to "red shirts." It also reduces the burden of manual categorization, reducing operational costs. Furthermore, it is expected to improve the user experience by ensuring consistent quality of search results.
[0838] The processing flow will be explained below.
[0839] Step 1:
[0840] The user enters a search query into the device. The user types in "red shirt" and presses the search button.
[0841] Step 2:
[0842] The device sends the search query entered by the user to the server. The query sent is "red shirt."
[0843] Step 3:
[0844] The server analyzes the received search query. The query analysis module extracts the keyword "red shirt."
[0845] Step 4:
[0846] The server executes a search query against the database to retrieve relevant image and text data, such as product information and images related to "red shirt."
[0847] Step 5:
[0848] The server sends the acquired image data to an image recognition module, which extracts features such as color, shape, and pattern from each image.
[0849] Step 6:
[0850] The image recognition module returns the extracted features to the server. Specifically, it recognizes that each image has features such as "red," "shirt," and "short sleeves."
[0851] Step 7:
[0852] The server sends the acquired text data to a natural language processing (NLP) module, which extracts important keywords and topics from the text data.
[0853] Step 8:
[0854] The NLP module returns the extracted keywords and topics to the server. Specifically, keywords such as "red shirt," "cotton," and "fashion" are extracted.
[0855] Step 9:
[0856] The server integrates the features, keywords, and topics returned by the image recognition and NLP modules.
[0857] Step 10:
[0858] The server uses the combined information to categorize the search results into specific categories, such as "clothes," "shirts," and "red."
[0859] Step 11:
[0860] The server sends the search results, organized by category, to a visual interface for review by the user or administrator.
[0861] Step 12:
[0862] A user or administrator visually inspects the results and corrects any incorrectly classified results.
[0863] Step 13:
[0864] After the final confirmation is completed, the server sends the final organized search results to the user terminal.
[0865] Step 14:
[0866] The device displays the final result on the user's screen, where the user can view accurate information about the "red shirt."
[0867] This series of steps allows for efficient delivery of search results to users and also reduces the burden of manual categorization.
[0868] Example 1
[0869] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0870] Existing search systems have had difficulty efficiently classifying and organizing information related to user-entered search queries. In particular, the integrated classification of image and text data relies on manual work, which is time-consuming and results in high operational costs. Furthermore, manual categorization often results in inconsistent quality of search results. The purpose of this invention is to solve the above problems and achieve efficient classification of search results while maintaining high quality.
[0871] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0872] In this invention, the server includes means for receiving a search query entered by a user, means for searching a database based on the search query to obtain related visual data and text data, and means for applying an image recognition algorithm to the obtained visual data to extract features, thereby enabling automatic acquisition of visual and text data related to the search query and efficient extraction of its features.
[0873] The server includes means for applying a natural language processing algorithm to the acquired text data to extract keywords and topics, means for classifying search results into specific categories based on the features, keywords, and topics, means for sending the classified search results to an interface for visual confirmation, and means for displaying the confirmed search results to the user. This enables automatic classification of search results and final confirmation by the user or administrator, and provides accurate, high-quality search results to the user.
[0874] A "user" is an entity that enters a search query into the system and receives the results.
[0875] A "search query" is an instruction for searching information that a user inputs to a system.
[0876] "Database" means a data storage that stores and retrieves relevant data based on a search query.
[0877] "Visual data" refers to visual information such as images and illustrations.
[0878] "Text data" refers to information expressed in text format.
[0879] An "image recognition algorithm" is a software process for extracting features such as color, shape, and pattern from image data.
[0880] A "natural language processing algorithm" is a software process for extracting keywords and topics from text data.
[0881] "Features" refer to visual attributes extracted by image recognition algorithms.
[0882] "Keywords" refer to important words or terms extracted by natural language processing algorithms.
[0883] "Topics" refer to themes or topics extracted by natural language processing algorithms.
[0884] A "category" is a group for classifying search results.
[0885] A "visual review interface" is a screen or tool that allows a user or administrator to review and modify search results.
[0886] "Means for displaying to the user" refers to the mechanism for displaying the verified search results on the user's device.
[0887] This invention relates to a system that automatically classifies and organizes related visual and textual data based on a user-entered search query. The system utilizes image recognition and natural language processing algorithms to achieve high-quality classification of search results.
[0888] Hardware and software used
[0889] The system hardware includes a device (such as a PC or smartphone) where users enter search queries, and a server that searches the database and processes the results. The server is a powerful machine that provides high-speed data processing and search functions.
[0890] The system's software includes the following components:
[0891] 1. Database search engine: An engine is needed to quickly retrieve relevant data based on a search query, such as various commercial databases.
[0892] 2. Image recognition algorithm: To extract features from the acquired visual data, we use image recognition libraries such as TensorFlow and OpenCV.
[0893] 3. Natural Language Processing (NLP) algorithms: We use libraries such as spaCy and hugface Transformers to extract keywords and topics from the acquired text data.
[0894] Specific examples
[0895] Take the example of a user entering the search query "red shirt" on a terminal.
[0896] 1. Enter your query:
[0897] A user enters the search query "red shirt" into a device.
[0898] The device sends a search query to the server.
[0899] 2. Obtaining and processing search results:
[0900] The server receives the search query and searches the database to retrieve image and text data related to "red shirt."
[0901] The server uses TensorFlow and OpenCV to perform image recognition on the images it acquires, extracting features such as color, shape, and pattern.
[0902] The server performs natural language processing on the text data it acquires using spaCy and hugface Transformers to extract important keywords and topics such as "red shirt," "cotton," and "fashion."
[0903] 3. Automatic Category Assignment:
[0904] The server classifies the search results into categories such as "clothes," "shirts," and "red color" based on the extracted features, keywords, and topics.
[0905] The server groups the search results by category and stores them in a new data structure.
[0906] 4. Final check and display:
[0907] The server then sends the compiled list of search results by category to an interface for visual review.
[0908] The user or administrator visually checks the results and makes any necessary corrections. Once the final check is complete, the server sends the organized search results to the user's device and displays them on the user's screen.
[0909] Prompt Sentence Examples
[0910] "Category search results based on the query: red shirt"
[0911] In this way, users can easily browse products and information related to "red shirts," significantly reducing the burden of manual categorization, and the quality of search results is maintained at a consistent level, improving the user experience.
[0912] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0913] Step 1:
[0914] A user enters a search query on a device.
[0915] Input: A search query, such as "red shirt," that a user types into the search bar.
[0916] What happens: The user opens a browser or application on their device and enters the desired information into the search bar.
[0917] Output: The entered search query is stored in the message to be sent by the terminal.
[0918] Step 2:
[0919] The device sends a search query to the server.
[0920] Input: The search query entered by the user in step 1.
[0921] Specific operation: The terminal generates an HTTP request and sends request data including a search query to the server.
[0922] Output: The server receives the search query.
[0923] Step 3:
[0924] The server receives the search query and searches the database.
[0925] Input: The search query received from the device.
[0926] Specific operation: The server generates a database query to search for relevant data.
[0927] Output: Visual data (images) and textual data (text) related to the search query are obtained.
[0928] Step 4:
[0929] The server retrieves the visual and textual data from the search results.
[0930] Input: Database search results containing visual and textual data that match the search query.
[0931] Specific operation: The server extracts relevant image and text data from the database.
[0932] Output: The extracted visual and textual data is stored in memory.
[0933] Step 5:
[0934] The server applies image recognition algorithms to the visual data to extract features.
[0935] Input: Acquired visual data.
[0936] Specific operation: The server uses image recognition libraries such as TensorFlow and OpenCV to analyze visual data.
[0937] Output: Features extracted from visual data (e.g., color, shape, pattern).
[0938] Step 6:
[0939] The server applies natural language processing algorithms to the text data to extract keywords and topics.
[0940] Input: Acquired sentence data.
[0941] Specific operation: The server analyzes the text data using natural language processing algorithms such as spaCy and HugFace Transformers.
[0942] Output: Keywords and topics extracted from the text data.
[0943] Step 7:
[0944] The server categorizes search results into specific categories based on features, keywords, and topics.
[0945] Input: Features extracted from visual data, and keywords and topics extracted from text data.
[0946] What it does: The server automatically categorizes the results based on a predefined list of categories.
[0947] Output: Search results organized by category.
[0948] Step 8:
[0949] The server sends the categorized search results to an interface for visual review.
[0950] Input: Search results organized by category.
[0951] Specific operation: The server performs data conversion to send the search results to the visual confirmation interface in an appropriate format.
[0952] Output: Search results displayed in a visual interface.
[0953] Step 9:
[0954] The user or administrator visually checks and corrects as necessary.
[0955] Input: Search results displayed on the visual interface.
[0956] Specific Action: A user or administrator reviews search results through the interface and corrects any incorrectly classified items.
[0957] Output: The final confirmed search results.
[0958] Step 10:
[0959] The server sends the final confirmed search results to the user's device.
[0960] Input: Last confirmed search results.
[0961] Specific operation: The server converts the search results into a data format suitable for transmission to the user terminal, and transmits the results.
[0962] Output: The final search results that are displayed on the user's device.
[0963] Step 11:
[0964] The user views the search results.
[0965] Input: The final search results displayed on the user's device.
[0966] Specific operation: The user views the search results displayed on the device and obtains the required information.
[0967] Output: Information that leads to a user's decision or next action.
[0968] (Application example 1)
[0969] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0970] Existing search systems face the challenge of effectively categorizing and displaying relevant image and text data in response to user queries. Manual categorization of search results is also labor-intensive and expensive. Smartphone shopping applications, in particular, need to efficiently manage large amounts of search results and provide a user-friendly experience.
[0971] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0972] In this invention, the server includes: means for receiving a search query entered by a user; means for searching a database based on the search query to acquire related images and text data; means for applying an image recognition algorithm to the acquired images to extract features; means for applying a natural language processing algorithm to the acquired text data to extract keywords and topics; means for categorizing search results into specific categories based on the features, keywords, and topics; means for sending the categorized search results to an interface for visual confirmation; means for displaying the confirmed search results to the user; and means for categorizing the search results and returning them to the user in JSON format. This automates the categorization of search results by category, reducing manual work and operational costs. Furthermore, users can quickly and efficiently browse related product information through a smartphone mail-order application.
[0973] "User" refers to a person who enters a search query using a device, including a smartphone.
[0974] A "search query" refers to a search term or phrase that a user enters into a search box.
[0975] "Database" refers to an information management system that stores related image and text data and provides data based on a search query.
[0976] An "image recognition algorithm" refers to a computational method for analyzing and extracting features such as color, shape, and pattern from acquired image data.
[0977] "Features" refer to specific data attributes such as color, shape, and pattern extracted using image recognition algorithms.
[0978] A "natural language processing algorithm" refers to a computational method for analyzing and extracting keywords and topics from acquired text data.
[0979] "Keywords" refer to important words and phrases extracted by natural language processing algorithms.
[0980] "Topic" refers to the main subject or theme of text data analyzed by natural language processing algorithms.
[0981] "Category" refers to various groups that categorize search results based on extracted features, keywords, and topics.
[0982] "Search Results" refers to a collection of relevant image and text data retrieved from a database based on a user's search query.
[0983] "Visual confirmation interface" refers to the display means by which a user or administrator can confirm and modify classified search results.
[0984] "JSON format" refers to a method of structuring and representing data in JavaScript Object Notation format.
[0985] "Smartphone online shopping application" refers to application software that runs on a smartphone and allows users to search and view related product information by entering a search query.
[0986] The present invention relates to a system that provides a mail-order application for smartphones that allows a user to input a search query, efficiently categorize related product information based on the query, and make the information available for browsing.
[0987] This system works as follows: When a user enters a search query using a smartphone, the data is sent to the server. Based on the entered query, the server searches a database to retrieve relevant image and text data. An image recognition algorithm is applied to the retrieved image data to extract features such as color, shape, and pattern. In parallel, a natural language processing algorithm is applied to the retrieved text data to extract important keywords and topics. Based on this, the server classifies the search results into specific categories. The classified search results are sent to a visual confirmation interface where the user or administrator can review and modify them. The confirmed search results are finally sent to the user's smartphone and displayed in JSON format.
[0988] The hardware used includes the following: a smartphone, which acts as a mobile device for entering search queries and displaying results; a server, which processes and stores the data; and finally, a database, which stores the images and text data to be searched.
[0989] The software used includes the following: Flask (Python) is used as the web application development framework; the OpenCV library is used for image recognition algorithms, and TfidfVectorizer from the Scikit-learn library is used for natural language processing; and the JSON format is used to organize and transmit data structures.
[0990] Examples:
[0991] A user enters the query "red shirt" into an online shopping application on their smartphone. The server retrieves image and text data related to "red shirt" from the database and analyzes the data. The retrieved image data is subjected to an image recognition algorithm to extract features such as "red," "shirt," and "design." The text data is analyzed using an NLP algorithm to extract keywords such as "red shirt," "cotton," and "fashion." Based on these features and keywords, the server classifies search results into categories such as "clothing," "shirt," and "red." The categorized search results are then sent to a visual review interface before being displayed to the user, where the user or administrator can make any necessary corrections. Finally, the reviewed and corrected search results are sent in JSON format to the user's smartphone and displayed.
[0992] Example prompt sentence:
[0993] When a user searches for "red shirt," related products are automatically categorized into categories such as "red," "shirt," and "clothing" based on the query and displayed to the user.
[0994] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0995] Step 1:
[0996] A user enters a search query into a shopping application on their smartphone and sends it to a server. The input data includes search terms such as "red shirt." The server receives the query and stores it for further data processing.
[0997] Step 2:
[0998] Based on the received search query, the server searches a database containing relevant image and text data. In this step, the server executes the database query to retrieve product data related to "red shirt." The output is a set of image and text data.
[0999] Step 3:
[1000] The server applies an image recognition algorithm to the acquired image data. Specifically, it uses the OpenCV library to extract features such as color, shape, and pattern from the image. The input to this process is the image data, and the output is the extracted feature data.
[1001] Step 4:
[1002] The server applies natural language processing (NLP) algorithms to the acquired text data. It uses the TfidfVectorizer from the Scikit-learn library to extract important keywords and topics from the text data. The input of this step is the text data, and the output is the extracted keywords and topics.
[1003] Step 5:
[1004] The server classifies search results into specific categories based on the extracted image features and text keywords and topics. Based on features such as "red," "shirt," and "cotton," products are automatically sorted into categories such as "clothing," "shirt," and "red." The input for this step is feature data, keywords, and topics, and the output is search results classified by category.
[1005] Step 6:
[1006] The categorized search results are sent to a visual review interface, where administrators or users can review the results and make corrections as needed. The input is the categorized search results, and the output is the reviewed and corrected search results.
[1007] Step 7:
[1008] The verified and corrected search results are sent from the server to the smartphone. The data is sent in JSON format, and the smartphone application parses it and displays it to the user. The input is the verified search results, and the output is the final search results displayed on the user's smartphone.
[1009] The above is the specific processing flow in the system of this application example. The specific data processing and calculations performed at each step enable the user to efficiently obtain product information.
[1010] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1011] This invention relates to a system that automatically categorizes search results based on a user-entered search query. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions, helping to filter and customize the search results.
[1012] When a user enters a search query from their device, the data is sent to the server. The server searches a database based on the received search query to retrieve relevant image and text data. It then applies image recognition algorithms to the retrieved image data to extract features such as color, shape, and patterns from each image. It also applies natural language processing (NLP) algorithms to the text data to extract important keywords and topics.
[1013] At the same time, the emotion engine analyzes the user's input behavior and facial expressions to recognize their emotions. This emotion information is used by the server to filter search results. For example, if the user is feeling stressed, such as "urgent," information requiring a quick response will be displayed first.
[1014] The server categorizes search results into specific categories based on image and text features, keywords, and topics, and then sends the categorized results to an interface for visual review, where users or administrators can participate and check the accuracy of the classification. After the final review is complete, the server sends the organized search results to the user's device for display.
[1015] Example: A user searches for "red shirt"
[1016] 1. Enter your query:
[1017] A user enters the search query "red shirt" into a device.
[1018] The device sends a search query to the server.
[1019] 2. Obtaining and processing search results:
[1020] The server receives the search query and searches the database to retrieve image and text data related to "red shirt."
[1021] The server applies an image recognition algorithm to the images it acquires and extracts features such as "red," "short sleeves," and "design."
[1022] The server applies NLP algorithms to the text data it acquires to extract important keywords and topics such as "red shirt," "cotton," and "fashion."
[1023] 3. Use of Emotion Engine:
[1024] When a user enters a search query, the emotion engine analyzes the user's facial expressions and typing speed to recognize their emotion.
[1025] The emotion engine provides emotion information such as whether the user is "excited" or "stressed" to the server.
[1026] 4. Automatic category assignment and filtering:
[1027] The server assigns categories such as "clothes," "shirt," and "red" based on the extracted features, keywords, and topics.
[1028] Based on the emotional information provided by the emotion engine, filtering is performed to prioritize and display search results that best suit the user's emotions.
[1029] 5. Final check and display:
[1030] The server then sends the compiled list of search results by category to an interface for visual review.
[1031] The user or administrator visually checks and corrects as necessary.
[1032] After final confirmation, the server sends the final search results to the user terminal and displays them on the user's screen.
[1033] This system allows users to easily browse products and information related to "red shirts" that are customized to their own emotions. It also reduces the burden of manual categorization, reducing operational costs. Furthermore, it is expected to improve the user experience by ensuring consistent quality of search results.
[1034] The processing flow will be explained below.
[1035] Step 1:
[1036] The user enters a search query into the device. The user types in "red shirt" and presses the search button.
[1037] Step 2:
[1038] The device sends the search query entered by the user to the server. The query sent is "red shirt."
[1039] Step 3:
[1040] The server analyzes the received search query. The query analysis module extracts the keyword "red shirt."
[1041] Step 4:
[1042] The server executes a search query against the database to retrieve relevant image and text data, such as product information and images related to "red shirt."
[1043] Step 5:
[1044] The server sends the acquired image data to an image recognition module, which extracts features such as color, shape, and pattern from each image.
[1045] Step 6:
[1046] The image recognition module returns the extracted features to the server. Specifically, it recognizes that each image has features such as "red," "shirt," and "short sleeves."
[1047] Step 7:
[1048] The server sends the acquired text data to a natural language processing (NLP) module, which extracts important keywords and topics from the text data.
[1049] Step 8:
[1050] The NLP module returns the extracted keywords and topics to the server. Specifically, keywords such as "red shirt," "cotton," and "fashion" are extracted.
[1051] Step 9:
[1052] The server integrates the features, keywords, and topics returned by the image recognition and NLP modules.
[1053] Step 10:
[1054] The server uses the combined information to categorize the search results into specific categories, such as "clothes," "shirts," and "red."
[1055] Step 11:
[1056] The server instructs the emotion engine to acquire emotion data when the user enters a search query. The emotion engine analyzes the user's facial expressions and typing speed to identify the emotion.
[1057] Step 12:
[1058] The emotion engine returns the analyzed emotion data to the server and recognizes whether the user is feeling "excited" or "stressed."
[1059] Step 13:
[1060] The server uses the emotion data to filter search results that match the recognized emotion. For example, if the user is excited, information that matches fashion trends will be displayed preferentially.
[1061] Step 14:
[1062] The server sends the categorized and filtered search results to a visual interface for review by the user or administrator.
[1063] Step 15:
[1064] A user or administrator visually inspects the results and corrects any incorrectly classified results.
[1065] Step 16:
[1066] After the final confirmation is completed, the server sends the final organized search results to the user terminal.
[1067] Step 17:
[1068] The device displays the final results on the user's screen, allowing them to view accurate and emotionally relevant information about the "red shirt."
[1069] This series of steps results in efficient search results for users, reduces the burden of manual categorization, and leverages the emotion engine to improve the quality of the user experience.
[1070] Example 2
[1071] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1072] Conventional search systems only retrieve information related to the query entered by the user, and do not adequately provide search results that reflect the user's emotions or circumstances. Furthermore, proper categorization of the retrieved search results requires manual work, resulting in high operational costs and inconsistent search accuracy. Furthermore, insufficient customization based on user emotions makes it difficult to improve the user experience.
[1073] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1074] In this invention, the server includes means for receiving a search query entered by a user, means for searching a database based on the search query to obtain related data, means for applying an image recognition algorithm to the obtained images to extract features, means for applying a natural language processing algorithm to the obtained text data to extract keywords and topics, means for classifying search results into specific categories based on the features, keywords, and topics, means for recognizing a user's emotion, means for filtering the search results based on the recognized user's emotion, means for sending the classified search results to an interface for visual confirmation, and means for displaying the confirmed search results to the user. This makes it possible to provide search results customized according to the user's emotion, reduce the burden of manual categorization, reduce operating costs, and improve the quality of search results.
[1075] A "search query" is a string of characters that a user enters to obtain information.
[1076] A "database" is a collection of information that is accessed based on a search query.
[1077] An "image recognition algorithm" is a technique for extracting features from image data.
[1078] "Features" are unique information extracted from image or text data.
[1079] "Natural language processing algorithms" are techniques for extracting keywords and topics from text data.
[1080] A "keyword" is an important word in text data.
[1081] A "topic" is the main theme or content that the text data deals with.
[1082] A "category" is a group for classifying search results.
[1083] "Emotion engine" is a technology that analyzes and recognizes the user's emotions.
[1084] "Filtering" is the process of selecting data based on specific criteria.
[1085] The "visual confirmation interface" is a display screen that allows users and administrators to check and modify search results.
[1086] "Means for displaying to the user" refers to the technology for displaying the final search results on the user's terminal.
[1087] This invention relates to a system that automatically categorizes search results based on a search query entered by a user, and then recognizes the user's emotions to provide customized search results. Specific processing for this includes the following steps:
[1088] When a user inputs a search query from a device, the device sends the data to the server. The server searches the database based on the received search query to retrieve related image and text data. For example, if the search query is "red shirt," the server extracts image and text data related to "red shirt" from the database.
[1089] The server applies image recognition algorithms to the acquired image data and extracts features such as color, shape, and pattern from each image. Libraries such as OpenCV and TensorFlow can be used for image recognition. For example, features such as "red," "short sleeves," and "design" can be extracted from an image.
[1090] The server then applies natural language processing (NLP) algorithms to the text data to extract important keywords and topics. NLP algorithms such as spaCy and NLTK are available. For example, keywords such as "red shirt," "cotton," and "fashion" are extracted from the text data.
[1091] At the same time, the emotion engine analyzes the user's input behavior and facial expressions to recognize their emotions. This emotion recognition is performed using software such as Emotion AI. As the user types a search query, the device's camera captures the user's facial expressions and identifies emotions such as "excitement" or "stress."
[1092] The server categorizes search results into specific categories based on image and text features, keywords, and topics, and also filters search results based on the user's emotional state. For example, if a user is feeling stressed, it prioritizes simple, easy-to-understand information.
[1093] The classified search results are sent to a visual review interface where the user or administrator can review the results and make any necessary corrections. The final review is then sent to the user's device and displayed on the user's screen.
[1094] Specific examples
[1095] The specific flow when a user searches for "red shirt" is as follows.
[1096] 1. Enter your query:
[1097] A user types "red shirt" into the device's search bar, and the device generates an HTTP request and sends it to the server.
[1098] 2. Obtaining and processing search results:
[1099] The server analyzes the received query and retrieves product data related to the "red shirt" from the database using an SQL query. The retrieved data includes images and text, and can also be searched using a search engine such as Apache Solr.
[1100] 3. Image Recognition and Text Analysis:
[1101] The server uses OpenCV to apply image recognition algorithms to the acquired images to extract features such as color, shape, and pattern.The acquired text data is then analyzed using spaCy and NLTK to extract important keywords.
[1102] 4. Use of Emotion Engine:
[1103] While the user is entering a query, the device's camera captures the user's facial expressions, which Emotion AI analyzes and sends emotional information to the server.
[1104] 5. Automatic category assignment and filtering:
[1105] The server assigns categories such as "clothing," "shirt," and "red" based on features obtained from image and text analysis, keywords, and emotional information provided by the emotion engine, and then prioritizes displaying search results that best match the user's emotions.
[1106] 6. Visual inspection:
[1107] The server generates search results and sends them to a visual review interface where they can be reviewed by the user or administrator, and corrections can be made if necessary.
[1108] 7. Final check and display:
[1109] The server sends the final confirmed result to the user terminal, which displays it.
[1110] This allows users to easily browse products and information related to "red shirts" that are customized to their own emotions. Furthermore, it reduces the burden of manual categorization, which is expected to reduce operational costs and improve the quality of search results.
[1111] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1112] The flow of this system's program processing
[1113] Step 1:
[1114] The user enters a search query into the search bar on the device. After completing the input, the entered search query is sent from the device to the server as an HTTP request. The request contains the query content and the user ID. Input data: The search query entered by the user in the search bar (e.g., "red shirt"). Output data: The HTTP request sent to the server.
[1115] Step 2:
[1116] The server extracts the search query from the HTTP request received and searches the database based on the query. Apache Solr or other search engines may be used, and related images and text data are obtained as search results. Input data: HTTP request (search query) sent to the server. Output data: Related images and text data.
[1117] Step 3:
[1118] The server applies an image recognition algorithm (e.g., OpenCV) to the image data it acquires to extract features such as color, shape, and pattern. This analyzes the image data and obtains each feature (e.g., "red," "short sleeves," "design"). Input data: acquired image data. Output data: extracted image feature information.
[1119] Step 4:
[1120] A natural language processing algorithm (e.g., spaCy or NLTK) is applied to the acquired text data to extract important keywords and topics. This allows related keywords (e.g., "red shirt," "cotton," "fashion") to be obtained from the text data. Input data: The acquired text data. Output data: The keywords and topics of the extracted text.
[1121] Step 5:
[1122] The emotion engine analyzes the user's input behavior and facial expression data to recognize the user's emotions. While the user is entering a query, the device's camera captures their facial expressions, which are then analyzed by emotion analysis software such as Emotion AI. Input data: User's input behavior and facial expression data. Output data: Recognized user emotional information (e.g., "excitement" or "stress").
[1123] Step 6:
[1124] The server classifies search results into specific categories based on image and text features, keywords and topics, and user sentiment information. For example, categories such as "clothing," "shirt," and "red" are assigned. Input data: Image and text feature information, and user sentiment information. Output data: Search results classified by category.
[1125] Step 7:
[1126] The server filters search results based on the user's emotional information and prioritizes displaying the information most appropriate for the user. Information is prioritized, taking into account emotions such as stress and excitement in particular. Input data: Search results categorized. Output data: Filtered search results.
[1127] Step 8:
[1128] The server sends the categorized search results to an interface for visual confirmation. Here, a web console or similar is used, allowing the user or administrator to check the search results and make corrections as necessary. Input data: Filtered search results. Output data: Search results displayed in the visual confirmation interface.
[1129] Step 9:
[1130] The user or administrator checks the search results and makes any necessary corrections. Based on this, the server generates the final confirmed search results. Input data: Search results displayed in the visual confirmation interface. Output data: Final corrected search results.
[1131] Step 10:
[1132] The server generates the final search results as an HTTP response and sends it to the user's device. The device displays the received search results on the screen, allowing the user to view customized search results. Input data: Final confirmed search results. Output data: Final search results displayed on the user's device.
[1133] This allows users to efficiently obtain and view search results that are customized to their own emotions. It also reduces the burden of manual categorization, which is expected to reduce operational costs and improve the quality of search results.
[1134] (Application example 2)
[1135] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1136] Conventional search systems can appropriately categorize related data in response to a user's search query, but they are unable to provide optimal search results that take the user's emotions into account. In particular, when a user is in a specific emotional state, such as stress or excitement, it is difficult to provide search results that are optimal for that emotion. This leads to problems such as reduced user satisfaction and an inability to obtain effective search results.
[1137] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1138] In this invention, the server includes means for receiving a search query entered by a user, means for searching a database based on the search query to obtain related images and text data, means for applying an image recognition algorithm to the obtained images to extract features, means for applying a natural language processing algorithm to the obtained text data to extract keywords and topics, means for recognizing the user's emotions, means for filtering search results based on the recognized emotions, means for classifying search results into specific categories based on the features, keywords, and topics, means for sending the classified search results to an interface for visual confirmation, and means for displaying the confirmed search results to the user, thereby enabling the provision of optimal search results that take the user's emotional state into consideration.
[1139] A "user-entered search query" is a word or phrase that a user enters into a system to request information.
[1140] A "database" is a collection of information that organizes and stores related information and provides data in response to search queries.
[1141] An "image recognition algorithm" is a computational procedure for automatically analyzing and extracting specific features or patterns from image data.
[1142] A "natural language processing algorithm" is a computational procedure for analyzing and extracting meaning, keywords, and topics from text data.
[1143] "Emotion recognition means" is a technology for determining and analyzing a user's emotional state from their facial expressions and input actions.
[1144] A "search result filtering method" is a technique for sorting and ordering search results based on specific criteria or conditions.
[1145] "Means for categorizing" refers to technology for automatically classifying search result images and text data into specific categories.
[1146] A "visual review interface" is a screen or tool that allows a user or administrator to review and correct the classification and accuracy of search results.
[1147] "Means of displaying to the user" refers to the technology used to display the organized and filtered search results on the user's device.
[1148] This invention relates to a system that automatically categorizes related information based on a search query entered by a user and provides search results that take the user's emotions into consideration. In this system, a server processes and analyzes data using various algorithms.
[1149] The server first receives a search query from the user's device. Based on the received search query, the server searches the database to retrieve relevant image and text data. Image recognition algorithms are applied to the retrieved image data to extract features such as color, shape, and pattern. Natural language processing (NLP) algorithms are applied to the retrieved text data to extract important keywords and topics.
[1150] To recognize the user's emotions, the server uses an emotion recognition means that captures the user's facial expressions with a camera and analyzes them using a machine learning model to determine the user's emotional state. It also analyzes the user's input behavior, such as typing speed on a keyboard or touchscreen.
[1151] Next, the server filters the search results based on the emotion information provided by the emotion recognition means. For example, if the user is in an "urgent" state, it prioritizes displaying information that requires a quick response, and if the user is "excited," it prioritizes displaying information about popular or new products.
[1152] The server classifies the search results into specific categories based on the extracted features and keywords. The classified search results are sent to a visual review interface for review by the user or administrator. Once review is complete, the final search results are sent to the user's terminal and displayed to the user.
[1153] The system uses the following hardware and software:
[1154] Hardware:
[1155] 1. Smartphone camera - used to capture the user's facial expressions.
[1156] 2. Server - Used to process and manage data.
[1157] software:
[1158] 1. OpenCV - Used for camera capture and face detection to recognize the user's facial expressions.
[1159] 2. TensorFlow / Keras - Used for machine learning models for emotion recognition.
[1160] 3. Transformers - Used for natural language processing algorithms.
[1161] 4. Requests - Used to communicate with the server.
[1162] As a specific example, if a user searches for "new smartphone cases" and it is recognized that the user is excited at that time, the search results will prioritize smartphone cases with high popularity rankings and the latest designs. An example of an input prompt for a generative AI model is, "Write the code for a Python program that will output search results appropriate for when the user is searching for new smartphone cases and is in an excited state."
[1163] This makes it possible to provide personalized search results that reflect the user's emotional state, improving the user experience.
[1164] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1165] Step 1:
[1166] A search query entered by a user is received at the terminal.
[1167] A user types "new phone case" into the search bar of their smartphone and presses the search button. This entered search query ("new phone case") is sent from the device to the server. The input data is a search query, and the server processes this query as received data.
[1168] Step 2:
[1169] The server searches a database based on the received search query to retrieve relevant image and text data.
[1170] Based on the search query "new smartphone case," the server searches for related information in the database. It retrieves multiple image and text data (product names, descriptions, etc.) from the database. The input data is the search query, and the output data is the related image and text data.
[1171] Step 3:
[1172] The server applies an image recognition algorithm to the image data it acquires and extracts features.
[1173] The server uses OpenCV to perform image recognition on the acquired image data, extracting features such as color, shape, and pattern from the image data. The input data is the image data, and the output data is the extracted feature data.
[1174] Step 4:
[1175] The server applies natural language processing algorithms to the text data it acquires to extract keywords and topics.
[1176] The server uses Transformers to perform natural language analysis on the acquired text data, extracting important keywords and topics from the text data. The input data is the text data, and the output data is the extracted keywords and topics.
[1177] Step 5:
[1178] The server uses an emotion recognition means to recognize the user's emotion.
[1179] When a user enters a search query, the device's camera captures their facial expression. The server uses TensorFlow / Keras to analyze the captured facial expression data and determine the user's emotional state. It also uses behavioral data such as input speed for analysis. The input data are facial expression capture data and input behavior data, and the output data is the user's emotional state.
[1180] Step 6:
[1181] The server filters the search results based on the emotion information provided by the emotion recognition means.
[1182] Based on the recognized emotion information, the server filters the extracted search results. For example, if the user is "excited," popular products and the latest designs are displayed preferentially. The input data are emotion information and search result data, and the output data are the filtered search results.
[1183] Step 7:
[1184] The server categorizes the search results into specific categories based on features, keywords and topics.
[1185] Based on the image and text data of the filtered search results, the server classifies each search result into an appropriate category (e.g., "smartphone cases," "new products," etc.). The input data is the search result data, and the output data is the search results classified by category.
[1186] Step 8:
[1187] The server sends the categorized search results to an interface for visual review.
[1188] The server sends the categorized search results to an interface where the user or administrator can check them. The input data is the categorized search results, and the output data is an interface for visual confirmation.
[1189] Step 9:
[1190] The server displays the verified search results to the user.
[1191] After the visual confirmation is completed, the server sends the final search results to the user's terminal and displays them on the screen. The input data are the search results after visual confirmation, and the output data are the search results displayed on the user's screen.
[1192] The above are the specific processing steps of the system that realizes the application example.
[1193] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1194] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1195] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1196] [Fourth embodiment]
[1197] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1198] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1199] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1200] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1201] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1202] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1203] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1204] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1205] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1206] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1207] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1208] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1209] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1210] This invention relates to a system that automatically categorizes search results based on a user-entered search query. By using an algorithm to efficiently categorize search results, it is possible to significantly reduce manual labor, lower operational costs, and maintain consistent search result quality. The invention operates as follows.
[1211] When a user enters a search query from their device, the data is sent to the server. The server searches a database based on the received search query to retrieve relevant image and text data. It then applies image recognition algorithms to the retrieved image data to extract features such as color, shape, and patterns from each image. In parallel, it applies natural language processing (NLP) algorithms to the text data to extract important keywords and topics.
[1212] The server classifies search results into specific categories based on image and text features, keywords, and topics. The classified search results are sent to an interface for visual review before being displayed to the user. Visual review involves the user or administrator, who corrects incorrectly classified results. After final review is complete, the server sends the organized search results to the user's device for display.
[1213] Specific examples are shown below.
[1214] Example: A user searches for "red shirt"
[1215] 1. Enter your query:
[1216] A user enters the search query "red shirt" into a device.
[1217] The device sends a search query to the server.
[1218] 2. Obtaining and processing search results:
[1219] The server receives the search query and searches the database to retrieve image and text data related to "red shirt."
[1220] The server applies an image recognition algorithm to the images it acquires and extracts features such as "red," "short sleeves," and "design."
[1221] The server applies NLP algorithms to the text data it acquires to extract important keywords and topics such as "red shirt," "cotton," and "fashion."
[1222] 3. Automatic Category Assignment:
[1223] The server assigns categories such as "clothes," "shirt," and "red" based on the extracted features, keywords, and topics.
[1224] Group search results by category and store them in a new data structure.
[1225] 4. Final check and display:
[1226] The server then sends the compiled list of search results by category to an interface for visual review.
[1227] The user or administrator visually checks and corrects as necessary.
[1228] After final confirmation, the server sends the final search results to the user terminal and displays them on the user's screen.
[1229] This system allows users to easily browse products and information related to "red shirts." It also reduces the burden of manual categorization, reducing operational costs. Furthermore, it is expected to improve the user experience by ensuring consistent quality of search results.
[1230] The processing flow will be explained below.
[1231] Step 1:
[1232] The user enters a search query into the device. The user types in "red shirt" and presses the search button.
[1233] Step 2:
[1234] The device sends the search query entered by the user to the server. The query sent is "red shirt."
[1235] Step 3:
[1236] The server analyzes the received search query. The query analysis module extracts the keyword "red shirt."
[1237] Step 4:
[1238] The server executes a search query against the database to retrieve relevant image and text data, such as product information and images related to "red shirt."
[1239] Step 5:
[1240] The server sends the acquired image data to an image recognition module, which extracts features such as color, shape, and pattern from each image.
[1241] Step 6:
[1242] The image recognition module returns the extracted features to the server. Specifically, it recognizes that each image has features such as "red," "shirt," and "short sleeves."
[1243] Step 7:
[1244] The server sends the acquired text data to a natural language processing (NLP) module, which extracts important keywords and topics from the text data.
[1245] Step 8:
[1246] The NLP module returns the extracted keywords and topics to the server. Specifically, keywords such as "red shirt," "cotton," and "fashion" are extracted.
[1247] Step 9:
[1248] The server integrates the features, keywords, and topics returned by the image recognition and NLP modules.
[1249] Step 10:
[1250] The server uses the combined information to categorize the search results into specific categories, such as "clothes," "shirts," and "red."
[1251] Step 11:
[1252] The server sends the search results, organized by category, to a visual interface for review by the user or administrator.
[1253] Step 12:
[1254] A user or administrator visually inspects the results and corrects any incorrectly classified results.
[1255] Step 13:
[1256] After the final confirmation is completed, the server sends the final organized search results to the user terminal.
[1257] Step 14:
[1258] The device displays the final result on the user's screen, where the user can view accurate information about the "red shirt."
[1259] This series of steps allows for efficient delivery of search results to users and also reduces the burden of manual categorization.
[1260] Example 1
[1261] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1262] Existing search systems have had difficulty efficiently classifying and organizing information related to user-entered search queries. In particular, the integrated classification of image and text data relies on manual work, which is time-consuming and results in high operational costs. Furthermore, manual categorization often results in inconsistent quality of search results. The purpose of this invention is to solve the above problems and achieve efficient classification of search results while maintaining high quality.
[1263] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1264] In this invention, the server includes means for receiving a search query entered by a user, means for searching a database based on the search query to obtain related visual data and text data, and means for applying an image recognition algorithm to the obtained visual data to extract features, thereby enabling automatic acquisition of visual and text data related to the search query and efficient extraction of its features.
[1265] The server includes means for applying a natural language processing algorithm to the acquired text data to extract keywords and topics, means for classifying search results into specific categories based on the features, keywords, and topics, means for sending the classified search results to an interface for visual confirmation, and means for displaying the confirmed search results to the user. This enables automatic classification of search results and final confirmation by the user or administrator, and provides accurate, high-quality search results to the user.
[1266] A "user" is an entity that enters a search query into the system and receives the results.
[1267] A "search query" is an instruction for searching information that a user inputs to a system.
[1268] "Database" means a data storage that stores and retrieves relevant data based on a search query.
[1269] "Visual data" refers to visual information such as images and illustrations.
[1270] "Text data" refers to information expressed in text format.
[1271] An "image recognition algorithm" is a software process for extracting features such as color, shape, and pattern from image data.
[1272] A "natural language processing algorithm" is a software process for extracting keywords and topics from text data.
[1273] "Features" refer to visual attributes extracted by image recognition algorithms.
[1274] "Keywords" refer to important words or terms extracted by natural language processing algorithms.
[1275] "Topics" refer to themes or topics extracted by natural language processing algorithms.
[1276] A "category" is a group for classifying search results.
[1277] A "visual review interface" is a screen or tool that allows a user or administrator to review and modify search results.
[1278] "Means for displaying to the user" refers to the mechanism for displaying the verified search results on the user's device.
[1279] This invention relates to a system that automatically classifies and organizes related visual and textual data based on a user-entered search query. The system utilizes image recognition and natural language processing algorithms to achieve high-quality classification of search results.
[1280] Hardware and software used
[1281] The system hardware includes a device (such as a PC or smartphone) where users enter search queries, and a server that searches the database and processes the results. The server is a powerful machine that provides high-speed data processing and search functions.
[1282] The system's software includes the following components:
[1283] 1. Database search engine: An engine is needed to quickly retrieve relevant data based on a search query, such as various commercial databases.
[1284] 2. Image recognition algorithm: To extract features from the acquired visual data, we use image recognition libraries such as TensorFlow and OpenCV.
[1285] 3. Natural Language Processing (NLP) algorithms: We use libraries such as spaCy and hugface Transformers to extract keywords and topics from the acquired text data.
[1286] Specific examples
[1287] Take the example of a user entering the search query "red shirt" on a terminal.
[1288] 1. Enter your query:
[1289] A user enters the search query "red shirt" into a device.
[1290] The device sends a search query to the server.
[1291] 2. Obtaining and processing search results:
[1292] The server receives the search query and searches the database to retrieve image and text data related to "red shirt."
[1293] The server uses TensorFlow and OpenCV to perform image recognition on the images it acquires, extracting features such as color, shape, and pattern.
[1294] The server performs natural language processing on the text data it acquires using spaCy and hugface Transformers to extract important keywords and topics such as "red shirt," "cotton," and "fashion."
[1295] 3. Automatic Category Assignment:
[1296] The server classifies the search results into categories such as "clothes," "shirts," and "red color" based on the extracted features, keywords, and topics.
[1297] The server groups the search results by category and stores them in a new data structure.
[1298] 4. Final check and display:
[1299] The server then sends the compiled list of search results by category to an interface for visual review.
[1300] The user or administrator visually checks the results and makes any necessary corrections. Once the final check is complete, the server sends the organized search results to the user's device and displays them on the user's screen.
[1301] Prompt Sentence Examples
[1302] "Category search results based on the query: red shirt"
[1303] In this way, users can easily browse products and information related to "red shirts," significantly reducing the burden of manual categorization, and the quality of search results is maintained at a consistent level, improving the user experience.
[1304] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1305] Step 1:
[1306] A user enters a search query on a device.
[1307] Input: A search query, such as "red shirt," that a user types into the search bar.
[1308] What happens: The user opens a browser or application on their device and enters the desired information into the search bar.
[1309] Output: The entered search query is stored in the message to be sent by the terminal.
[1310] Step 2:
[1311] The device sends a search query to the server.
[1312] Input: The search query entered by the user in step 1.
[1313] Specific operation: The terminal generates an HTTP request and sends request data including a search query to the server.
[1314] Output: The server receives the search query.
[1315] Step 3:
[1316] The server receives the search query and searches the database.
[1317] Input: The search query received from the device.
[1318] Specific operation: The server generates a database query to search for relevant data.
[1319] Output: Visual data (images) and textual data (text) related to the search query are obtained.
[1320] Step 4:
[1321] The server retrieves the visual and textual data from the search results.
[1322] Input: Database search results containing visual and textual data that match the search query.
[1323] Specific operation: The server extracts relevant image and text data from the database.
[1324] Output: The extracted visual and textual data is stored in memory.
[1325] Step 5:
[1326] The server applies image recognition algorithms to the visual data to extract features.
[1327] Input: Acquired visual data.
[1328] Specific operation: The server uses image recognition libraries such as TensorFlow and OpenCV to analyze visual data.
[1329] Output: Features extracted from visual data (e.g., color, shape, pattern).
[1330] Step 6:
[1331] The server applies natural language processing algorithms to the text data to extract keywords and topics.
[1332] Input: Acquired sentence data.
[1333] Specific operation: The server analyzes the text data using natural language processing algorithms such as spaCy and HugFace Transformers.
[1334] Output: Keywords and topics extracted from the text data.
[1335] Step 7:
[1336] The server categorizes search results into specific categories based on features, keywords, and topics.
[1337] Input: Features extracted from visual data, and keywords and topics extracted from text data.
[1338] What it does: The server automatically categorizes the results based on a predefined list of categories.
[1339] Output: Search results organized by category.
[1340] Step 8:
[1341] The server sends the categorized search results to an interface for visual review.
[1342] Input: Search results organized by category.
[1343] Specific operation: The server performs data conversion to send the search results to the visual confirmation interface in an appropriate format.
[1344] Output: Search results displayed in a visual interface.
[1345] Step 9:
[1346] The user or administrator visually checks and corrects as necessary.
[1347] Input: Search results displayed on the visual interface.
[1348] Specific Action: A user or administrator reviews search results through the interface and corrects any incorrectly classified items.
[1349] Output: The final confirmed search results.
[1350] Step 10:
[1351] The server sends the final confirmed search results to the user's device.
[1352] Input: Last confirmed search results.
[1353] Specific operation: The server converts the search results into a data format suitable for transmission to the user terminal, and transmits the results.
[1354] Output: The final search results that are displayed on the user's device.
[1355] Step 11:
[1356] The user views the search results.
[1357] Input: The final search results displayed on the user's device.
[1358] Specific operation: The user views the search results displayed on the device and obtains the required information.
[1359] Output: Information that leads to a user's decision or next action.
[1360] (Application example 1)
[1361] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1362] Existing search systems face the challenge of effectively categorizing and displaying relevant image and text data in response to user queries. Manual categorization of search results is also labor-intensive and expensive. Smartphone shopping applications, in particular, need to efficiently manage large amounts of search results and provide a user-friendly experience.
[1363] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1364] In this invention, the server includes: means for receiving a search query entered by a user; means for searching a database based on the search query to acquire related images and text data; means for applying an image recognition algorithm to the acquired images to extract features; means for applying a natural language processing algorithm to the acquired text data to extract keywords and topics; means for categorizing search results into specific categories based on the features, keywords, and topics; means for sending the categorized search results to an interface for visual confirmation; means for displaying the confirmed search results to the user; and means for categorizing the search results and returning them to the user in JSON format. This automates the categorization of search results by category, reducing manual work and operational costs. Furthermore, users can quickly and efficiently browse related product information through a smartphone mail-order application.
[1365] "User" refers to a person who enters a search query using a device, including a smartphone.
[1366] A "search query" refers to a search term or phrase that a user enters into a search box.
[1367] "Database" refers to an information management system that stores related image and text data and provides data based on a search query.
[1368] An "image recognition algorithm" refers to a computational method for analyzing and extracting features such as color, shape, and pattern from acquired image data.
[1369] "Features" refer to specific data attributes such as color, shape, and pattern extracted using image recognition algorithms.
[1370] A "natural language processing algorithm" refers to a computational method for analyzing and extracting keywords and topics from acquired text data.
[1371] "Keywords" refer to important words and phrases extracted by natural language processing algorithms.
[1372] "Topic" refers to the main subject or theme of text data analyzed by natural language processing algorithms.
[1373] "Category" refers to various groups that categorize search results based on extracted features, keywords, and topics.
[1374] "Search Results" refers to a collection of relevant image and text data retrieved from a database based on a user's search query.
[1375] "Visual confirmation interface" refers to the display means by which a user or administrator can confirm and modify classified search results.
[1376] "JSON format" refers to a method of structuring and representing data in JavaScript Object Notation format.
[1377] "Smartphone online shopping application" refers to application software that runs on a smartphone and allows users to search and view related product information by entering a search query.
[1378] The present invention relates to a system that provides a mail-order application for smartphones that allows a user to input a search query, efficiently categorize related product information based on the query, and make the information available for browsing.
[1379] This system works as follows: When a user enters a search query using a smartphone, the data is sent to the server. Based on the entered query, the server searches a database to retrieve relevant image and text data. An image recognition algorithm is applied to the retrieved image data to extract features such as color, shape, and pattern. In parallel, a natural language processing algorithm is applied to the retrieved text data to extract important keywords and topics. Based on this, the server classifies the search results into specific categories. The classified search results are sent to a visual confirmation interface where the user or administrator can review and modify them. The confirmed search results are finally sent to the user's smartphone and displayed in JSON format.
[1380] The hardware used includes the following: a smartphone, which acts as a mobile device for entering search queries and displaying results; a server, which processes and stores the data; and finally, a database, which stores the images and text data to be searched.
[1381] The software used includes the following: Flask (Python) is used as the web application development framework; the OpenCV library is used for image recognition algorithms, and TfidfVectorizer from the Scikit-learn library is used for natural language processing; and the JSON format is used to organize and transmit data structures.
[1382] Examples:
[1383] A user enters the query "red shirt" into an online shopping application on their smartphone. The server retrieves image and text data related to "red shirt" from the database and analyzes the data. The retrieved image data is subjected to an image recognition algorithm to extract features such as "red," "shirt," and "design." The text data is analyzed using an NLP algorithm to extract keywords such as "red shirt," "cotton," and "fashion." Based on these features and keywords, the server classifies search results into categories such as "clothing," "shirt," and "red." The categorized search results are then sent to a visual review interface before being displayed to the user, where the user or administrator can make any necessary corrections. Finally, the reviewed and corrected search results are sent in JSON format to the user's smartphone and displayed.
[1384] Example prompt sentence:
[1385] When a user searches for "red shirt," related products are automatically categorized into categories such as "red," "shirt," and "clothing" based on the query and displayed to the user.
[1386] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1387] Step 1:
[1388] A user enters a search query into a shopping application on their smartphone and sends it to a server. The input data includes search terms such as "red shirt." The server receives the query and stores it for further data processing.
[1389] Step 2:
[1390] Based on the received search query, the server searches a database containing relevant image and text data. In this step, the server executes the database query to retrieve product data related to "red shirt." The output is a set of image and text data.
[1391] Step 3:
[1392] The server applies an image recognition algorithm to the acquired image data. Specifically, it uses the OpenCV library to extract features such as color, shape, and pattern from the image. The input to this process is the image data, and the output is the extracted feature data.
[1393] Step 4:
[1394] The server applies natural language processing (NLP) algorithms to the acquired text data. It uses the TfidfVectorizer from the Scikit-learn library to extract important keywords and topics from the text data. The input of this step is the text data, and the output is the extracted keywords and topics.
[1395] Step 5:
[1396] The server classifies search results into specific categories based on the extracted image features and text keywords and topics. Based on features such as "red," "shirt," and "cotton," products are automatically sorted into categories such as "clothing," "shirt," and "red." The input for this step is feature data, keywords, and topics, and the output is search results classified by category.
[1397] Step 6:
[1398] The categorized search results are sent to a visual review interface, where administrators or users can review the results and make corrections as needed. The input is the categorized search results, and the output is the reviewed and corrected search results.
[1399] Step 7:
[1400] The verified and corrected search results are sent from the server to the smartphone. The data is sent in JSON format, and the smartphone application parses it and displays it to the user. The input is the verified search results, and the output is the final search results displayed on the user's smartphone.
[1401] The above is the specific processing flow in the system of this application example. The specific data processing and calculations performed at each step enable the user to efficiently obtain product information.
[1402] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1403] This invention relates to a system that automatically categorizes search results based on a user-entered search query. Furthermore, the system incorporates an emotion engine that recognizes the user's emotions, helping to filter and customize the search results.
[1404] When a user enters a search query from their device, the data is sent to the server. The server searches a database based on the received search query to retrieve relevant image and text data. It then applies image recognition algorithms to the retrieved image data to extract features such as color, shape, and patterns from each image. It also applies natural language processing (NLP) algorithms to the text data to extract important keywords and topics.
[1405] At the same time, the emotion engine analyzes the user's input behavior and facial expressions to recognize their emotions. This emotion information is used by the server to filter search results. For example, if the user is feeling stressed, such as "urgent," information requiring a quick response will be displayed first.
[1406] The server categorizes search results into specific categories based on image and text features, keywords, and topics, and then sends the categorized results to an interface for visual review, where users or administrators can participate and check the accuracy of the classification. After the final review is complete, the server sends the organized search results to the user's device for display.
[1407] Example: A user searches for "red shirt"
[1408] 1. Enter your query:
[1409] A user enters the search query "red shirt" into a device.
[1410] The device sends a search query to the server.
[1411] 2. Obtaining and processing search results:
[1412] The server receives the search query and searches the database to retrieve image and text data related to "red shirt."
[1413] The server applies an image recognition algorithm to the images it acquires and extracts features such as "red," "short sleeves," and "design."
[1414] The server applies NLP algorithms to the text data it acquires to extract important keywords and topics such as "red shirt," "cotton," and "fashion."
[1415] 3. Use of Emotion Engine:
[1416] When a user enters a search query, the emotion engine analyzes the user's facial expressions and typing speed to recognize their emotion.
[1417] The emotion engine provides emotion information such as whether the user is "excited" or "stressed" to the server.
[1418] 4. Automatic category assignment and filtering:
[1419] The server assigns categories such as "clothes," "shirt," and "red" based on the extracted features, keywords, and topics.
[1420] Based on the emotional information provided by the emotion engine, filtering is performed to prioritize and display search results that best suit the user's emotions.
[1421] 5. Final check and display:
[1422] The server then sends the compiled list of search results by category to an interface for visual review.
[1423] The user or administrator visually checks and corrects as necessary.
[1424] After final confirmation, the server sends the final search results to the user terminal and displays them on the user's screen.
[1425] This system allows users to easily browse products and information related to "red shirts" that are customized to their own emotions. It also reduces the burden of manual categorization, reducing operational costs. Furthermore, it is expected to improve the user experience by ensuring consistent quality of search results.
[1426] The processing flow will be explained below.
[1427] Step 1:
[1428] The user enters a search query into the device. The user types in "red shirt" and presses the search button.
[1429] Step 2:
[1430] The device sends the search query entered by the user to the server. The query sent is "red shirt."
[1431] Step 3:
[1432] The server analyzes the received search query. The query analysis module extracts the keyword "red shirt."
[1433] Step 4:
[1434] The server executes a search query against the database to retrieve relevant image and text data, such as product information and images related to "red shirt."
[1435] Step 5:
[1436] The server sends the acquired image data to an image recognition module, which extracts features such as color, shape, and pattern from each image.
[1437] Step 6:
[1438] The image recognition module returns the extracted features to the server. Specifically, it recognizes that each image has features such as "red," "shirt," and "short sleeves."
[1439] Step 7:
[1440] The server sends the acquired text data to a natural language processing (NLP) module, which extracts important keywords and topics from the text data.
[1441] Step 8:
[1442] The NLP module returns the extracted keywords and topics to the server. Specifically, keywords such as "red shirt," "cotton," and "fashion" are extracted.
[1443] Step 9:
[1444] The server integrates the features, keywords, and topics returned by the image recognition and NLP modules.
[1445] Step 10:
[1446] The server uses the combined information to categorize the search results into specific categories, such as "clothes," "shirts," and "red."
[1447] Step 11:
[1448] The server instructs the emotion engine to acquire emotion data when the user enters a search query. The emotion engine analyzes the user's facial expressions and typing speed to identify the emotion.
[1449] Step 12:
[1450] The emotion engine returns the analyzed emotion data to the server and recognizes whether the user is feeling "excited" or "stressed."
[1451] Step 13:
[1452] The server uses the emotion data to filter search results that match the recognized emotion. For example, if the user is excited, information that matches fashion trends will be displayed preferentially.
[1453] Step 14:
[1454] The server sends the categorized and filtered search results to a visual interface for review by the user or administrator.
[1455] Step 15:
[1456] A user or administrator visually inspects the results and corrects any incorrectly classified results.
[1457] Step 16:
[1458] After the final confirmation is completed, the server sends the final organized search results to the user terminal.
[1459] Step 17:
[1460] The device displays the final results on the user's screen, allowing them to view accurate and emotionally relevant information about the "red shirt."
[1461] This series of steps results in efficient search results for users, reduces the burden of manual categorization, and leverages the emotion engine to improve the quality of the user experience.
[1462] Example 2
[1463] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1464] Conventional search systems only retrieve information related to the query entered by the user, and do not adequately provide search results that reflect the user's emotions or circumstances. Furthermore, proper categorization of the retrieved search results requires manual work, resulting in high operational costs and inconsistent search accuracy. Furthermore, insufficient customization based on user emotions makes it difficult to improve the user experience.
[1465] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1466] In this invention, the server includes means for receiving a search query entered by a user, means for searching a database based on the search query to obtain related data, means for applying an image recognition algorithm to the obtained images to extract features, means for applying a natural language processing algorithm to the obtained text data to extract keywords and topics, means for classifying search results into specific categories based on the features, keywords, and topics, means for recognizing a user's emotion, means for filtering the search results based on the recognized user's emotion, means for sending the classified search results to an interface for visual confirmation, and means for displaying the confirmed search results to the user. This makes it possible to provide search results customized according to the user's emotion, reduce the burden of manual categorization, reduce operating costs, and improve the quality of search results.
[1467] A "search query" is a string of characters that a user enters to obtain information.
[1468] A "database" is a collection of information that is accessed based on a search query.
[1469] An "image recognition algorithm" is a technique for extracting features from image data.
[1470] "Features" are unique information extracted from image or text data.
[1471] "Natural language processing algorithms" are techniques for extracting keywords and topics from text data.
[1472] A "keyword" is an important word in text data.
[1473] A "topic" is the main theme or content that the text data deals with.
[1474] A "category" is a group for classifying search results.
[1475] "Emotion engine" is a technology that analyzes and recognizes the user's emotions.
[1476] "Filtering" is the process of selecting data based on specific criteria.
[1477] The "visual confirmation interface" is a display screen that allows users and administrators to check and modify search results.
[1478] "Means for displaying to the user" refers to the technology for displaying the final search results on the user's terminal.
[1479] This invention relates to a system that automatically categorizes search results based on a search query entered by a user, and then recognizes the user's emotions to provide customized search results. Specific processing for this includes the following steps:
[1480] When a user inputs a search query from a device, the device sends the data to the server. The server searches the database based on the received search query to retrieve related image and text data. For example, if the search query is "red shirt," the server extracts image and text data related to "red shirt" from the database.
[1481] The server applies image recognition algorithms to the acquired image data and extracts features such as color, shape, and pattern from each image. Libraries such as OpenCV and TensorFlow can be used for image recognition. For example, features such as "red," "short sleeves," and "design" can be extracted from an image.
[1482] The server then applies natural language processing (NLP) algorithms to the text data to extract important keywords and topics. NLP algorithms such as spaCy and NLTK are available. For example, keywords such as "red shirt," "cotton," and "fashion" are extracted from the text data.
[1483] At the same time, the emotion engine analyzes the user's input behavior and facial expressions to recognize their emotions. This emotion recognition is performed using software such as Emotion AI. As the user types a search query, the device's camera captures the user's facial expressions and identifies emotions such as "excitement" or "stress."
[1484] The server categorizes search results into specific categories based on image and text features, keywords, and topics, and also filters search results based on the user's emotional state. For example, if a user is feeling stressed, it prioritizes simple, easy-to-understand information.
[1485] The classified search results are sent to a visual review interface where the user or administrator can review the results and make any necessary corrections. The final review is then sent to the user's device and displayed on the user's screen.
[1486] Specific examples
[1487] The specific flow when a user searches for "red shirt" is as follows.
[1488] 1. Enter your query:
[1489] A user types "red shirt" into the device's search bar, and the device generates an HTTP request and sends it to the server.
[1490] 2. Obtaining and processing search results:
[1491] The server analyzes the received query and retrieves product data related to the "red shirt" from the database using an SQL query. The retrieved data includes images and text, and can also be searched using a search engine such as Apache Solr.
[1492] 3. Image Recognition and Text Analysis:
[1493] The server uses OpenCV to apply image recognition algorithms to the acquired images to extract features such as color, shape, and pattern.The acquired text data is then analyzed using spaCy and NLTK to extract important keywords.
[1494] 4. Use of Emotion Engine:
[1495] While the user is entering a query, the device's camera captures the user's facial expressions, which Emotion AI analyzes and sends emotional information to the server.
[1496] 5. Automatic category assignment and filtering:
[1497] The server assigns categories such as "clothing," "shirt," and "red" based on features obtained from image and text analysis, keywords, and emotional information provided by the emotion engine, and then prioritizes displaying search results that best match the user's emotions.
[1498] 6. Visual inspection:
[1499] The server generates search results and sends them to a visual review interface where they can be reviewed by the user or administrator, and corrections can be made if necessary.
[1500] 7. Final check and display:
[1501] The server sends the final confirmed result to the user terminal, which displays it.
[1502] This allows users to easily browse products and information related to "red shirts" that are customized to their own emotions. Furthermore, it reduces the burden of manual categorization, which is expected to reduce operational costs and improve the quality of search results.
[1503] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1504] The flow of this system's program processing
[1505] Step 1:
[1506] The user enters a search query into the search bar on the device. After completing the input, the entered search query is sent from the device to the server as an HTTP request. The request contains the query content and the user ID. Input data: The search query entered by the user in the search bar (e.g., "red shirt"). Output data: The HTTP request sent to the server.
[1507] Step 2:
[1508] The server extracts the search query from the HTTP request received and searches the database based on the query. Apache Solr or other search engines may be used, and related images and text data are obtained as search results. Input data: HTTP request (search query) sent to the server. Output data: Related images and text data.
[1509] Step 3:
[1510] The server applies an image recognition algorithm (e.g., OpenCV) to the image data it acquires to extract features such as color, shape, and pattern. This analyzes the image data and obtains each feature (e.g., "red," "short sleeves," "design"). Input data: acquired image data. Output data: extracted image feature information.
[1511] Step 4:
[1512] A natural language processing algorithm (e.g., spaCy or NLTK) is applied to the acquired text data to extract important keywords and topics. This allows related keywords (e.g., "red shirt," "cotton," "fashion") to be obtained from the text data. Input data: The acquired text data. Output data: The keywords and topics of the extracted text.
[1513] Step 5:
[1514] The emotion engine analyzes the user's input behavior and facial expression data to recognize the user's emotions. While the user is entering a query, the device's camera captures their facial expressions, which are then analyzed by emotion analysis software such as Emotion AI. Input data: User's input behavior and facial expression data. Output data: Recognized user emotional information (e.g., "excitement" or "stress").
[1515] Step 6:
[1516] The server classifies search results into specific categories based on image and text features, keywords and topics, and user sentiment information. For example, categories such as "clothing," "shirt," and "red" are assigned. Input data: Image and text feature information, and user sentiment information. Output data: Search results classified by category.
[1517] Step 7:
[1518] The server filters search results based on the user's emotional information and prioritizes displaying the information most appropriate for the user. Information is prioritized, taking into account emotions such as stress and excitement in particular. Input data: Search results categorized. Output data: Filtered search results.
[1519] Step 8:
[1520] The server sends the categorized search results to an interface for visual confirmation. Here, a web console or similar is used, allowing the user or administrator to check the search results and make corrections as necessary. Input data: Filtered search results. Output data: Search results displayed in the visual confirmation interface.
[1521] Step 9:
[1522] The user or administrator checks the search results and makes any necessary corrections. Based on this, the server generates the final confirmed search results. Input data: Search results displayed in the visual confirmation interface. Output data: Final corrected search results.
[1523] Step 10:
[1524] The server generates the final search results as an HTTP response and sends it to the user's device. The device displays the received search results on the screen, allowing the user to view customized search results. Input data: Final confirmed search results. Output data: Final search results displayed on the user's device.
[1525] This allows users to efficiently obtain and view search results that are customized to their own emotions. It also reduces the burden of manual categorization, which is expected to reduce operational costs and improve the quality of search results.
[1526] (Application example 2)
[1527] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1528] Conventional search systems can appropriately categorize related data in response to a user's search query, but they are unable to provide optimal search results that take the user's emotions into account. In particular, when a user is in a specific emotional state, such as stress or excitement, it is difficult to provide search results that are optimal for that emotion. This leads to problems such as reduced user satisfaction and an inability to obtain effective search results.
[1529] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1530] In this invention, the server includes means for receiving a search query entered by a user, means for searching a database based on the search query to obtain related images and text data, means for applying an image recognition algorithm to the obtained images to extract features, means for applying a natural language processing algorithm to the obtained text data to extract keywords and topics, means for recognizing the user's emotions, means for filtering search results based on the recognized emotions, means for classifying search results into specific categories based on the features, keywords, and topics, means for sending the classified search results to an interface for visual confirmation, and means for displaying the confirmed search results to the user, thereby enabling the provision of optimal search results that take the user's emotional state into consideration.
[1531] A "user-entered search query" is a word or phrase that a user enters into a system to request information.
[1532] A "database" is a collection of information that organizes and stores related information and provides data in response to search queries.
[1533] An "image recognition algorithm" is a computational procedure for automatically analyzing and extracting specific features or patterns from image data.
[1534] A "natural language processing algorithm" is a computational procedure for analyzing and extracting meaning, keywords, and topics from text data.
[1535] "Emotion recognition means" is a technology for determining and analyzing a user's emotional state from their facial expressions and input actions.
[1536] A "search result filtering method" is a technique for sorting and ordering search results based on specific criteria or conditions.
[1537] "Means for categorizing" refers to technology for automatically classifying search result images and text data into specific categories.
[1538] A "visual review interface" is a screen or tool that allows a user or administrator to review and correct the classification and accuracy of search results.
[1539] "Means of displaying to the user" refers to the technology used to display the organized and filtered search results on the user's device.
[1540] This invention relates to a system that automatically categorizes related information based on a search query entered by a user and provides search results that take the user's emotions into consideration. In this system, a server processes and analyzes data using various algorithms.
[1541] The server first receives a search query from the user's device. Based on the received search query, the server searches the database to retrieve relevant image and text data. Image recognition algorithms are applied to the retrieved image data to extract features such as color, shape, and pattern. Natural language processing (NLP) algorithms are applied to the retrieved text data to extract important keywords and topics.
[1542] To recognize the user's emotions, the server uses an emotion recognition means that captures the user's facial expressions with a camera and analyzes them using a machine learning model to determine the user's emotional state. It also analyzes the user's input behavior, such as typing speed on a keyboard or touchscreen.
[1543] Next, the server filters the search results based on the emotion information provided by the emotion recognition means. For example, if the user is in an "urgent" state, it prioritizes displaying information that requires a quick response, and if the user is "excited," it prioritizes displaying information about popular or new products.
[1544] The server classifies the search results into specific categories based on the extracted features and keywords. The classified search results are sent to a visual review interface for review by the user or administrator. Once review is complete, the final search results are sent to the user's terminal and displayed to the user.
[1545] The system uses the following hardware and software:
[1546] Hardware:
[1547] 1. Smartphone camera - used to capture the user's facial expressions.
[1548] 2. Server - Used to process and manage data.
[1549] software:
[1550] 1. OpenCV - Used for camera capture and face detection to recognize the user's facial expressions.
[1551] 2. TensorFlow / Keras - Used for machine learning models for emotion recognition.
[1552] 3. Transformers - Used for natural language processing algorithms.
[1553] 4. Requests - Used to communicate with the server.
[1554] As a specific example, if a user searches for "new smartphone cases" and it is recognized that the user is excited at that time, the search results will prioritize smartphone cases with high popularity rankings and the latest designs. An example of an input prompt for a generative AI model is, "Write the code for a Python program that will output search results appropriate for when the user is searching for new smartphone cases and is in an excited state."
[1555] This makes it possible to provide personalized search results that reflect the user's emotional state, improving the user experience.
[1556] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1557] Step 1:
[1558] A search query entered by a user is received at the terminal.
[1559] A user types "new phone case" into the search bar of their smartphone and presses the search button. This entered search query ("new phone case") is sent from the device to the server. The input data is a search query, and the server processes this query as received data.
[1560] Step 2:
[1561] The server searches a database based on the received search query to retrieve relevant image and text data.
[1562] Based on the search query "new smartphone case," the server searches for related information in the database. It retrieves multiple image and text data (product names, descriptions, etc.) from the database. The input data is the search query, and the output data is the related image and text data.
[1563] Step 3:
[1564] The server applies an image recognition algorithm to the image data it acquires and extracts features.
[1565] The server uses OpenCV to perform image recognition on the acquired image data, extracting features such as color, shape, and pattern from the image data. The input data is the image data, and the output data is the extracted feature data.
[1566] Step 4:
[1567] The server applies natural language processing algorithms to the text data it acquires to extract keywords and topics.
[1568] The server uses Transformers to perform natural language analysis on the acquired text data, extracting important keywords and topics from the text data. The input data is the text data, and the output data is the extracted keywords and topics.
[1569] Step 5:
[1570] The server uses an emotion recognition means to recognize the user's emotion.
[1571] When a user enters a search query, the device's camera captures their facial expression. The server uses TensorFlow / Keras to analyze the captured facial expression data and determine the user's emotional state. It also uses behavioral data such as input speed for analysis. The input data are facial expression capture data and input behavior data, and the output data is the user's emotional state.
[1572] Step 6:
[1573] The server filters the search results based on the emotion information provided by the emotion recognition means.
[1574] Based on the recognized emotion information, the server filters the extracted search results. For example, if the user is "excited," popular products and the latest designs are displayed preferentially. The input data are emotion information and search result data, and the output data are the filtered search results.
[1575] Step 7:
[1576] The server categorizes the search results into specific categories based on features, keywords and topics.
[1577] Based on the image and text data of the filtered search results, the server classifies each search result into an appropriate category (e.g., "smartphone cases," "new products," etc.). The input data is the search result data, and the output data is the search results classified by category.
[1578] Step 8:
[1579] The server sends the categorized search results to an interface for visual review.
[1580] The server sends the categorized search results to an interface where the user or administrator can check them. The input data is the categorized search results, and the output data is an interface for visual confirmation.
[1581] Step 9:
[1582] The server displays the verified search results to the user.
[1583] After the visual confirmation is completed, the server sends the final search results to the user's terminal and displays them on the screen. The input data are the search results after visual confirmation, and the output data are the search results displayed on the user's screen.
[1584] The above are the specific processing steps of the system that realizes the application example.
[1585] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1586] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1587] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1588] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1589] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1590] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1591] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1592] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1593] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1594] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1595] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1596] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1597] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1598] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1599] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1600] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1601] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1602] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1603] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1604] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1605] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1606] The following is further disclosed regarding the above embodiment.
[1607] (Claim 1)
[1608] means for receiving a search query entered by a user;
[1609] means for searching a database based on the search query to obtain relevant image and text data;
[1610] means for applying an image recognition algorithm to the acquired image to extract features;
[1611] A means for applying a natural language processing algorithm to the acquired text data to extract keywords and topics;
[1612] means for classifying search results into specific categories based on said features, keywords and topics;
[1613] means for transmitting the classified search results to an interface for visual review;
[1614] means for displaying the confirmed search results to the user;
[1615] A system including:
[1616] (Claim 2)
[1617] The system of claim 1, wherein the image recognition algorithm is used to identify color, shape, and pattern.
[1618] (Claim 3)
[1619] 10. The system of claim 1, wherein the natural language processing algorithm uses topic modeling techniques to extract keywords and topics from text data.
[1620] "Example 1"
[1621] (Claim 1)
[1622] means for receiving a search query entered by a user;
[1623] means for searching a database based on the search query to obtain relevant visual and textual data;
[1624] means for applying an image recognition algorithm to the acquired visual data to extract features;
[1625] A means for applying a natural language processing algorithm to the acquired text data to extract keywords and topics;
[1626] means for classifying search results into specific categories based on said features, keywords and topics;
[1627] means for transmitting the classified search results to an interface for visual review;
[1628] means for displaying the confirmed search results to the user;
[1629] A system including:
[1630] (Claim 2)
[1631] 10. The system of claim 1, wherein the image recognition algorithm is used to identify colors, shapes, and patterns.
[1632] (Claim 3)
[1633] 10. The system of claim 1, wherein the natural language processing algorithm uses a topic modeling technique to extract keywords and topics from text data.
[1634] "Application Example 1"
[1635] (Claim 1)
[1636] means for receiving a search query entered by a user;
[1637] means for searching a database based on the search query to obtain relevant image and text data;
[1638] means for applying an image recognition algorithm to the acquired image to extract features;
[1639] A means for applying a natural language processing algorithm to the acquired text data to extract keywords and topics;
[1640] means for classifying search results into specific categories based on said features, keywords and topics;
[1641] means for transmitting the classified search results to an interface for visual review;
[1642] means for displaying the confirmed search results to the user;
[1643] A way to categorize search results and return them to the user in JSON format.
[1644] A system including:
[1645] (Claim 2)
[1646] 10. The system of claim 1, wherein the image recognition algorithm is used to identify color, shape, pattern, and to extract product features in response to a search query.
[1647] (Claim 3)
[1648] 2. The system of claim 1, wherein the natural language processing algorithm uses a topic modeling technique to extract keywords and topics from text data, and classifies products based on the extracted keywords and features.
[1649] "Example 2: Combining Emotion Engines"
[1650] (Claim 1)
[1651] means for receiving a search query entered by a user;
[1652] means for searching a database based on the search query to obtain relevant data;
[1653] means for applying an image recognition algorithm to the acquired image to extract features;
[1654] A means for applying a natural language processing algorithm to the acquired text data to extract keywords and topics;
[1655] means for classifying search results into specific categories based on said features, keywords and topics;
[1656] means for recognizing a user's emotion;
[1657] means for filtering search results based on the recognized user sentiment;
[1658] means for transmitting the classified search results to an interface for visual review;
[1659] means for displaying the confirmed search results to the user;
[1660] A system including:
[1661] (Claim 2)
[1662] The system of claim 1, wherein the image recognition algorithm is used to identify color, shape, and pattern.
[1663] (Claim 3)
[1664] 10. The system of claim 1, wherein the natural language processing algorithm uses topic modeling techniques to extract keywords and topics from text data.
[1665] "Application example 2 when combining emotion engines"
[1666] (Claim 1)
[1667] means for receiving a search query entered by a user;
[1668] means for searching a database based on the search query to obtain relevant image and text data;
[1669] means for applying an image recognition algorithm to the acquired image to extract features;
[1670] A means for applying a natural language processing algorithm to the acquired text data to extract keywords and topics;
[1671] emotion recognition means for recognizing an emotion of a user;
[1672] a means for filtering search results based on the recognized sentiment;
[1673] means for classifying search results into specific categories based on said features, keywords and topics;
[1674] means for transmitting the classified search results to an interface for visual review;
[1675] means for displaying the confirmed search results to the user;
[1676] A system including:
[1677] (Claim 2)
[1678] The system of claim 1, wherein the image recognition algorithm is used to identify color, shape, and pattern.
[1679] (Claim 3)
[1680] 10. The system of claim 1, wherein the natural language processing algorithm uses topic modeling techniques to extract keywords and topics from text data. [Explanation of symbols]
[1681] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for receiving a search query entered by a user; means for searching a database based on the search query to obtain relevant image and text data; means for applying an image recognition algorithm to the acquired image to extract features; A means for applying a natural language processing algorithm to the acquired text data to extract keywords and topics; means for classifying search results into specific categories based on said features, keywords and topics; means for transmitting the classified search results to an interface for visual review; means for displaying the confirmed search results to the user; A system including:
2. The system of claim 1 , wherein the image recognition algorithm is used to identify color, shape, and pattern.
3. The system of claim 1 , wherein the natural language processing algorithm uses topic modeling techniques to extract keywords and topics from text data.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A