system

The system addresses inefficiencies in image search by using AI to recognize and tag objects and text in images, enabling rapid retrieval of relevant data.

JP2026070205APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Conventional image search systems struggle to efficiently search for specific objects from user-captured photos due to reliance on keyword-based searches, failing to utilize visual features and text information effectively, leading to inefficiencies in industries like construction where past examples and materials cannot be quickly searched.

Method used

A system that uses AI to recognize objects and extract text from images, generating detailed tags, and stores them in a database for efficient retrieval, allowing users to search for relevant images based on keywords.

Benefits of technology

Enables rapid and efficient retrieval of image data by automatically identifying and tagging objects and text, improving work efficiency by quickly finding necessary photos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070205000001_ABST
    Figure 2026070205000001_ABST
Patent Text Reader

Abstract

Provide a system. 【Solution means】 Means for receiving image data captured by a user, Means for using artificial intelligence to identify an object in the image data, Means for generating a tag corresponding to the identified object, Means for extracting character information in the image data using OCR technology and generating additional tags based on this character information, Means for storing the generated tags in a database, Means for receiving a text-based search query and searching for related image data based on the stored tags, Means for presenting search results to the user, A system including the above.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of this disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In conventional image search systems, it is difficult for users to efficiently search for specific objects from photos taken by themselves. Many systems are limited to keyword-based searches and cannot fully utilize visual features and text information in images. For this reason, especially in the construction industry, there is a problem that past construction examples and related materials cannot be quickly searched, resulting in a decline in work efficiency.

Means for Solving the Problems

[0005] This invention provides a system that receives image data captured by a user and automatically recognizes objects within the image using AI technology. This system generates tags related to the objects and further extracts textual information from the image using OCR technology to generate additional tags, thereby producing more detailed and searchable data. Furthermore, by storing these tags in a database and quickly searching for relevant images in response to the user's search query, and presenting the results, the system achieves efficient image retrieval.

[0006] A "user" is an individual or organization that operates the system and uses the interface to upload or search for image data.

[0007] "Image data" refers to digital image files that are taken by users and uploaded to the system.

[0008] An "object" refers to a visual element, such as a specific object or equipment, included in image data.

[0009] "Artificial intelligence" is a technology that includes machine learning algorithms used to automatically identify objects in images.

[0010] A "tag" is a keyword or label associated with image data, used to describe an object or feature.

[0011] "OCR technology" is a technology that mechanically recognizes textual information contained in image data and converts its content into digital text data.

[0012] A "database" is an information system used to store and manage generated tags and associated image data.

[0013] A "search query" is a text-based request that a user enters to search for specific information.

[0014] "Search results" refer to a list or collection of relevant image data extracted and presented by the system based on the search query submitted by the user. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc. <00,00109> In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention relates to an image search system that efficiently handles image data captured by users using their personal devices. This system operates primarily through a server, and its embodiments are described below.

[0037] When a user uploads image data taken with their device to the system, the image data is sent to the server. Upon receiving this image data, the server immediately begins analysis. First, the server applies artificial intelligence technology to the image data to identify objects within the image. This is done by an artificial intelligence model that has been pre-trained to recognize specific objects. For example, since photographs of construction sites often contain objects such as cranes and rebar, the system is configured to identify these.

[0038] Next, the server uses OCR technology to extract text information present in the image. This extraction process optically recognizes characters in the image and saves them as digital text. For example, it can extract information from signs at construction sites or text within safety signs.

[0039] Based on the identified objects and the text extracted by OCR, the server generates tags associated with the photographs. These tags serve as keywords representing the object names or summaries of the extracted text. The generated tags are stored in a database, linked to the photograph data, to facilitate searching.

[0040] When a user searches for a specific object or element, they send a search query from their device to the server. The server searches its database based on the keywords specified in this search query and quickly extracts image data with relevant tags. The resulting image data is sent to the user's device and displayed on the screen. This allows the user to quickly find the information they are looking for.

[0041] As a concrete example, when a user in the construction industry searches for photos of past construction sites using the keyword "crane," the server automatically extracts photos tagged with "crane" from the database and quickly sends them to the user's terminal. This process allows users to find the necessary photos in a short amount of time, dramatically improving work efficiency.

[0042] The following describes the processing flow.

[0043] Step 1:

[0044] The user selects image data to search using their device and uploads it to the server. The device then sends the image file to the server as an HTTP request.

[0045] Step 2:

[0046] The server prepares to analyze the received image data. The server loads an artificial intelligence model and processes the image data using a computer vision algorithm to identify objects within the image.

[0047] Step 3:

[0048] The server automatically generates tags based on the features of objects extracted from images. These tags represent the type and properties of the objects and are used for subsequent searches.

[0049] Step 4:

[0050] The server uses OCR technology to analyze the text information present in the image and extracts it as digital text. This resulting text is also used as a tag.

[0051] Step 5:

[0052] The server associates the generated tags with image data and stores them in a database. This enables efficient searching.

[0053] Step 6:

[0054] To search for a specific image, the user enters a search query (keyword) on their device and sends it to the server.

[0055] Step 7:

[0056] The server searches the database based on the user's search query and extracts image data tagged with the query.

[0057] Step 8:

[0058] The server generates search results and sends them to the user's device. The user can then view the search results on their device and browse images that interest them.

[0059] (Example 1)

[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0061] Conventional image search systems struggle to efficiently find specific information from large amounts of image data. For example, manual tagging is required, placing a significant burden on human resources. Furthermore, it has been difficult to automatically extract and appropriately classify information based on the objects and text information contained within the original images. This results in the challenge of users being unable to quickly search for the image data they need.

[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0063] In this invention, the server includes means for receiving image data captured from a user's information processing device, means for using machine intelligence to identify objects in the image data, and means for generating relevant tags based on the identified objects and extracted text. This enables the user to efficiently and quickly search for image data related to specific keywords.

[0064] A "user information processing device" is a device, such as a computer or mobile information terminal, used by a user that enables the capture and uploading of image data to a system.

[0065] "Image data" refers to digital visual information captured by a user using an information processing device, and is stored as data in the form of still images or videos.

[0066] "Machine intelligence" refers to artificial intelligence technology used to automatically identify objects within image data, and is implemented through models trained using learning algorithms.

[0067] An "identified object" refers to a specific item or element recognized within image data using machine intelligence, and each object is identified based on a pre-defined classification.

[0068] "Optical character recognition technology" is a technology that mechanically extracts character information contained within image data as digital text, and is a process executed by an optical processor.

[0069] A "tag" refers to keywords or metadata generated based on identified objects and extracted text, and is information added to facilitate searching image data.

[0070] A "storage device" is a device or system that functions as a database or storage system and is used to hold generated tags and associated image data.

[0071] A "text-based information request" is a search query in string format entered by a user, used to communicate information about specific keywords to the system.

[0072] "Searching" is the process of searching for data within a storage device based on a user's request and finding relevant image data.

[0073] This invention is an image search system that efficiently manages image data captured by users using an information processing device and allows for rapid retrieval of specific objects or text information. It is primarily server-based and functions as follows:

[0074] The user first takes an image using an information processing device. This device is a common device such as a computer or mobile terminal with a camera function. This image data is uploaded to a server via an internet connection using a secure protocol (e.g., HTTPS).

[0075] The server temporarily stores the received image data in file storage. Next, it utilizes a generative AI model to identify objects within the image data. This AI model is trained using machine learning algorithms and can utilize deep learning techniques. For example, in a photograph of a construction site, it can be configured to recognize specific objects such as cranes and reinforcing bars.

[0076] The server also uses optical character recognition (OCR) technology to extract text information from images. The OCR process employs common optical character recognition software (e.g., Tesseract) to detect characters from image data and save them as digital text. For example, it can retrieve the content written on signs and notices.

[0077] Based on the identified objects and extracted text information, the server generates relevant tags and stores them in a database. This allows users to quickly retrieve relevant information when searching for images using specific keywords. SQL or NoSQL databases can be used for this purpose.

[0078] As a concrete example, if a user in the construction industry searches for images of past construction sites using the keyword "crane," the server extracts images tagged with "crane" and provides them to the user quickly. This process allows users to find the necessary information in a short amount of time, significantly improving work efficiency.

[0079] An example prompt might be a request like, "Search for photos containing cranes in construction site images and generate related text information as tags." This system could be a powerful tool, especially in industries that need to handle large amounts of image data.

[0080] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0081] Step 1:

[0082] The user captures image data with an information processing device. The device is a camera-equipped device that acquires the captured image data. The input is the captured image, and the output is that image data. This step involves the user capturing a specific scene or object and saving it in a digital format.

[0083] Step 2:

[0084] The device uploads captured image data to the server. During this process, the device transmits the image data using a secure communication protocol. The input is the image data provided by the user's device, and the output is the data transferred to the server. This step involves data transmission via an internet connection.

[0085] Step 3:

[0086] The server saves the received image data to file storage. The input is the uploaded image data, and the output is the image data stored in storage. This step involves the operation of saving the data to the storage device hosted by the server.

[0087] Step 4:

[0088] The server generates image data and analyzes it using an AI model to identify objects within the image. The input is stored image data, and the output is a list of identified objects. At this stage, the AI ​​model performs the process of feature extraction and classification.

[0089] Step 5:

[0090] The server extracts text information from an image using optical character recognition (OCR) technology. The input is the image data after identification is complete, and the output is the extracted text information. In this step, the text is digitized using the OCR process.

[0091] Step 6:

[0092] The server generates relevant tags based on identified objects and extracted text. The input is a list of objects and text information, and the output is the generated tags. This process also generates object names and keywords.

[0093] Step 7:

[0094] The server saves the generated tags to the database. The input is the newly generated tags, and the output is the updated database entry. This step involves registering and maintaining information in the database.

[0095] Step 8:

[0096] A user sends a search query from their device to the server using specific keywords. The input is the search query entered by the user, and the output is the search request forwarded to the server. In this process, query input and submission occur via the user interface.

[0097] Step 9:

[0098] The server extracts image data from the database that have tags matching the search query. The input is the search query and the database, and the output is image data as the search results. At this stage, a database search algorithm is applied.

[0099] Step 10:

[0100] The server sends the search results to the user's terminal. The input is image data obtained through the search, and the output is the information displayed on the user's terminal. This step involves data transfer over the internet and display on the user's screen.

[0101] (Application Example 1)

[0102] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0103] Managing goods and machinery within a factory generates a large amount of image data, and there is a need to efficiently analyze and search this data. Conventional systems have the problem of requiring significant time and effort to manually classify and search image data. Furthermore, the lack of readily available information on-site makes rapid decision-making and inventory management difficult.

[0104] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0105] In this invention, the server includes means for receiving image data captured by a user, means for using artificial intelligence to identify objects in the image data, means for generating tags corresponding to the identified objects, means for extracting text information from the image data using OCR technology and generating additional tags based on this text information, means for storing the generated tags in a database, means for receiving text-based search queries and searching for related image data based on the stored tags, means for presenting the search results to the user, means for utilizing a remote information display device or a mobile automated machine as a shooting terminal, and means for providing a system for efficiently managing object and machine information based on the generated tags. This enables real-time information acquisition within the factory, efficient inventory management, and rapid decision-making.

[0106] A "user" is an individual or group that operates an image capture device or information display device to acquire or view image data.

[0107] "Image data" refers to data that digitally represents the visual information of objects or scenes captured by a user.

[0108] Artificial intelligence is a collection of algorithms and models that computer systems use to enable learning and recognition.

[0109] A "tag" is a keyword or label used to identify and summarize objects and text information contained in image data.

[0110] "OCR technology" is a technology that optically recognizes text within an image and converts it into digital text.

[0111] A "database" is a digital information management system that systematically stores generated tags and other related information, enabling information retrieval.

[0112] A "search query" is a text-based request that a user enters into a system to retrieve specific information.

[0113] A "remote information display device" is a visual device that allows users to acquire information remotely by wearing or carrying it.

[0114] A "mobile automated machine" is a machine or device that operates autonomously in an environment such as a factory and is capable of taking images and collecting data.

[0115] "Inventory management" refers to activities aimed at efficiently understanding and managing the location, quantity, and condition of items within an environment such as a factory or warehouse.

[0116] The system for implementing this invention efficiently manages image data captured by the user and quickly searches for and retrieves necessary information. The user uses a remote information display device or a mobile automated machine as the imaging device. For example, a worker wearing smart glasses or a self-propelled robot patrols the factory and takes pictures of products and machinery.

[0117] The server receives and analyzes captured image data in real time. Specifically, the server uses artificial intelligence technology to identify objects within the image data and uses OCR technology to optically recognize text information within the image and extract it as digital text. Based on the identified objects and extracted text information, the server generates relevant tags and stores them in a database.

[0118] This system utilizes generative AI models based on TENSORFLOW® and PyTorch to perform image analysis. It also performs OCR processing using the Google® Cloud Vision API. A database management system such as MySQL® is used for the database.

[0119] When a user searches for specific information, this is achieved by sending a text-based search query from the terminal to the server. For example, by entering a prompt such as "I want to check the location of the crane," the server quickly searches for relevant image data and displays it on the terminal. This allows users to efficiently manage items and obtain machine information within the factory.

[0120] Specific examples of prompt statements include "Show me the inventory image of part A" and "I want to check the maintenance information for machine 123." Through these prompt statements, users can instantly obtain the information they need on-site and improve the efficiency of their work.

[0121] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0122] Step 1:

[0123] The user takes images within the factory using a remote information display device or mobile automated machine. The input is visual information of the environment, and the output is captured image data. This image data is immediately uploaded to a cloud server.

[0124] Step 2:

[0125] The server retrieves the received image data and applies a generative AI model using TensorFlow or PyTorch. The input is the captured image data, and the output is a result in which objects are identified. Specifically, the server inputs the image into the model, analyzes the features of each pixel, and identifies the type of object.

[0126] Step 3:

[0127] The server performs OCR processing using the Google Cloud Vision API. The input is image data, and the output is extracted text information. Specifically, the server identifies text regions within the image and converts their content into digital text.

[0128] Step 4:

[0129] The server generates relevant tags based on the obtained object identification results and text information, and stores them in the database. The input is the object identification results and text information, and the output is a database entry containing the generated tags. Specifically, the server extracts keywords useful for searching based on the identified features.

[0130] Step 5:

[0131] When a user wants to retrieve specific information, they send a search query from their terminal to the server as a prompt. The input is a text-based query, and the output is a request for the corresponding image data. Specifically, the user might enter a command such as "Show me the inventory image of part A" into their terminal.

[0132] Step 6:

[0133] The server searches the database based on the received search query and extracts relevant image data. The input consists of a prompt and database tag information, and the output is the corresponding image data. Specifically, the server compares the query based on the stored tags and quickly extracts matching data.

[0134] Step 7:

[0135] The server sends the search results to the user's terminal, and the terminal displays the images. The input is extracted image data, and the output is the image that the user displays. Specifically, the server transfers the image data to the user's terminal, and the terminal displays that data on its screen.

[0136] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0137] This invention is an innovative image analysis system that incorporates an emotion engine to recognize user emotions, in addition to a conventional image search system. The system begins with the user uploading image data captured using a terminal to a server.

[0138] The server analyzes the received image data. First, it uses artificial intelligence to automatically identify objects within the image. This generates tags for each unique object in the image, which are then stored in a database. Additionally, OCR technology is used to extract text information contained in the image, and further tags are generated based on this information.

[0139] Furthermore, a key feature of this system is its use of an emotion engine to analyze the user's emotional state from image data. The server analyzes facial expressions and various biometric information in the images to identify the user's emotions. This emotional information is also generated as a tag and added to the relevant data in the database. This enables searches that consider not only physical information but also emotional context.

[0140] When a user searches for a specific image or related information, they send a text-based search query from their device to the server. The server quickly extracts image data tagged with the query and presents the search results to the user's device. This system can also provide recommendations based on information obtained through sentiment analysis, for example, by identifying emotional responses to specific statuses or work environments on a construction site.

[0141] For example, if a user searches for past event photos based on the emotion of "happiness," the emotion engine will extract photos tagged with the emotion "happiness" from the database. Users can find the most suitable images based on their emotional state, going beyond simply finding physical objects. This makes it easier to obtain information from a new perspective that would have been difficult to find with conventional search systems.

[0142] The following describes the processing flow.

[0143] Step 1:

[0144] The user uses their device to select image data containing emotions and sends an upload request to the server. The device then transfers the selected image files to the server via the HTTP protocol.

[0145] Step 2:

[0146] The server stores the received image data in a buffer for analysis. Then, it runs artificial intelligence to automatically identify objects in the image. This involves using an object recognition model to detect identifiable objects in the image and extract the associated object names.

[0147] Step 3:

[0148] The server generates relevant tags based on the identified objects. These generated tags indicate the type or category of the object and serve to improve search efficiency within the database.

[0149] Step 4:

[0150] The server uses OCR technology to analyze text information within image data. It detects textual information contained in the scene and converts its content into digital text. This text information is also generated as tags and associated with other elements.

[0151] Step 5:

[0152] The server recognizes the user's emotions by analyzing facial expressions and biometric information within images using an emotion engine. Based on the recognized emotion information, it generates additional emotion tags and associates them with the image data.

[0153] Step 6:

[0154] The generated object tags, text tags, and sentiment tags are stored in a database, and an index is created to allow for efficient searching of matching images.

[0155] Step 7:

[0156] To search for images related to a specific object or emotion, the user enters a text-based search query on their device and sends it to the server.

[0157] Step 8:

[0158] The server extracts image data from the database that matches the appropriate tags based on the received search query. If necessary, it refines the results by considering sentiment tags.

[0159] Step 9:

[0160] The server generates search results and sends them to the user's device. The user can view the search results on their device, select images of interest, and view more details.

[0161] (Example 2)

[0162] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0163] Modern image search systems extract only physical information from user-submitted image data, failing to consider emotions or context. This makes it difficult to perform image searches based on the emotional context desired by the user. Specifically, the inability to quickly and accurately search for images based on a particular emotional state or mood limits the user's search experience.

[0164] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0165] In this invention, the server includes means for receiving image data captured by the user, means for using artificial intelligence to identify objects in the image data, means for extracting textual information from the image data using OCR technology, and means for identifying the user's emotional state from the image data using an emotion analysis engine and generating emotion-based tags. This enables the user to perform image searches based on emotional context in addition to physical information.

[0166] "Means for receiving image data captured by a user" refers to a function that allows a server to receive image data captured by a user with their device via a network.

[0167] "Means of using artificial intelligence to identify objects in image data" refers to a process that utilizes image analysis algorithms to automatically detect and identify objects within an image.

[0168] "Means for extracting text information from image data using OCR technology" refers to a technology that analyzes text contained in an image using optical character recognition technology and extracts it as digital text.

[0169] "Methods for identifying a user's emotional state from image data using an emotion analysis engine" refers to algorithms that analyze facial expressions and biometric information in images to identify the user's emotions.

[0170] "Means for generating emotion-based tags" refers to a function that generates relevant tags based on user emotion information obtained through emotion analysis and stores them in a database.

[0171] This invention is a system that enhances image search by processing image data captured by a user using a terminal and performing object recognition and sentiment analysis. Specific embodiments for carrying out the invention are described below.

[0172] Users take pictures using devices such as smartphones or personal computers and upload the image data to the system. The uploaded data is first received by the server. On the server, the image data is preprocessed using the image processing library "OpenCV" and converted to an appropriate format for analysis.

[0173] Next, the server utilizes deep learning frameworks such as "TensorFlow" and "PyTorch" to identify objects in the image using models (e.g., VGG and ResNet). A corresponding tag is generated for each identified object, and these tags are stored in a database.

[0174] Furthermore, the server uses OCR technology such as "Tesseract" to extract text information from the image and generates additional tags based on the text. This process efficiently analyzes the characters in the image, allowing the necessary information to be obtained.

[0175] In particular, in this invention, the server is equipped with an emotion analysis engine that analyzes facial expressions and biometric information using tools such as "Haar Cascade" and "Keras." Through this analysis, the user's emotional state is identified, and tags indicating that emotion are also stored in a database.

[0176] For example, if a user wants to search for photos of events related to happiness, the system will quickly extract and present images tagged with the emotion "happiness." This allows users to perform advanced image searches based on emotional context.

[0177] Examples of prompts for a generative AI model include: "Create a description of the algorithm for an image search system that takes user emotions into consideration. Include the specific process of emotion analysis performed by the emotion engine."

[0178] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0179] Step 1:

[0180] The user takes image data using their device and uploads it to the system. This involves clicking an "upload" button via an application on the device, sending the image data to the server over the internet. The input is the image data captured by the device, and the output is the image data sent to the server.

[0181] Step 2:

[0182] The server receives the image data and first performs image preprocessing using the "OpenCV" library. This preprocessing includes standardizing the image size and format. The input is the received image data, and the output is the preprocessed image data.

[0183] Step 3:

[0184] The server then applies an object recognition model (e.g., VGG or ResNet) using a deep learning framework such as TensorFlow or PyTorch. This identifies objects in the image and generates corresponding tags. The input is preprocessed image data, and the output is the object recognition result and the generated tags.

[0185] Step 4:

[0186] The server uses OCR technology such as "Tesseract" to extract text information from image data. This process obtains text from the image as a digital string and also generates tags. The input is image data, and the output is the extracted text information and additional tags.

[0187] Step 5:

[0188] The server uses an emotion analysis engine to analyze emotional information from images. Specifically, it uses "Haar Cascade" and "Keras" to analyze facial expressions and biometric information to identify the user's emotions. The input is image data, and the output is an emotion status and emotion tag.

[0189] Step 6:

[0190] The server saves all generated tags to the database. Here, object tags, text tags, and sentiment tags are all saved together in association. The input is the various tags that have been generated, and the output is the status of the data being saved to the database.

[0191] Step 7:

[0192] The user sends a text-based search query from their device to the server. Specifically, they enter keywords into the application's search bar and click the "Search" button. The input is the query information entered by the user, and the output is the query sent to the server.

[0193] Step 8:

[0194] The server extracts relevant image data from the database based on the received search query. It searches for highly relevant images from stored tags and formats the results. The input is the received search query, and the output is the search results.

[0195] Step 9:

[0196] The server displays the extracted search results on the user's device. Here, the results are formatted to be easily viewed by the user as an image list. The input is the search results formatted by the server, and the output is the image list displayed on the device screen.

[0197] (Application Example 2)

[0198] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0199] There is a challenge in providing appropriate information and recommendations based on images taken by users, taking into account their emotions and the context at the time. Conventional image search systems primarily rely on information retrieval based on physical characteristics and are unable to provide information that takes into account the user's emotions or psychological state.

[0200] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0201] In this invention, the server includes means for receiving image data captured by the user, means for identifying objects in the image data and generating tags, and means for generating tags based on emotional information using an emotion engine for analyzing the user's emotional state from the image data. This makes it possible to provide and recommend information that takes the user's emotional state into consideration.

[0202] "Means for receiving image data captured by a user" refers to a processing device that uses communication or an interface to acquire image information captured by a user from a terminal and incorporate it into the system.

[0203] "Methods using artificial intelligence" refer to methods that use computer-based algorithms to recognize and identify objects from input data, and utilize machine learning techniques.

[0204] "Means for generating tags corresponding to identified objects" refers to a device or software that performs a process of assigning textual information or identification information to objects recognized within image data.

[0205] "Means of using OCR technology" refers to a method or apparatus for optically recognizing text information within an image and converting it into digital data.

[0206] "Means of using an emotion engine" refers to an algorithm or program that analyzes the user's facial expressions and biometric information contained in an image to determine the user's emotional state.

[0207] "Means for storing generated tags in a database" refers to a database or storage device that is permanently maintained within the system for managing identified tags and sentiment information.

[0208] "A means for receiving text-based search queries and searching for related image data" refers to a search processing device that receives textual inquiries from users and searches for and extracts related image information based on those inquiries.

[0209] "Means of providing recommended information to users based on search results" refers to notifications or display devices that analyze data obtained through searches and present the user with the most optimal or relevant information.

[0210] "Means for presenting search results to the user" refers to an interface or output device for displaying or transmitting the results obtained by the system through a search to the user's terminal.

[0211] The system for implementing this invention begins with a user sending image data captured using a smart device to a server. The server analyzes the received images using an artificial intelligence module. This AI module is a model trained using machine learning platforms such as TensorFlow and PyTorch, and has the ability to identify objects in the image. Furthermore, it uses a Tesseract engine employing OCR (optical character recognition) technology to extract text information from the image and generate tags.

[0212] Subsequently, the image data is processed by an emotion engine, which uses Microsoft® Azure® facial recognition APIs and other tools to analyze the user's facial expressions and biometric indicators, thereby identifying their emotional state. This process generates emotion tags such as "happiness" and "surprise," which are then stored in a database.

[0213] When a user searches for specific information, the device sends a text-based search query to the server. The server quickly searches for relevant image data based on tags in its database and provides information related to a specific emotional state. For example, if the user's emotion is identified as "happy," advertisements related to entertainment and leisure will be displayed.

[0214] For example, if a user sends a photo of themselves happily with a friend, it could be tagged with "happiness," and then they could be offered movie ticket promotions or event information.

[0215] An example of a prompt message might be: "The user took a photo of themselves smiling with a friend. Sentiment analysis determined that they were feeling 'happy.' Please create an advertisement that matches this emotion."

[0216] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0217] Step 1:

[0218] The user takes an image with a smart device and sends the image data from the device to the server. The input is the image data taken by the user, and the output is the image data received by the server. In this process, the image is acquired using the smart device's camera API and the data is uploaded to the server via the HTTP protocol over an internet connection.

[0219] Step 2:

[0220] The server analyzes received image data using a machine learning model to identify objects within the image. The input is the received image data, and the output is a list of tags corresponding to the identified objects. Specifically, it uses generative AI models such as TensorFlow or PyTorch to analyze features in the image and generate tags.

[0221] Step 3:

[0222] The server extracts text information from the image using an OCR engine. The input is still the received image data, and the output is additional tag information based on the extracted text. An OCR engine such as Tesseract is used here to extract string data from the image and perform analysis.

[0223] Step 4:

[0224] The server uses an emotion engine to analyze facial expressions from images and identify the user's emotional state. The input is image data containing faces, and the output is tag information based on the user's emotions. For emotion analysis, the server uses Microsoft Azure's facial recognition API, among others, to infer emotions from the identified facial expression data.

[0225] Step 5:

[0226] The server saves all generated tags to a database. The input is the generated tag information, and the output is the registration status of the tags in the program's database. The tag information is recorded in a database system such as MySQL.

[0227] Step 6:

[0228] The user sends a text-based search query from their terminal to the server to find specific information. The input is the user's query text, and the output is that text data received by the server. The user enters the query through the application's search interface.

[0229] Step 7:

[0230] The server searches the database for appropriate image data based on the tags associated with the received query. The input is the search query and tag information from the database, and the output is a set of related image data. Here, an SQL query is used to filter the data for matching tags.

[0231] Step 8:

[0232] The server generates and presents recommended information to the user based on the search results. The input is image data and related information as search results, and the output is recommended advertisements and information displayed on the user's device. Specifically, it sets matching advertisement content based on the search results and displays the results on the user's interface.

[0233] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0234] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0235] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0236] [Second Embodiment]

[0237] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0238] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0239] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0240] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0241] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0242] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0243] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0244] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0245] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0246] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0247] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0248] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0249] This invention relates to an image search system that efficiently handles image data captured by users using their personal devices. This system operates primarily through a server, and its embodiments are described below.

[0250] When a user uploads image data taken with their device to the system, the image data is sent to the server. Upon receiving this image data, the server immediately begins analysis. First, the server applies artificial intelligence technology to the image data to identify objects within the image. This is done by an artificial intelligence model that has been pre-trained to recognize specific objects. For example, since photographs of construction sites often contain objects such as cranes and rebar, the system is configured to identify these.

[0251] Next, the server uses OCR technology to extract text information present in the image. This extraction process optically recognizes characters in the image and saves them as digital text. For example, it can extract information from signs at construction sites or text within safety signs.

[0252] Based on the identified objects and the text extracted by OCR, the server generates tags associated with the photographs. These tags serve as keywords representing the object names or summaries of the extracted text. The generated tags are stored in a database, linked to the photograph data, to facilitate searching.

[0253] When a user searches for a specific object or element, they send a search query from their device to the server. The server searches its database based on the keywords specified in this search query and quickly extracts image data with relevant tags. The resulting image data is sent to the user's device and displayed on the screen. This allows the user to quickly find the information they are looking for.

[0254] For example, if a user in the construction industry searches for photos of past construction sites using the keyword "crane," the server automatically extracts photos tagged with "crane" from the database and quickly sends them to the user's terminal. This process allows users to find the necessary photos in a short amount of time, dramatically improving work efficiency.

[0255] The following describes the processing flow.

[0256] Step 1:

[0257] The user selects image data to search using their device and uploads it to the server. The device then sends the image file to the server as an HTTP request.

[0258] Step 2:

[0259] The server prepares to analyze the received image data. The server loads an artificial intelligence model and processes the image data using a computer vision algorithm to identify objects within the image.

[0260] Step 3:

[0261] The server automatically generates tags based on the features of objects extracted from images. These tags represent the type and properties of the objects and are used for subsequent searches.

[0262] Step 4:

[0263] The server uses OCR technology to analyze the text information present in the image and extracts it as digital text. This resulting text is also used as a tag.

[0264] Step 5:

[0265] The server associates the generated tags with image data and stores them in a database. This enables efficient searching.

[0266] Step 6:

[0267] To search for a specific image, the user enters a search query (keyword) on their device and sends it to the server.

[0268] Step 7:

[0269] The server searches the database based on the user's search query and extracts image data tagged with the query.

[0270] Step 8:

[0271] The server generates search results and sends them to the user's device. The user can then view the search results on their device and browse images that interest them.

[0272] (Example 1)

[0273] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0274] Conventional image search systems struggle to efficiently find specific information from large amounts of image data. For example, manual tagging is required, placing a significant burden on human resources. Furthermore, it has been difficult to automatically extract and appropriately classify information based on the objects and text information contained within the original images. This results in the challenge of users being unable to quickly search for the image data they need.

[0275] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0276] In this invention, the server includes means for receiving image data captured from a user's information processing device, means for using machine intelligence to identify objects in the image data, and means for generating relevant tags based on the identified objects and extracted text. This enables the user to efficiently and quickly search for image data related to specific keywords.

[0277] A "user information processing device" is a device, such as a computer or mobile information terminal, used by a user that enables the capture and uploading of image data to a system.

[0278] "Image data" refers to digital visual information captured by a user using an information processing device, and is data stored as still images or videos.

[0279] "Machine intelligence" refers to artificial intelligence technology for automatically identifying objects within image data, and is a technology realized by a model trained using a learning algorithm.

[0280] "Identified object" refers to a specific article or element recognized within image data using machine intelligence, and each object is identified based on a preset classification.

[0281] "Optical character recognition technology" is a technology for mechanically extracting character information contained within image data as digital text, and is a process executed by an optical processor.

[0282] "Tag" refers to keywords or metadata generated based on identified objects and extracted text, and is information assigned to facilitate the search of image data.

[0283] "Memory device" is a device or system that functions as a database or storage system, and is used to hold generated tags and related image data.

[0284] "Text-based information request" is a search query in string format input by a user, and is used to transmit information regarding specific keywords to the system.

[0285] "Search" is a process of searching data within a memory device based on a user's request and finding related image data.

[0286] This invention is an image search system that efficiently manages image data captured by a user using an information processing device and quickly searches for specific objects and text information. It is mainly configured around a server and functions as follows.

[0287] First, the user uses an information processing device to take a picture. This device is a general device such as a computer or mobile terminal with a camera function. This image data is uploaded to the server via an Internet connection using a secure protocol (e.g., HTTPS).

[0288] The server temporarily stores the received image data in a file storage. Next, it utilizes a generative AI model to identify the objects in the image data. This AI model is trained by a machine learning algorithm and can use deep learning techniques. For example, in a photo of a construction site, it is set to recognize specific objects such as cranes and steel bars.

[0289] In addition, the server uses optical character recognition technology to extract text information in the image. For the OCR process, general optical character recognition software (e.g., Tesseract) is used to detect characters from the image data and save them as digital text. As an example, the content written on a signboard or a label can be obtained.

[0290] Based on the identified objects and the extracted text information, the server generates relevant tags and saves them in a database. This enables the server to quickly retrieve the corresponding information when the user searches for images with a specific keyword. SQL databases or NoSQL databases are used for the database.

[0291] As a specific example, when a user in the construction industry searches for images of past construction sites with the keyword "crane", the server extracts the images with the tag "crane" and quickly provides them to the user. Through this process, the user can discover the necessary information in a short time and significantly improve the efficiency of their work.

[0292] An example prompt might be a request like, "Search for photos containing cranes in construction site images and generate related text information as tags." This system could be a powerful tool, especially in industries that need to handle large amounts of image data.

[0293] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0294] Step 1:

[0295] The user captures image data with an information processing device. The device is a camera-equipped device that acquires the captured image data. The input is the captured image, and the output is that image data. This step involves the user capturing a specific scene or object and saving it in a digital format.

[0296] Step 2:

[0297] The device uploads captured image data to the server. During this process, the device transmits the image data using a secure communication protocol. The input is the image data provided by the user's device, and the output is the data transferred to the server. This step involves data transmission via an internet connection.

[0298] Step 3:

[0299] The server saves the received image data to file storage. The input is the uploaded image data, and the output is the image data stored in storage. This step involves the operation of saving the data to the storage device hosted by the server.

[0300] Step 4:

[0301] The server analyzes the image data using a generative AI model to identify the objects within the image. The input is the saved image data, and the output is a list of identified objects. At this stage, the processes of feature extraction and classification by the AI model are executed.

[0302] Step 5:

[0303] The server extracts the character information within the image using optical character recognition technology. The input is the image data for which identification has been completed, and the output is the extracted text information. At this step, the digitization of the text by the OCR process is performed.

[0304] Step 6:

[0305] The server generates relevant tags based on the identified objects and the extracted text. The input is the list of objects and the text information, and the output is the generated tags. In this process, the generation of object names and keywords is performed.

[0306] Step 7:

[0307] The server saves the generated tags to the database. The input is the newly generated tags, and the output is the updated database entry. At this step, the registration and retention of information in the database are performed.

[0308] Step 8:

[0309] The user sends a search query from the terminal to the server with specific keywords. The input is the search query entered by the user, and the output is the search request transferred to the server. In this process, the input and transmission of the query are performed via the user interface.

[0310] Step 9:

[0311] The server extracts image data from the database that have tags matching the search query. The input is the search query and the database, and the output is image data as the search results. At this stage, a database search algorithm is applied.

[0312] Step 10:

[0313] The server sends the search results to the user's terminal. The input is image data obtained through the search, and the output is the information displayed on the user's terminal. This step involves data transfer over the internet and display on the user's screen.

[0314] (Application Example 1)

[0315] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0316] Managing goods and machinery within a factory generates a large amount of image data, and there is a need to efficiently analyze and search this data. Conventional systems have the problem of requiring significant time and effort to manually classify and search image data. Furthermore, the lack of readily available information on-site makes rapid decision-making and inventory management difficult.

[0317] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0318] In this invention, the server includes means for receiving image data captured by a user, means for using artificial intelligence to identify objects in the image data, means for generating tags corresponding to the identified objects, means for extracting text information from the image data using OCR technology and generating additional tags based on this text information, means for storing the generated tags in a database, means for receiving text-based search queries and searching for related image data based on the stored tags, means for presenting the search results to the user, means for utilizing a remote information display device or a mobile automated machine as a shooting terminal, and means for providing a system for efficiently managing object and machine information based on the generated tags. This enables real-time information acquisition within the factory, efficient inventory management, and rapid decision-making.

[0319] A "user" is an individual or group that operates an image capture device or information display device to acquire or view image data.

[0320] "Image data" refers to data that digitally represents the visual information of objects or scenes captured by a user.

[0321] Artificial intelligence is a collection of algorithms and models that computer systems use to enable learning and recognition.

[0322] A "tag" is a keyword or label used to identify and summarize objects and text information contained in image data.

[0323] "OCR technology" is a technology that optically recognizes text within an image and converts it into digital text.

[0324] A "database" is a digital information management system that systematically stores generated tags and other related information, enabling information retrieval.

[0325] A "search query" is a text-based request that a user enters into a system to retrieve specific information.

[0326] A "remote information display device" is a visual device that allows users to acquire information remotely by wearing or carrying it.

[0327] A "mobile automated machine" is a machine or device that operates autonomously in an environment such as a factory and is capable of taking images and collecting data.

[0328] "Inventory management" refers to activities aimed at efficiently understanding and managing the location, quantity, and condition of items within an environment such as a factory or warehouse.

[0329] The system for implementing this invention efficiently manages image data captured by the user and quickly searches for and retrieves necessary information. The user uses a remote information display device or a mobile automated machine as the imaging device. For example, a worker wearing smart glasses or a self-propelled robot patrols the factory and takes pictures of products and machinery.

[0330] The server receives and analyzes captured image data in real time. Specifically, the server uses artificial intelligence technology to identify objects within the image data and uses OCR technology to optically recognize text information within the image and extract it as digital text. Based on the identified objects and extracted text information, the server generates relevant tags and stores them in a database.

[0331] This system utilizes generative AI models based on TensorFlow and PyTorch to perform image analysis. It also uses the Google Cloud Vision API for OCR processing. A database management system such as MySQL is used for the database.

[0332] When a user searches for specific information, this is achieved by sending a text-based search query from the terminal to the server. For example, by entering a prompt such as "I want to check the location of the crane," the server quickly searches for relevant image data and displays it on the terminal. This allows users to efficiently manage items and obtain machine information within the factory.

[0333] Specific examples of prompt statements include "Show me the inventory image of part A" and "I want to check the maintenance information for machine 123." Through these prompt statements, users can instantly obtain the information they need on-site and improve the efficiency of their work.

[0334] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0335] Step 1:

[0336] The user takes images within the factory using a remote information display device or mobile automated machine. The input is visual information of the environment, and the output is captured image data. This image data is immediately uploaded to a cloud server.

[0337] Step 2:

[0338] The server retrieves the received image data and applies a generative AI model using TensorFlow or PyTorch. The input is the captured image data, and the output is a result in which objects are identified. Specifically, the server inputs the image into the model, analyzes the features of each pixel, and identifies the type of object.

[0339] Step 3:

[0340] The server performs OCR processing using the Google Cloud Vision API. The input is image data, and the output is extracted text information. Specifically, the server identifies text regions within the image and converts their content into digital text.

[0341] Step 4:

[0342] The server generates relevant tags based on the obtained object identification results and text information, and stores them in the database. The input is the object identification results and text information, and the output is a database entry containing the generated tags. Specifically, the server extracts keywords useful for searching based on the identified features.

[0343] Step 5:

[0344] When a user wants to retrieve specific information, they send a search query from their terminal to the server as a prompt. The input is a text-based query, and the output is a request for the corresponding image data. Specifically, the user might enter a command such as "Show me the inventory image of part A" into their terminal.

[0345] Step 6:

[0346] The server searches the database based on the received search query and extracts relevant image data. The input consists of a prompt and database tag information, and the output is the corresponding image data. Specifically, the server compares the query based on the stored tags and quickly extracts matching data.

[0347] Step 7:

[0348] The server sends the search results to the user's terminal, and the terminal displays the images. The input is extracted image data, and the output is the image that the user displays. Specifically, the server transfers the image data to the user's terminal, and the terminal displays that data on its screen.

[0349] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0350] This invention is an innovative image analysis system that incorporates an emotion engine to recognize user emotions, in addition to a conventional image search system. The system begins with the user uploading image data captured using a terminal to a server.

[0351] The server analyzes the received image data. First, it uses artificial intelligence to automatically identify objects within the image. This generates tags for each unique object in the image, which are then stored in a database. Additionally, OCR technology is used to extract text information contained in the image, and further tags are generated based on this information.

[0352] Furthermore, a key feature of this system is its use of an emotion engine to analyze the user's emotional state from image data. The server analyzes facial expressions and various biometric information in the images to identify the user's emotions. This emotional information is also generated as a tag and added to the relevant data in the database. This enables searches that consider not only physical information but also emotional context.

[0353] When a user searches for a specific image or related information, they send a text-based search query from their device to the server. The server quickly extracts image data tagged with the query and presents the search results to the user's device. This system can also provide recommendations based on information obtained through sentiment analysis, for example, by identifying emotional responses to specific statuses or work environments on a construction site.

[0354] For example, if a user searches for past event photos based on the emotion of "happiness," the emotion engine will extract photos tagged with the emotion "happiness" from the database. Users can find the most suitable images based on their emotional state, going beyond simply finding physical objects. This makes it easier to obtain information from a new perspective that would have been difficult to find with conventional search systems.

[0355] The following describes the processing flow.

[0356] Step 1:

[0357] The user uses their device to select image data containing emotions and sends an upload request to the server. The device then transfers the selected image files to the server via the HTTP protocol.

[0358] Step 2:

[0359] The server stores the received image data in a buffer for analysis. Then, it runs artificial intelligence to automatically identify objects in the image. This involves using an object recognition model to detect identifiable objects in the image and extract the associated object names.

[0360] Step 3:

[0361] The server generates relevant tags based on the identified objects. These generated tags indicate the type or category of the object and serve to improve search efficiency within the database.

[0362] Step 4:

[0363] The server uses OCR technology to analyze text information within image data. It detects textual information contained in the scene and converts its content into digital text. This text information is also generated as tags and associated with other elements.

[0364] Step 5:

[0365] The server recognizes the user's emotions by analyzing facial expressions and biometric information within images using an emotion engine. Based on the recognized emotion information, it generates additional emotion tags and associates them with the image data.

[0366] Step 6:

[0367] The generated object tags, text tags, and sentiment tags are stored in a database, and an index is created to allow for efficient searching of matching images.

[0368] Step 7:

[0369] To search for images related to a specific object or emotion, the user enters a text-based search query on their device and sends it to the server.

[0370] Step 8:

[0371] The server extracts image data from the database that matches the appropriate tags based on the received search query. If necessary, it refines the results by considering sentiment tags.

[0372] Step 9:

[0373] The server generates search results and sends them to the user's device. The user can view the search results on their device, select images of interest, and view more details.

[0374] (Example 2)

[0375] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0376] Modern image search systems extract only physical information from user-submitted image data, failing to consider emotions or context. This makes it difficult to perform image searches based on the emotional context desired by the user. Specifically, the inability to quickly and accurately search for images based on a particular emotional state or mood limits the user's search experience.

[0377] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0378] In this invention, the server includes means for receiving image data captured by the user, means for using artificial intelligence to identify objects in the image data, means for extracting textual information from the image data using OCR technology, and means for identifying the user's emotional state from the image data using an emotion analysis engine and generating emotion-based tags. This enables the user to perform image searches based on emotional context in addition to physical information.

[0379] "Means for receiving image data captured by a user" refers to a function that allows a server to receive image data captured by a user with their device via a network.

[0380] "Means of using artificial intelligence to identify objects in image data" refers to a process that utilizes image analysis algorithms to automatically detect and identify objects within an image.

[0381] "Means for extracting text information from image data using OCR technology" refers to a technology that analyzes text contained in an image using optical character recognition technology and extracts it as digital text.

[0382] "Methods for identifying a user's emotional state from image data using an emotion analysis engine" refers to algorithms that analyze facial expressions and biometric information in images to identify the user's emotions.

[0383] "Means for generating emotion-based tags" refers to a function that generates relevant tags based on user emotion information obtained through emotion analysis and stores them in a database.

[0384] This invention is a system that enhances image search by processing image data captured by a user using a terminal and performing object recognition and sentiment analysis. Specific embodiments for carrying out the invention are described below.

[0385] Users take pictures using devices such as smartphones or personal computers and upload the image data to the system. The uploaded data is first received by the server. On the server, the image data is preprocessed using the image processing library "OpenCV" and converted to an appropriate format for analysis.

[0386] Next, the server utilizes deep learning frameworks such as "TensorFlow" and "PyTorch" to identify objects in the image using models (e.g., VGG and ResNet). A corresponding tag is generated for each identified object, and these tags are stored in a database.

[0387] Furthermore, the server uses OCR technology such as "Tesseract" to extract text information from the image and generates additional tags based on the text. This process efficiently analyzes the characters in the image, allowing the necessary information to be obtained.

[0388] In particular, in this invention, the server is equipped with an emotion analysis engine that analyzes facial expressions and biometric information using tools such as "Haar Cascade" and "Keras." Through this analysis, the user's emotional state is identified, and tags indicating that emotion are also stored in a database.

[0389] For example, if a user wants to search for photos of events related to happiness, the system will quickly extract and present images tagged with the emotion "happiness." This allows users to perform advanced image searches based on emotional context.

[0390] Examples of prompts for a generative AI model include: "Create a description of the algorithm for an image search system that takes user emotions into consideration. Include the specific process of emotion analysis performed by the emotion engine."

[0391] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0392] Step 1:

[0393] The user takes image data using their device and uploads it to the system. This involves clicking an "upload" button via an application on the device, sending the image data to the server over the internet. The input is the image data captured by the device, and the output is the image data sent to the server.

[0394] Step 2:

[0395] The server receives the image data and first performs image preprocessing using the "OpenCV" library. This preprocessing includes standardizing the image size and format. The input is the received image data, and the output is the preprocessed image data.

[0396] Step 3:

[0397] The server then applies an object recognition model (e.g., VGG or ResNet) using a deep learning framework such as TensorFlow or PyTorch. This identifies objects in the image and generates corresponding tags. The input is preprocessed image data, and the output is the object recognition result and the generated tags.

[0398] Step 4:

[0399] The server uses OCR technology such as "Tesseract" to extract text information from image data. This process obtains text from the image as a digital string and also generates tags. The input is image data, and the output is the extracted text information and additional tags.

[0400] Step 5:

[0401] The server uses an emotion analysis engine to analyze emotional information from images. Specifically, it uses "Haar Cascade" and "Keras" to analyze facial expressions and biometric information to identify the user's emotions. The input is image data, and the output is an emotion status and emotion tag.

[0402] Step 6:

[0403] The server saves all generated tags to the database. Here, object tags, text tags, and sentiment tags are all saved together in association. The input is the various tags that have been generated, and the output is the status of the data being saved to the database.

[0404] Step 7:

[0405] The user sends a text-based search query from their device to the server. Specifically, they enter keywords into the application's search bar and click the "Search" button. The input is the query information entered by the user, and the output is the query sent to the server.

[0406] Step 8:

[0407] The server extracts relevant image data from the database based on the received search query. It searches for highly relevant images from stored tags and formats the results. The input is the received search query, and the output is the search results.

[0408] Step 9:

[0409] The server displays the extracted search results on the user's device. Here, the results are formatted to be easily viewed by the user as an image list. The input is the search results formatted by the server, and the output is the image list displayed on the device screen.

[0410] (Application Example 2)

[0411] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0412] There is a challenge in providing appropriate information and recommendations based on images taken by users, taking into account their emotions and the context at the time. Conventional image search systems primarily rely on information retrieval based on physical characteristics and are unable to provide information that takes into account the user's emotions or psychological state.

[0413] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0414] In this invention, the server includes means for receiving image data captured by the user, means for identifying objects in the image data and generating tags, and means for generating tags based on emotional information using an emotion engine for analyzing the user's emotional state from the image data. This makes it possible to provide and recommend information that takes the user's emotional state into consideration.

[0415] "Means for receiving image data captured by a user" refers to a processing device that uses communication or an interface to acquire image information captured by a user from a terminal and incorporate it into the system.

[0416] "Methods using artificial intelligence" refer to methods that use computer-based algorithms to recognize and identify objects from input data, and utilize machine learning techniques.

[0417] "Means for generating tags corresponding to identified objects" refers to a device or software that performs a process of assigning textual information or identification information to objects recognized within image data.

[0418] "Means of using OCR technology" refers to a method or apparatus for optically recognizing text information within an image and converting it into digital data.

[0419] "Means of using an emotion engine" refers to an algorithm or program that analyzes the user's facial expressions and biometric information contained in an image to determine the user's emotional state.

[0420] "Means for storing generated tags in a database" refers to a database or storage device that is permanently maintained within the system for managing identified tags and sentiment information.

[0421] "A means for receiving text-based search queries and searching for related image data" refers to a search processing device that receives textual inquiries from users and searches for and extracts related image information based on those inquiries.

[0422] "Means of providing recommended information to users based on search results" refers to notifications or display devices that analyze data obtained through searches and present the user with the most optimal or relevant information.

[0423] "Means for presenting search results to the user" refers to an interface or output device for displaying or transmitting the results obtained by the system through a search to the user's terminal.

[0424] The system for implementing this invention begins with a user sending image data captured using a smart device to a server. The server analyzes the received images using an artificial intelligence module. This AI module is a model trained using machine learning platforms such as TensorFlow and PyTorch, and has the ability to identify objects in the image. Furthermore, it uses a Tesseract engine employing OCR (optical character recognition) technology to extract text information from the image and generate tags.

[0425] The image data is then fed into an emotion engine, which uses Microsoft Azure's facial recognition API and other tools to analyze the user's facial expressions and biometric indicators, thereby identifying their emotional state. This process generates emotion tags such as "happiness" and "surprise," which are then stored in a database.

[0426] When a user searches for specific information, the device sends a text-based search query to the server. The server quickly searches for relevant image data based on tags in its database and provides information related to a specific emotional state. For example, if the user's emotion is identified as "happy," advertisements related to entertainment and leisure will be displayed.

[0427] For example, if a user sends a photo of themselves happily with a friend, it could be tagged with "happiness," and then they could be offered movie ticket promotions or event information.

[0428] An example of a prompt message might be: "The user took a photo of themselves smiling with a friend. Sentiment analysis determined that they were feeling 'happy.' Please create an advertisement that matches this emotion."

[0429] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0430] Step 1:

[0431] The user takes an image with a smart device and sends the image data from the device to the server. The input is the image data taken by the user, and the output is the image data received by the server. In this process, the image is acquired using the smart device's camera API and the data is uploaded to the server via the HTTP protocol over an internet connection.

[0432] Step 2:

[0433] The server analyzes received image data using a machine learning model to identify objects within the image. The input is the received image data, and the output is a list of tags corresponding to the identified objects. Specifically, it uses generative AI models such as TensorFlow or PyTorch to analyze features in the image and generate tags.

[0434] Step 3:

[0435] The server extracts text information from the image using an OCR engine. The input is still the received image data, and the output is additional tag information based on the extracted text. An OCR engine such as Tesseract is used here to extract string data from the image and perform analysis.

[0436] Step 4:

[0437] The server uses an emotion engine to analyze facial expressions from images and identify the user's emotional state. The input is image data containing faces, and the output is tag information based on the user's emotions. For emotion analysis, the server uses Microsoft Azure's facial recognition API, among others, to infer emotions from the identified facial expression data.

[0438] Step 5:

[0439] The server saves all generated tags to a database. The input is the generated tag information, and the output is the registration status of the tags in the program's database. The tag information is recorded in a database system such as MySQL.

[0440] Step 6:

[0441] The user sends a text-based search query from their terminal to the server to find specific information. The input is the user's query text, and the output is that text data received by the server. The user enters the query through the application's search interface.

[0442] Step 7:

[0443] The server searches the database for appropriate image data based on the tags associated with the received query. The input is the search query and tag information from the database, and the output is a set of related image data. Here, an SQL query is used to filter the data for matching tags.

[0444] Step 8:

[0445] The server generates and presents recommended information to the user based on the search results. The input is image data and related information as search results, and the output is recommended advertisements and information displayed on the user's device. Specifically, it sets matching advertisement content based on the search results and displays the results on the user's interface.

[0446] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0447] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0448] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0449] [Third Embodiment]

[0450] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0451] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0452] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0453] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0454] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0455] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0456] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0457] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0458] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0459] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0460] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0461] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0462] This invention relates to an image search system that efficiently handles image data captured by users using their personal devices. This system operates primarily through a server, and its embodiments are described below.

[0463] When a user uploads image data taken with their device to the system, the image data is sent to the server. Upon receiving this image data, the server immediately begins analysis. First, the server applies artificial intelligence technology to the image data to identify objects within the image. This is done by an artificial intelligence model that has been pre-trained to recognize specific objects. For example, since photographs of construction sites often contain objects such as cranes and rebar, the system is configured to identify these.

[0464] Next, the server uses OCR technology to extract text information present in the image. This extraction process optically recognizes characters in the image and saves them as digital text. For example, it can extract information from signs at construction sites or text within safety signs.

[0465] Based on the identified objects and the text extracted by OCR, the server generates tags associated with the photographs. These tags serve as keywords representing the object names or summaries of the extracted text. The generated tags are stored in a database, linked to the photograph data, to facilitate searching.

[0466] When a user searches for a specific object or element, they send a search query from their device to the server. The server searches its database based on the keywords specified in this search query and quickly extracts image data with relevant tags. The resulting image data is sent to the user's device and displayed on the screen. This allows the user to quickly find the information they are looking for.

[0467] For example, if a user in the construction industry searches for photos of past construction sites using the keyword "crane," the server automatically extracts photos tagged with "crane" from the database and quickly sends them to the user's terminal. This process allows users to find the necessary photos in a short amount of time, dramatically improving work efficiency.

[0468] The following describes the processing flow.

[0469] Step 1:

[0470] The user selects image data to search using their device and uploads it to the server. The device then sends the image file to the server as an HTTP request.

[0471] Step 2:

[0472] The server prepares to analyze the received image data. The server loads an artificial intelligence model and processes the image data using a computer vision algorithm to identify objects within the image.

[0473] Step 3:

[0474] The server automatically generates tags based on the features of objects extracted from images. These tags represent the type and properties of the objects and are used for subsequent searches.

[0475] Step 4:

[0476] The server uses OCR technology to analyze the text information present in the image and extracts it as digital text. This resulting text is also used as a tag.

[0477] Step 5:

[0478] The server associates the generated tags with image data and stores them in a database. This enables efficient searching.

[0479] Step 6:

[0480] To search for a specific image, the user enters a search query (keyword) on their device and sends it to the server.

[0481] Step 7:

[0482] The server searches the database based on the user's search query and extracts image data tagged with the query.

[0483] Step 8:

[0484] The server generates search results and sends them to the user's device. The user can then view the search results on their device and browse images that interest them.

[0485] (Example 1)

[0486] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0487] Conventional image search systems struggle to efficiently find specific information from large amounts of image data. For example, manual tagging is required, placing a significant burden on human resources. Furthermore, it has been difficult to automatically extract and appropriately classify information based on the objects and text information contained within the original images. This results in the challenge of users being unable to quickly search for the image data they need.

[0488] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0489] In this invention, the server includes means for receiving image data captured from a user's information processing device, means for using machine intelligence to identify objects in the image data, and means for generating relevant tags based on the identified objects and extracted text. This enables the user to efficiently and quickly search for image data related to specific keywords.

[0490] A "user information processing device" is a device, such as a computer or mobile information terminal, used by a user that enables the capture and uploading of image data to a system.

[0491] "Image data" refers to digital visual information captured by a user using an information processing device, and is stored as data in the form of still images or videos.

[0492] "Machine intelligence" refers to artificial intelligence technology used to automatically identify objects within image data, and is implemented through models trained using learning algorithms.

[0493] An "identified object" refers to a specific item or element recognized within image data using machine intelligence, and each object is identified based on a pre-defined classification.

[0494] "Optical character recognition technology" is a technology that mechanically extracts character information contained within image data as digital text, and is a process executed by an optical processor.

[0495] A "tag" refers to keywords or metadata generated based on identified objects and extracted text, and is information added to facilitate searching image data.

[0496] A "storage device" is a device or system that functions as a database or storage system and is used to hold generated tags and associated image data.

[0497] A "text-based information request" is a search query in string format entered by a user, used to communicate information about specific keywords to the system.

[0498] "Searching" is the process of searching for data within a storage device based on a user's request and finding relevant image data.

[0499] This invention is an image search system that efficiently manages image data captured by users using an information processing device and allows for rapid retrieval of specific objects or text information. It is primarily server-based and functions as follows:

[0500] The user first takes an image using an information processing device. This device is a common device such as a computer or mobile terminal with a camera function. This image data is uploaded to a server via an internet connection using a secure protocol (e.g., HTTPS).

[0501] The server temporarily stores the received image data in file storage. Next, it utilizes a generative AI model to identify objects within the image data. This AI model is trained using machine learning algorithms and can utilize deep learning techniques. For example, in a photograph of a construction site, it can be configured to recognize specific objects such as cranes and reinforcing bars.

[0502] The server also uses optical character recognition (OCR) technology to extract text information from images. The OCR process employs common optical character recognition software (e.g., Tesseract) to detect characters from image data and save them as digital text. For example, it can retrieve the content written on signs and notices.

[0503] Based on the identified objects and extracted text information, the server generates relevant tags and stores them in a database. This allows users to quickly retrieve relevant information when searching for images using specific keywords. SQL or NoSQL databases can be used for this purpose.

[0504] As a concrete example, if a user in the construction industry searches for images of past construction sites using the keyword "crane," the server extracts images tagged with "crane" and provides them to the user quickly. This process allows users to find the necessary information in a short amount of time, significantly improving work efficiency.

[0505] An example prompt might be a request like, "Search for photos containing cranes in construction site images and generate related text information as tags." This system could be a powerful tool, especially in industries that need to handle large amounts of image data.

[0506] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0507] Step 1:

[0508] The user captures image data with an information processing device. The device is a camera-equipped device that acquires the captured image data. The input is the captured image, and the output is that image data. This step involves the user capturing a specific scene or object and saving it in a digital format.

[0509] Step 2:

[0510] The device uploads captured image data to the server. During this process, the device transmits the image data using a secure communication protocol. The input is the image data provided by the user's device, and the output is the data transferred to the server. This step involves data transmission via an internet connection.

[0511] Step 3:

[0512] The server saves the received image data to file storage. The input is the uploaded image data, and the output is the image data stored in storage. This step involves the operation of saving the data to the storage device hosted by the server.

[0513] Step 4:

[0514] The server generates image data and analyzes it using an AI model to identify objects within the image. The input is stored image data, and the output is a list of identified objects. At this stage, the AI ​​model performs the process of feature extraction and classification.

[0515] Step 5:

[0516] The server extracts text information from an image using optical character recognition (OCR) technology. The input is the image data after identification is complete, and the output is the extracted text information. In this step, the text is digitized using the OCR process.

[0517] Step 6:

[0518] The server generates relevant tags based on identified objects and extracted text. The input is a list of objects and text information, and the output is the generated tags. This process also generates object names and keywords.

[0519] Step 7:

[0520] The server saves the generated tags to the database. The input is the newly generated tags, and the output is the updated database entry. This step involves registering and maintaining information in the database.

[0521] Step 8:

[0522] A user sends a search query from their device to the server using specific keywords. The input is the search query entered by the user, and the output is the search request forwarded to the server. In this process, query input and submission occur via the user interface.

[0523] Step 9:

[0524] The server extracts image data from the database that have tags matching the search query. The input is the search query and the database, and the output is image data as the search results. At this stage, a database search algorithm is applied.

[0525] Step 10:

[0526] The server sends the search results to the user's terminal. The input is image data obtained through the search, and the output is the information displayed on the user's terminal. This step involves data transfer over the internet and display on the user's screen.

[0527] (Application Example 1)

[0528] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0529] Managing goods and machinery within a factory generates a large amount of image data, and there is a need to efficiently analyze and search this data. Conventional systems have the problem of requiring significant time and effort to manually classify and search image data. Furthermore, the lack of readily available information on-site makes rapid decision-making and inventory management difficult.

[0530] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0531] In this invention, the server includes means for receiving image data captured by a user, means for using artificial intelligence to identify objects in the image data, means for generating tags corresponding to the identified objects, means for extracting text information from the image data using OCR technology and generating additional tags based on this text information, means for storing the generated tags in a database, means for receiving text-based search queries and searching for related image data based on the stored tags, means for presenting the search results to the user, means for utilizing a remote information display device or a mobile automated machine as a shooting terminal, and means for providing a system for efficiently managing object and machine information based on the generated tags. This enables real-time information acquisition within the factory, efficient inventory management, and rapid decision-making.

[0532] A "user" is an individual or group that operates an image capture device or information display device to acquire or view image data.

[0533] "Image data" refers to data that digitally represents the visual information of objects or scenes captured by a user.

[0534] Artificial intelligence is a collection of algorithms and models that computer systems use to enable learning and recognition.

[0535] A "tag" is a keyword or label used to identify and summarize objects and text information contained in image data.

[0536] "OCR technology" is a technology that optically recognizes text within an image and converts it into digital text.

[0537] A "database" is a digital information management system that systematically stores generated tags and other related information, enabling information retrieval.

[0538] A "search query" is a text-based request that a user enters into a system to retrieve specific information.

[0539] A "remote information display device" is a visual device that allows users to acquire information remotely by wearing or carrying it.

[0540] A "mobile automated machine" is a machine or device that operates autonomously in an environment such as a factory and is capable of taking images and collecting data.

[0541] "Inventory management" refers to activities aimed at efficiently understanding and managing the location, quantity, and condition of items within an environment such as a factory or warehouse.

[0542] The system for implementing this invention efficiently manages image data captured by the user and quickly searches for and retrieves necessary information. The user uses a remote information display device or a mobile automated machine as the imaging device. For example, a worker wearing smart glasses or a self-propelled robot patrols the factory and takes pictures of products and machinery.

[0543] The server receives and analyzes captured image data in real time. Specifically, the server uses artificial intelligence technology to identify objects within the image data and uses OCR technology to optically recognize text information within the image and extract it as digital text. Based on the identified objects and extracted text information, the server generates relevant tags and stores them in a database.

[0544] This system utilizes generative AI models based on TensorFlow and PyTorch to perform image analysis. It also uses the Google Cloud Vision API for OCR processing. A database management system such as MySQL is used for the database.

[0545] When a user searches for specific information, this is achieved by sending a text-based search query from the terminal to the server. For example, by entering a prompt such as "I want to check the location of the crane," the server quickly searches for relevant image data and displays it on the terminal. This allows users to efficiently manage items and obtain machine information within the factory.

[0546] Specific examples of prompt statements include "Show me the inventory image of part A" and "I want to check the maintenance information for machine 123." Through these prompt statements, users can instantly obtain the information they need on-site and improve the efficiency of their work.

[0547] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0548] Step 1:

[0549] The user takes images within the factory using a remote information display device or mobile automated machine. The input is visual information of the environment, and the output is captured image data. This image data is immediately uploaded to a cloud server.

[0550] Step 2:

[0551] The server retrieves the received image data and applies a generative AI model using TensorFlow or PyTorch. The input is the captured image data, and the output is a result in which objects are identified. Specifically, the server inputs the image into the model, analyzes the features of each pixel, and identifies the type of object.

[0552] Step 3:

[0553] The server performs OCR processing using the Google Cloud Vision API. The input is image data, and the output is extracted text information. Specifically, the server identifies text regions within the image and converts their content into digital text.

[0554] Step 4:

[0555] The server generates relevant tags based on the obtained object identification results and text information, and stores them in the database. The input is the object identification results and text information, and the output is a database entry containing the generated tags. Specifically, the server extracts keywords useful for searching based on the identified features.

[0556] Step 5:

[0557] When a user wants to retrieve specific information, they send a search query from their terminal to the server as a prompt. The input is a text-based query, and the output is a request for the corresponding image data. Specifically, the user might enter a command such as "Show me the inventory image of part A" into their terminal.

[0558] Step 6:

[0559] The server searches the database based on the received search query and extracts relevant image data. The input consists of a prompt and database tag information, and the output is the corresponding image data. Specifically, the server compares the query based on the stored tags and quickly extracts matching data.

[0560] Step 7:

[0561] The server sends the search results to the user's terminal, and the terminal displays the images. The input is extracted image data, and the output is the image that the user displays. Specifically, the server transfers the image data to the user's terminal, and the terminal displays that data on its screen.

[0562] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0563] This invention is an innovative image analysis system that incorporates an emotion engine to recognize user emotions, in addition to a conventional image search system. The system begins with the user uploading image data captured using a terminal to a server.

[0564] The server analyzes the received image data. First, it uses artificial intelligence to automatically identify objects within the image. This generates tags for each unique object in the image, which are then stored in a database. Additionally, OCR technology is used to extract text information contained in the image, and further tags are generated based on this information.

[0565] Furthermore, a key feature of this system is its use of an emotion engine to analyze the user's emotional state from image data. The server analyzes facial expressions and various biometric information in the images to identify the user's emotions. This emotional information is also generated as a tag and added to the relevant data in the database. This enables searches that consider not only physical information but also emotional context.

[0566] When a user searches for a specific image or related information, they send a text-based search query from their device to the server. The server quickly extracts image data tagged with the query and presents the search results to the user's device. This system can also provide recommendations based on information obtained through sentiment analysis, for example, by identifying emotional responses to specific statuses or work environments on a construction site.

[0567] For example, if a user searches for past event photos based on the emotion of "happiness," the emotion engine will extract photos tagged with the emotion "happiness" from the database. Users can find the most suitable images based on their emotional state, going beyond simply finding physical objects. This makes it easier to obtain information from a new perspective that would have been difficult to find with conventional search systems.

[0568] The following describes the processing flow.

[0569] Step 1:

[0570] The user uses their device to select image data containing emotions and sends an upload request to the server. The device then transfers the selected image files to the server via the HTTP protocol.

[0571] Step 2:

[0572] The server stores the received image data in a buffer for analysis. Then, it runs artificial intelligence to automatically identify objects in the image. This involves using an object recognition model to detect identifiable objects in the image and extract the associated object names.

[0573] Step 3:

[0574] The server generates relevant tags based on the identified objects. These generated tags indicate the type or category of the object and serve to improve search efficiency within the database.

[0575] Step 4:

[0576] The server uses OCR technology to analyze text information within image data. It detects textual information contained in the scene and converts its content into digital text. This text information is also generated as tags and associated with other elements.

[0577] Step 5:

[0578] The server recognizes the user's emotions by analyzing facial expressions and biometric information within images using an emotion engine. Based on the recognized emotion information, it generates additional emotion tags and associates them with the image data.

[0579] Step 6:

[0580] The generated object tags, text tags, and sentiment tags are stored in a database, and an index is created to allow for efficient searching of matching images.

[0581] Step 7:

[0582] To search for images related to a specific object or emotion, the user enters a text-based search query on their device and sends it to the server.

[0583] Step 8:

[0584] The server extracts image data from the database that matches the appropriate tags based on the received search query. If necessary, it refines the results by considering sentiment tags.

[0585] Step 9:

[0586] The server generates search results and sends them to the user's device. The user can view the search results on their device, select images of interest, and view more details.

[0587] (Example 2)

[0588] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0589] Modern image search systems extract only physical information from user-submitted image data, failing to consider emotions or context. This makes it difficult to perform image searches based on the emotional context desired by the user. Specifically, the inability to quickly and accurately search for images based on a particular emotional state or mood limits the user's search experience.

[0590] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0591] In this invention, the server includes means for receiving image data captured by the user, means for using artificial intelligence to identify objects in the image data, means for extracting textual information from the image data using OCR technology, and means for identifying the user's emotional state from the image data using an emotion analysis engine and generating emotion-based tags. This enables the user to perform image searches based on emotional context in addition to physical information.

[0592] "Means for receiving image data captured by a user" refers to a function that allows a server to receive image data captured by a user with their device via a network.

[0593] "Means of using artificial intelligence to identify objects in image data" refers to a process that utilizes image analysis algorithms to automatically detect and identify objects within an image.

[0594] "Means for extracting text information from image data using OCR technology" refers to a technology that analyzes text contained in an image using optical character recognition technology and extracts it as digital text.

[0595] "Methods for identifying a user's emotional state from image data using an emotion analysis engine" refers to algorithms that analyze facial expressions and biometric information in images to identify the user's emotions.

[0596] "Means for generating emotion-based tags" refers to a function that generates relevant tags based on user emotion information obtained through emotion analysis and stores them in a database.

[0597] This invention is a system that enhances image search by processing image data captured by a user using a terminal and performing object recognition and sentiment analysis. Specific embodiments for carrying out the invention are described below.

[0598] Users take pictures using devices such as smartphones or personal computers and upload the image data to the system. The uploaded data is first received by the server. On the server, the image data is preprocessed using the image processing library "OpenCV" and converted to an appropriate format for analysis.

[0599] Next, the server utilizes deep learning frameworks such as "TensorFlow" and "PyTorch" to identify objects in the image using models (e.g., VGG and ResNet). A corresponding tag is generated for each identified object, and these tags are stored in a database.

[0600] Furthermore, the server uses OCR technology such as "Tesseract" to extract text information from the image and generates additional tags based on the text. This process efficiently analyzes the characters in the image, allowing the necessary information to be obtained.

[0601] In particular, in this invention, the server is equipped with an emotion analysis engine that analyzes facial expressions and biometric information using "Haar Cascade" or "Keras," among others. Through this analysis, the user's emotional state is identified, and tags indicating that emotion are also stored in a database.

[0602] For example, if a user wants to search for photos of events related to happiness, the system will quickly extract and present images tagged with the emotion "happiness." This allows users to perform advanced image searches based on emotional context.

[0603] Examples of prompts for a generative AI model include: "Create a description of the algorithm for an image search system that takes user emotions into consideration. Include the specific process of emotion analysis performed by the emotion engine."

[0604] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0605] Step 1:

[0606] The user takes image data using their device and uploads it to the system. This involves clicking an "upload" button via an application on the device, sending the image data to the server over the internet. The input is the image data captured by the device, and the output is the image data sent to the server.

[0607] Step 2:

[0608] The server receives the image data and first performs image preprocessing using the "OpenCV" library. This preprocessing includes standardizing the image size and format. The input is the received image data, and the output is the preprocessed image data.

[0609] Step 3:

[0610] The server then applies an object recognition model (e.g., VGG or ResNet) using a deep learning framework such as TensorFlow or PyTorch. This identifies objects in the image and generates corresponding tags. The input is preprocessed image data, and the output is the object recognition result and the generated tags.

[0611] Step 4:

[0612] The server uses OCR technology such as "Tesseract" to extract text information from image data. This process obtains text from the image as a digital string and also generates tags. The input is image data, and the output is the extracted text information and additional tags.

[0613] Step 5:

[0614] The server uses an emotion analysis engine to analyze emotional information from images. Specifically, it uses "Haar Cascade" and "Keras" to analyze facial expressions and biometric information to identify the user's emotions. The input is image data, and the output is an emotion status and emotion tag.

[0615] Step 6:

[0616] The server saves all generated tags to the database. Here, object tags, text tags, and sentiment tags are all saved together in association. The input is the various tags that have been generated, and the output is the status of the data being saved to the database.

[0617] Step 7:

[0618] The user sends a text-based search query from their device to the server. Specifically, they enter keywords into the application's search bar and click the "Search" button. The input is the query information entered by the user, and the output is the query sent to the server.

[0619] Step 8:

[0620] The server extracts relevant image data from the database based on the received search query. It searches for highly relevant images from stored tags and formats the results. The input is the received search query, and the output is the search results.

[0621] Step 9:

[0622] The server displays the extracted search results on the user's device. Here, the results are formatted to be easily viewed by the user as an image list. The input is the search results formatted by the server, and the output is the image list displayed on the device screen.

[0623] (Application Example 2)

[0624] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0625] There is a challenge in providing appropriate information and recommendations based on images taken by users, taking into account their emotions and the context at the time. Conventional image search systems primarily rely on information retrieval based on physical characteristics and are unable to provide information that takes into account the user's emotions or psychological state.

[0626] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0627] In this invention, the server includes means for receiving image data captured by the user, means for identifying objects in the image data and generating tags, and means for generating tags based on emotional information using an emotion engine for analyzing the user's emotional state from the image data. This makes it possible to provide and recommend information that takes the user's emotional state into consideration.

[0628] "Means for receiving image data captured by a user" refers to a processing device that uses communication or an interface to acquire image information captured by a user from a terminal and incorporate it into the system.

[0629] "Methods using artificial intelligence" refer to methods that use computer-based algorithms to recognize and identify objects from input data, and utilize machine learning techniques.

[0630] "Means for generating tags corresponding to identified objects" refers to a device or software that performs a process of assigning textual information or identification information to objects recognized within image data.

[0631] "Means of using OCR technology" refers to a method or apparatus for optically recognizing text information within an image and converting it into digital data.

[0632] "Means of using an emotion engine" refers to an algorithm or program that analyzes the user's facial expressions and biometric information contained in an image to determine the user's emotional state.

[0633] "Means for storing generated tags in a database" refers to a database or storage device that is permanently maintained within the system for managing identified tags and sentiment information.

[0634] A "means for receiving text-based search queries and searching for related image data" refers to a search processing device that receives textual inquiries from users and searches for and extracts related image information based on those inquiries.

[0635] "Means of providing recommended information to users based on search results" refers to notifications or display devices that analyze data obtained through searches and present the user with the most optimal or relevant information.

[0636] "Means for presenting search results to the user" refers to an interface or output device for displaying or transmitting the results obtained by the system through a search to the user's terminal.

[0637] The system for implementing this invention begins with a user sending image data captured using a smart device to a server. The server analyzes the received images using an artificial intelligence module. This AI module is a model trained using machine learning platforms such as TensorFlow and PyTorch, and has the ability to identify objects in the image. Furthermore, it uses a Tesseract engine employing OCR (optical character recognition) technology to extract text information from the image and generate tags.

[0638] The image data is then fed into an emotion engine, which uses Microsoft Azure's facial recognition API and other tools to analyze the user's facial expressions and biometric indicators, thereby identifying their emotional state. This process generates emotion tags such as "happiness" and "surprise," which are then stored in a database.

[0639] When a user searches for specific information, the device sends a text-based search query to the server. The server quickly searches for relevant image data based on tags in its database and provides information related to a specific emotional state. For example, if the user's emotion is identified as "happy," advertisements related to entertainment and leisure will be displayed.

[0640] For example, if a user sends a photo of themselves happily with a friend, it could be tagged with "happiness," and then they could be offered movie ticket promotions or event information.

[0641] An example of a prompt message might be: "The user took a photo of themselves smiling with a friend. Sentiment analysis determined that they were feeling 'happy.' Please create an advertisement that matches this emotion."

[0642] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0643] Step 1:

[0644] The user takes an image with a smart device and sends the image data from the device to the server. The input is the image data taken by the user, and the output is the image data received by the server. In this process, the image is acquired using the smart device's camera API and the data is uploaded to the server via the HTTP protocol over an internet connection.

[0645] Step 2:

[0646] The server analyzes received image data using a machine learning model to identify objects within the image. The input is the received image data, and the output is a list of tags corresponding to the identified objects. Specifically, it uses generative AI models such as TensorFlow or PyTorch to analyze features in the image and generate tags.

[0647] Step 3:

[0648] The server extracts text information from the image using an OCR engine. The input is still the received image data, and the output is additional tag information based on the extracted text. An OCR engine such as Tesseract is used here to extract string data from the image and perform analysis.

[0649] Step 4:

[0650] The server uses an emotion engine to analyze facial expressions from images and identify the user's emotional state. The input is image data containing faces, and the output is tag information based on the user's emotions. For emotion analysis, the server uses Microsoft Azure's facial recognition API, among others, to infer emotions from the identified facial expression data.

[0651] Step 5:

[0652] The server saves all generated tags to a database. The input is the generated tag information, and the output is the registration status of the tags in the program's database. The tag information is recorded in a database system such as MySQL.

[0653] Step 6:

[0654] The user sends a text-based search query from their terminal to the server to find specific information. The input is the user's query text, and the output is that text data received by the server. The user enters the query through the application's search interface.

[0655] Step 7:

[0656] The server searches the database for appropriate image data based on the tags associated with the received query. The input is the search query and tag information from the database, and the output is a set of related image data. Here, an SQL query is used to filter the data for matching tags.

[0657] Step 8:

[0658] The server generates and presents recommended information to the user based on the search results. The input is image data and related information as search results, and the output is recommended advertisements and information displayed on the user's device. Specifically, it sets matching advertisement content based on the search results and displays the results on the user's interface.

[0659] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0660] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0661] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0662] [Fourth Embodiment]

[0663] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0664] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0665] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0666] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0667] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0668] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0669] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0670] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0671] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0672] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0673] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0674] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0675] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0676] This invention relates to an image search system that efficiently handles image data captured by users using their personal devices. This system operates primarily through a server, and its embodiments are described below.

[0677] When a user uploads image data taken with their device to the system, the image data is sent to the server. Upon receiving this image data, the server immediately begins analysis. First, the server applies artificial intelligence technology to the image data to identify objects within the image. This is done by an artificial intelligence model that has been pre-trained to recognize specific objects. For example, since photographs of construction sites often contain objects such as cranes and rebar, the system is configured to identify these.

[0678] Next, the server uses OCR technology to extract text information present in the image. This extraction process optically recognizes characters in the image and saves them as digital text. For example, it can extract information from signs at construction sites or text within safety signs.

[0679] Based on the identified objects and the text extracted by OCR, the server generates tags associated with the photographs. These tags serve as keywords representing the object names or summaries of the extracted text. The generated tags are stored in a database, linked to the photograph data, to facilitate searching.

[0680] When a user searches for a specific object or element, they send a search query from their device to the server. The server searches its database based on the keywords specified in this search query and quickly extracts image data with relevant tags. The resulting image data is sent to the user's device and displayed on the screen. This allows the user to quickly find the information they are looking for.

[0681] For example, if a user in the construction industry searches for photos of past construction sites using the keyword "crane," the server automatically extracts photos tagged with "crane" from the database and quickly sends them to the user's terminal. This process allows users to find the necessary photos in a short amount of time, dramatically improving work efficiency.

[0682] The following describes the processing flow.

[0683] Step 1:

[0684] The user selects image data to search using their device and uploads it to the server. The device then sends the image file to the server as an HTTP request.

[0685] Step 2:

[0686] The server prepares to analyze the received image data. The server loads an artificial intelligence model and processes the image data using a computer vision algorithm to identify objects within the image.

[0687] Step 3:

[0688] The server automatically generates tags based on the features of objects extracted from images. These tags represent the type and properties of the objects and are used for subsequent searches.

[0689] Step 4:

[0690] The server uses OCR technology to analyze the text information present in the image and extracts it as digital text. This resulting text is also used as a tag.

[0691] Step 5:

[0692] The server associates the generated tags with image data and stores them in a database. This enables efficient searching.

[0693] Step 6:

[0694] To search for a specific image, the user enters a search query (keyword) on their device and sends it to the server.

[0695] Step 7:

[0696] The server searches the database based on the user's search query and extracts image data tagged with the query.

[0697] Step 8:

[0698] The server generates search results and sends them to the user's device. The user can then view the search results on their device and browse images that interest them.

[0699] (Example 1)

[0700] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0701] Conventional image search systems struggle to efficiently find specific information from large amounts of image data. For example, manual tagging is required, placing a significant burden on human resources. Furthermore, it has been difficult to automatically extract and appropriately classify information based on the objects and text information contained within the original images. This results in the challenge of users being unable to quickly search for the image data they need.

[0702] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0703] In this invention, the server includes means for receiving image data captured from a user's information processing device, means for using machine intelligence to identify objects in the image data, and means for generating relevant tags based on the identified objects and extracted text. This enables the user to efficiently and quickly search for image data related to specific keywords.

[0704] A "user information processing device" is a device, such as a computer or mobile information terminal, used by a user that enables the capture and uploading of image data to a system.

[0705] "Image data" refers to digital visual information captured by a user using an information processing device, and is stored as data in the form of still images or videos.

[0706] "Machine intelligence" refers to artificial intelligence technology used to automatically identify objects within image data, and is implemented through models trained using learning algorithms.

[0707] An "identified object" refers to a specific item or element recognized within image data using machine intelligence, and each object is identified based on a pre-defined classification.

[0708] "Optical character recognition technology" is a technology that mechanically extracts character information contained within image data as digital text, and is a process executed by an optical processor.

[0709] A "tag" refers to keywords or metadata generated based on identified objects and extracted text, and is information added to facilitate searching image data.

[0710] A "storage device" is a device or system that functions as a database or storage system and is used to hold generated tags and associated image data.

[0711] A "text-based information request" is a search query in string format entered by a user, used to communicate information about specific keywords to the system.

[0712] "Searching" is the process of searching for data within a storage device based on a user's request and finding relevant image data.

[0713] This invention is an image search system that efficiently manages image data captured by users using an information processing device and allows for rapid retrieval of specific objects or text information. It is primarily server-based and functions as follows:

[0714] The user first takes an image using an information processing device. This device is a common device such as a computer or mobile terminal with a camera function. This image data is uploaded to a server via an internet connection using a secure protocol (e.g., HTTPS).

[0715] The server temporarily stores the received image data in file storage. Next, it utilizes a generative AI model to identify objects within the image data. This AI model is trained using machine learning algorithms and can utilize deep learning techniques. For example, in a photograph of a construction site, it can be configured to recognize specific objects such as cranes and reinforcing bars.

[0716] The server also uses optical character recognition (OCR) technology to extract text information from images. The OCR process employs common optical character recognition software (e.g., Tesseract) to detect characters from image data and save them as digital text. For example, it can retrieve the content written on signs and notices.

[0717] Based on the identified objects and extracted text information, the server generates relevant tags and stores them in a database. This allows users to quickly retrieve relevant information when searching for images using specific keywords. SQL or NoSQL databases can be used for this purpose.

[0718] As a concrete example, if a user in the construction industry searches for images of past construction sites using the keyword "crane," the server extracts images tagged with "crane" and provides them to the user quickly. This process allows users to find the necessary information in a short amount of time, significantly improving work efficiency.

[0719] An example prompt might be a request like, "Search for photos containing cranes in construction site images and generate related text information as tags." This system could be a powerful tool, especially in industries that need to handle large amounts of image data.

[0720] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0721] Step 1:

[0722] The user captures image data with an information processing device. The device is a camera-equipped device that acquires the captured image data. The input is the captured image, and the output is that image data. This step involves the user capturing a specific scene or object and saving it in a digital format.

[0723] Step 2:

[0724] The device uploads captured image data to the server. During this process, the device transmits the image data using a secure communication protocol. The input is the image data provided by the user's device, and the output is the data transferred to the server. This step involves data transmission via an internet connection.

[0725] Step 3:

[0726] The server saves the received image data to file storage. The input is the uploaded image data, and the output is the image data stored in storage. This step involves the operation of saving the data to the storage device hosted by the server.

[0727] Step 4:

[0728] The server generates image data and analyzes it using an AI model to identify objects within the image. The input is stored image data, and the output is a list of identified objects. At this stage, the AI ​​model performs the process of feature extraction and classification.

[0729] Step 5:

[0730] The server extracts text information from an image using optical character recognition (OCR) technology. The input is the image data after identification is complete, and the output is the extracted text information. In this step, the text is digitized using the OCR process.

[0731] Step 6:

[0732] The server generates relevant tags based on identified objects and extracted text. The input is a list of objects and text information, and the output is the generated tags. This process also generates object names and keywords.

[0733] Step 7:

[0734] The server saves the generated tags to the database. The input is the newly generated tags, and the output is the updated database entry. This step involves registering and maintaining information in the database.

[0735] Step 8:

[0736] A user sends a search query from their device to the server using specific keywords. The input is the search query entered by the user, and the output is the search request forwarded to the server. In this process, query input and submission occur via the user interface.

[0737] Step 9:

[0738] The server extracts image data from the database that have tags matching the search query. The input is the search query and the database, and the output is image data as the search results. At this stage, a database search algorithm is applied.

[0739] Step 10:

[0740] The server sends the search results to the user's terminal. The input is image data obtained through the search, and the output is the information displayed on the user's terminal. This step involves data transfer over the internet and display on the user's screen.

[0741] (Application Example 1)

[0742] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0743] Managing goods and machinery within a factory generates a large amount of image data, and there is a need to efficiently analyze and search this data. Conventional systems have the problem of requiring significant time and effort to manually classify and search image data. Furthermore, the lack of readily available information on-site makes rapid decision-making and inventory management difficult.

[0744] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0745] In this invention, the server includes means for receiving image data captured by a user, means for using artificial intelligence to identify objects in the image data, means for generating tags corresponding to the identified objects, means for extracting text information from the image data using OCR technology and generating additional tags based on this text information, means for storing the generated tags in a database, means for receiving text-based search queries and searching for related image data based on the stored tags, means for presenting the search results to the user, means for utilizing a remote information display device or a mobile automated machine as a shooting terminal, and means for providing a system for efficiently managing object and machine information based on the generated tags. This enables real-time information acquisition within the factory, efficient inventory management, and rapid decision-making.

[0746] A "user" is an individual or group that operates an image capture device or information display device to acquire or view image data.

[0747] "Image data" refers to data that digitally represents the visual information of objects or scenes captured by a user.

[0748] Artificial intelligence is a collection of algorithms and models that computer systems use to enable learning and recognition.

[0749] A "tag" is a keyword or label used to identify and summarize objects and text information contained in image data.

[0750] "OCR technology" is a technology that optically recognizes text within an image and converts it into digital text.

[0751] A "database" is a digital information management system that systematically stores generated tags and other related information, enabling information retrieval.

[0752] A "search query" is a text-based request that a user enters into a system to retrieve specific information.

[0753] A "remote information display device" is a visual device that allows users to acquire information remotely by wearing or carrying it.

[0754] A "mobile automated machine" is a machine or device that operates autonomously in an environment such as a factory and is capable of taking images and collecting data.

[0755] "Inventory management" refers to activities aimed at efficiently understanding and managing the location, quantity, and condition of items within an environment such as a factory or warehouse.

[0756] The system for implementing this invention efficiently manages image data captured by the user and quickly searches for and retrieves necessary information. The user uses a remote information display device or a mobile automated machine as the imaging device. For example, a worker wearing smart glasses or a self-propelled robot patrols the factory and takes pictures of products and machinery.

[0757] The server receives and analyzes captured image data in real time. Specifically, the server uses artificial intelligence technology to identify objects within the image data and uses OCR technology to optically recognize text information within the image and extract it as digital text. Based on the identified objects and extracted text information, the server generates relevant tags and stores them in a database.

[0758] This system utilizes generative AI models based on TensorFlow and PyTorch to perform image analysis. It also uses the Google Cloud Vision API for OCR processing. A database management system such as MySQL is used for the database.

[0759] When a user searches for specific information, this is achieved by sending a text-based search query from the terminal to the server. For example, by entering a prompt such as "I want to check the location of the crane," the server quickly searches for relevant image data and displays it on the terminal. This allows users to efficiently manage items and obtain machine information within the factory.

[0760] Specific examples of prompt statements include "Show me the inventory image of part A" and "I want to check the maintenance information for machine 123." Through these prompt statements, users can instantly obtain the information they need on-site and improve the efficiency of their work.

[0761] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0762] Step 1:

[0763] The user takes images within the factory using a remote information display device or mobile automated machine. The input is visual information of the environment, and the output is captured image data. This image data is immediately uploaded to a cloud server.

[0764] Step 2:

[0765] The server retrieves the received image data and applies a generative AI model using TensorFlow or PyTorch. The input is the captured image data, and the output is a result in which objects are identified. Specifically, the server inputs the image into the model, analyzes the features of each pixel, and identifies the type of object.

[0766] Step 3:

[0767] The server performs OCR processing using the Google Cloud Vision API. The input is image data, and the output is extracted text information. Specifically, the server identifies text regions within the image and converts their content into digital text.

[0768] Step 4:

[0769] The server generates relevant tags based on the obtained object identification results and text information, and stores them in the database. The input is the object identification results and text information, and the output is a database entry containing the generated tags. Specifically, the server extracts keywords useful for searching based on the identified features.

[0770] Step 5:

[0771] When a user wants to retrieve specific information, they send a search query from their terminal to the server as a prompt. The input is a text-based query, and the output is a request for the corresponding image data. Specifically, the user might enter a command such as "Show me the inventory image of part A" into their terminal.

[0772] Step 6:

[0773] The server searches the database based on the received search query and extracts relevant image data. The input consists of a prompt and database tag information, and the output is the corresponding image data. Specifically, the server compares the query based on the stored tags and quickly extracts matching data.

[0774] Step 7:

[0775] The server sends the search results to the user's terminal, and the terminal displays the images. The input is extracted image data, and the output is the image that the user displays. Specifically, the server transfers the image data to the user's terminal, and the terminal displays that data on its screen.

[0776] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0777] This invention is an innovative image analysis system that incorporates an emotion engine to recognize user emotions, in addition to a conventional image search system. The system begins with the user uploading image data captured using a terminal to a server.

[0778] The server analyzes the received image data. First, it uses artificial intelligence to automatically identify objects within the image. This generates tags for each unique object in the image, which are then stored in a database. Additionally, OCR technology is used to extract text information contained in the image, and further tags are generated based on this information.

[0779] Furthermore, a key feature of this system is its use of an emotion engine to analyze the user's emotional state from image data. The server analyzes facial expressions and various biometric information in the images to identify the user's emotions. This emotional information is also generated as a tag and added to the relevant data in the database. This enables searches that consider not only physical information but also emotional context.

[0780] When a user searches for a specific image or related information, they send a text-based search query from their device to the server. The server quickly extracts image data tagged with the query and presents the search results to the user's device. This system can also provide recommendations based on information obtained through sentiment analysis, for example, by identifying emotional responses to specific statuses or work environments on a construction site.

[0781] For example, if a user searches for past event photos based on the emotion of "happiness," the emotion engine will extract photos tagged with the emotion "happiness" from the database. Users can find the most suitable images based on their emotional state, going beyond simply finding physical objects. This makes it easier to obtain information from a new perspective that would have been difficult to find with conventional search systems.

[0782] The following describes the processing flow.

[0783] Step 1:

[0784] The user uses their device to select image data containing emotions and sends an upload request to the server. The device then transfers the selected image files to the server via the HTTP protocol.

[0785] Step 2:

[0786] The server stores the received image data in a buffer for analysis. Then, it runs artificial intelligence to automatically identify objects in the image. This involves using an object recognition model to detect identifiable objects in the image and extract the associated object names.

[0787] Step 3:

[0788] The server generates relevant tags based on the identified objects. These generated tags indicate the type or category of the object and serve to improve search efficiency within the database.

[0789] Step 4:

[0790] The server uses OCR technology to analyze text information within image data. It detects textual information contained in the scene and converts its content into digital text. This text information is also generated as tags and associated with other elements.

[0791] Step 5:

[0792] The server recognizes the user's emotions by analyzing facial expressions and biometric information within images using an emotion engine. Based on the recognized emotion information, it generates additional emotion tags and associates them with the image data.

[0793] Step 6:

[0794] The generated object tags, text tags, and sentiment tags are stored in a database, and an index is created to allow for efficient searching of matching images.

[0795] Step 7:

[0796] To search for images related to a specific object or emotion, the user enters a text-based search query on their device and sends it to the server.

[0797] Step 8:

[0798] The server extracts image data from the database that matches the appropriate tags based on the received search query. If necessary, it refines the results by considering sentiment tags.

[0799] Step 9:

[0800] The server generates search results and sends them to the user's device. The user can view the search results on their device, select images of interest, and view more details.

[0801] (Example 2)

[0802] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0803] Modern image search systems extract only physical information from user-submitted image data, failing to consider emotions or context. This makes it difficult to perform image searches based on the emotional context desired by the user. Specifically, the inability to quickly and accurately search for images based on a particular emotional state or mood limits the user's search experience.

[0804] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0805] In this invention, the server includes means for receiving image data captured by the user, means for using artificial intelligence to identify objects in the image data, means for extracting textual information from the image data using OCR technology, and means for identifying the user's emotional state from the image data using an emotion analysis engine and generating emotion-based tags. This enables the user to perform image searches based on emotional context in addition to physical information.

[0806] "Means for receiving image data captured by a user" refers to a function that allows a server to receive image data captured by a user with their device via a network.

[0807] "Means of using artificial intelligence to identify objects in image data" refers to a process that utilizes image analysis algorithms to automatically detect and identify objects within an image.

[0808] "Means for extracting text information from image data using OCR technology" refers to a technology that analyzes text contained in an image using optical character recognition technology and extracts it as digital text.

[0809] "Methods for identifying a user's emotional state from image data using an emotion analysis engine" refers to algorithms that analyze facial expressions and biometric information in images to identify the user's emotions.

[0810] "Means for generating emotion-based tags" refers to a function that generates relevant tags based on user emotion information obtained through emotion analysis and stores them in a database.

[0811] This invention is a system that enhances image search by processing image data captured by a user using a terminal and performing object recognition and sentiment analysis. Specific embodiments for carrying out the invention are described below.

[0812] Users take pictures using devices such as smartphones or personal computers and upload the image data to the system. The uploaded data is first received by the server. On the server, the image data is preprocessed using the image processing library "OpenCV" and converted to an appropriate format for analysis.

[0813] Next, the server utilizes deep learning frameworks such as "TensorFlow" and "PyTorch" to identify objects in the image using models (e.g., VGG and ResNet). A corresponding tag is generated for each identified object, and these tags are stored in a database.

[0814] Furthermore, the server uses OCR technology such as "Tesseract" to extract text information from the image and generates additional tags based on the text. This process efficiently analyzes the characters in the image, allowing the necessary information to be obtained.

[0815] In particular, in this invention, the server is equipped with an emotion analysis engine that analyzes facial expressions and biometric information using "Haar Cascade" or "Keras," among others. Through this analysis, the user's emotional state is identified, and tags indicating that emotion are also stored in a database.

[0816] For example, if a user wants to search for photos of events related to happiness, the system will quickly extract and present images tagged with the emotion "happiness." This allows users to perform advanced image searches based on emotional context.

[0817] Examples of prompts for a generative AI model include: "Create a description of the algorithm for an image search system that takes user emotions into consideration. Include the specific process of emotion analysis performed by the emotion engine."

[0818] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0819] Step 1:

[0820] The user takes image data using their device and uploads it to the system. This involves clicking an "upload" button via an application on the device, sending the image data to the server over the internet. The input is the image data captured by the device, and the output is the image data sent to the server.

[0821] Step 2:

[0822] The server receives the image data and first performs image preprocessing using the "OpenCV" library. This preprocessing includes standardizing the image size and format. The input is the received image data, and the output is the preprocessed image data.

[0823] Step 3:

[0824] The server then applies an object recognition model (e.g., VGG or ResNet) using a deep learning framework such as TensorFlow or PyTorch. This identifies objects in the image and generates corresponding tags. The input is preprocessed image data, and the output is the object recognition result and the generated tags.

[0825] Step 4:

[0826] The server uses OCR technology such as "Tesseract" to extract text information from image data. This process obtains text from the image as a digital string and also generates tags. The input is image data, and the output is the extracted text information and additional tags.

[0827] Step 5:

[0828] The server uses an emotion analysis engine to analyze emotional information from images. Specifically, it uses "Haar Cascade" and "Keras" to analyze facial expressions and biometric information to identify the user's emotions. The input is image data, and the output is an emotion status and emotion tag.

[0829] Step 6:

[0830] The server saves all generated tags to the database. Here, object tags, text tags, and sentiment tags are all saved together in association. The input is the various tags that have been generated, and the output is the status of the data being saved to the database.

[0831] Step 7:

[0832] The user sends a text-based search query from their device to the server. Specifically, they enter keywords into the application's search bar and click the "Search" button. The input is the query information entered by the user, and the output is the query sent to the server.

[0833] Step 8:

[0834] The server extracts relevant image data from the database based on the received search query. It searches for highly relevant images from stored tags and formats the results. The input is the received search query, and the output is the search results.

[0835] Step 9:

[0836] The server displays the extracted search results on the user's device. Here, the results are formatted to be easily viewed by the user as an image list. The input is the search results formatted by the server, and the output is the image list displayed on the device screen.

[0837] (Application Example 2)

[0838] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0839] There is a challenge in providing appropriate information and recommendations based on images taken by users, taking into account their emotions and the context at the time. Conventional image search systems primarily rely on information retrieval based on physical characteristics and are unable to provide information that takes into account the user's emotions or psychological state.

[0840] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0841] In this invention, the server includes means for receiving image data captured by the user, means for identifying objects in the image data and generating tags, and means for generating tags based on emotional information using an emotion engine for analyzing the user's emotional state from the image data. This makes it possible to provide and recommend information that takes the user's emotional state into consideration.

[0842] "Means for receiving image data captured by a user" refers to a processing device that uses communication or an interface to acquire image information captured by a user from a terminal and incorporate it into the system.

[0843] "Methods using artificial intelligence" refer to methods that use computer-based algorithms to recognize and identify objects from input data, and utilize machine learning techniques.

[0844] "Means for generating tags corresponding to identified objects" refers to a device or software that performs a process of assigning textual information or identification information to objects recognized within image data.

[0845] "Means of using OCR technology" refers to a method or apparatus for optically recognizing text information within an image and converting it into digital data.

[0846] "Means of using an emotion engine" refers to an algorithm or program that analyzes the user's facial expressions and biometric information contained in an image to determine the user's emotional state.

[0847] "Means for storing generated tags in a database" refers to a database or storage device that is permanently maintained within the system for managing identified tags and sentiment information.

[0848] A "means for receiving text-based search queries and searching for related image data" refers to a search processing device that receives textual inquiries from users and searches for and extracts related image information based on those inquiries.

[0849] "Means of providing recommended information to users based on search results" refers to notifications or display devices that analyze data obtained through searches and present the user with the most optimal or relevant information.

[0850] "Means for presenting search results to the user" refers to an interface or output device for displaying or transmitting the results obtained by the system through a search to the user's terminal.

[0851] The system for implementing this invention begins with a user sending image data captured using a smart device to a server. The server analyzes the received images using an artificial intelligence module. This AI module is a model trained using machine learning platforms such as TensorFlow and PyTorch, and has the ability to identify objects in the image. Furthermore, it uses a Tesseract engine employing OCR (optical character recognition) technology to extract text information from the image and generate tags.

[0852] The image data is then fed into an emotion engine, which uses Microsoft Azure's facial recognition API and other tools to analyze the user's facial expressions and biometric indicators, thereby identifying their emotional state. This process generates emotion tags such as "happiness" and "surprise," which are then stored in a database.

[0853] When a user searches for specific information, the device sends a text-based search query to the server. The server quickly searches for relevant image data based on tags in its database and provides information related to a specific emotional state. For example, if the user's emotion is identified as "happy," advertisements related to entertainment and leisure will be displayed.

[0854] For example, if a user sends a photo of themselves happily with a friend, it could be tagged with "happiness," and then they could be offered movie ticket promotions or event information.

[0855] An example of a prompt message might be: "The user took a photo of themselves smiling with a friend. Sentiment analysis determined that they were feeling 'happy.' Please create an advertisement that matches this emotion."

[0856] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0857] Step 1:

[0858] The user takes an image with a smart device and sends the image data from the device to the server. The input is the image data taken by the user, and the output is the image data received by the server. In this process, the image is acquired using the smart device's camera API and the data is uploaded to the server via the HTTP protocol over an internet connection.

[0859] Step 2:

[0860] The server analyzes received image data using a machine learning model to identify objects within the image. The input is the received image data, and the output is a list of tags corresponding to the identified objects. Specifically, it uses generative AI models such as TensorFlow or PyTorch to analyze features in the image and generate tags.

[0861] Step 3:

[0862] The server extracts text information from the image using an OCR engine. The input is still the received image data, and the output is additional tag information based on the extracted text. An OCR engine such as Tesseract is used here to extract string data from the image and perform analysis.

[0863] Step 4:

[0864] The server uses an emotion engine to analyze facial expressions from images and identify the user's emotional state. The input is image data containing faces, and the output is tag information based on the user's emotions. For emotion analysis, the server uses Microsoft Azure's facial recognition API, among others, to infer emotions from the identified facial expression data.

[0865] Step 5:

[0866] The server saves all generated tags to a database. The input is the generated tag information, and the output is the registration status of the tags in the program's database. The tag information is recorded in a database system such as MySQL.

[0867] Step 6:

[0868] The user sends a text-based search query from their terminal to the server to find specific information. The input is the user's query text, and the output is that text data received by the server. The user enters the query through the application's search interface.

[0869] Step 7:

[0870] The server searches the database for appropriate image data based on the tags associated with the received query. The input is the search query and tag information from the database, and the output is a set of related image data. Here, an SQL query is used to filter the data for matching tags.

[0871] Step 8:

[0872] The server generates and presents recommended information to the user based on the search results. The input is image data and related information as search results, and the output is recommended advertisements and information displayed on the user's device. Specifically, it sets matching advertisement content based on the search results and displays the results on the user's interface.

[0873] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0874] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0875] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0876] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0877] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0878] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0879] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0880] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0881] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0882] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0883] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0884] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0885] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0886] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0887] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0888] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0889] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0890] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0891] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0892] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0893] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0894] The following is further disclosed regarding the embodiments described above.

[0895] (Claim 1)

[0896] Receives image data captured by the user.

[0897] means and

[0898] Artificial intelligence is used to identify objects in the aforementioned image data.

[0899] means and

[0900] Generate a tag corresponding to the identified object.

[0901] means and

[0902] OCR technology is used to extract text information from the image data, and additional tags are generated based on this text information.

[0903] means and

[0904] The generated tags are saved to the database.

[0905] means and

[0906] The system receives text-based search queries and searches for relevant image data based on the stored tags.

[0907] means and

[0908] Present search results to the user.

[0909] means and

[0910] A system that includes this.

[0911] (Claim 2)

[0912] The artificial intelligence uses a model trained based on a machine learning algorithm to identify objects in image data.

[0913] The system according to claim 1.

[0914] (Claim 3)

[0915] The aforementioned OCR technology uses an optical character recognition engine to identify and analyze the text contained in the image data.

[0916] The system according to claim 1.

[0917] "Example 1"

[0918] (Claim 1)

[0919] Receive image data captured from the user's information processing device.

[0920] means and

[0921] Machine intelligence is used to identify objects within the aforementioned image data.

[0922] means and

[0923] Generate relevant tags based on identified objects and extracted text.

[0924] means and

[0925] Optical character recognition technology is used to extract character information from the image data.

[0926] means and

[0927] The generated tag is stored in a storage device.

[0928] means and

[0929] It receives a text-based information request and searches for related image data based on the stored tags.

[0930] means and

[0931] Present the search results to the user.

[0932] means and

[0933] A system that includes this.

[0934] (Claim 2)

[0935] The aforementioned machine intelligence uses a model trained on a data classification algorithm to identify objects in image data.

[0936] The system according to claim 1.

[0937] (Claim 3)

[0938] The optical character recognition technology uses an optical character recognition processor to identify and analyze character information contained in the image data.

[0939] The system according to claim 1.

[0940] "Application Example 1"

[0941] (Claim 1)

[0942] Receives image data captured by the user.

[0943] means and

[0944] Artificial intelligence is used to identify objects in the aforementioned image data.

[0945] means and

[0946] Generate a tag corresponding to the identified object.

[0947] means and

[0948] OCR technology is used to extract text information from the image data, and additional tags are generated based on this text information.

[0949] means and

[0950] The generated tags are saved to the database.

[0951] means and

[0952] The system receives text-based search queries and searches for relevant image data based on the stored tags.

[0953] means and

[0954] Present search results to the user.

[0955] means and

[0956] A remote information display device or a mobile automated machine can be used as the shooting terminal.

[0957] means and

[0958] This system provides a way to efficiently manage object and machine information based on the generated tags.

[0959] means and

[0960] A system that includes this.

[0961] (Claim 2)

[0962] The artificial intelligence uses a model trained based on a machine learning algorithm to identify objects in image data.

[0963] The system according to claim 1.

[0964] (Claim 3)

[0965] The aforementioned OCR technology uses an optical character recognition engine to identify and analyze the text contained in the image data.

[0966] The system according to claim 1.

[0967] "Example 2 of combining an emotion engine"

[0968] (Claim 1)

[0969] Receives image data captured by the user.

[0970] means and

[0971] Artificial intelligence is used to identify objects in the aforementioned image data.

[0972] means and

[0973] Generate a tag corresponding to the identified object.

[0974] means and

[0975] OCR technology is used to extract text information from the image data, and additional tags are generated based on this text information.

[0976] means and

[0977] An emotion analysis engine is used to identify the user's emotional state from the image data and generate emotion-based tags.

[0978] means and

[0979] The generated tags are saved to the database.

[0980] means and

[0981] The system receives text-based search queries and searches for relevant image data based on the stored tags.

[0982] means and

[0983] Present search results to the user.

[0984] means and

[0985] A system that includes this.

[0986] (Claim 2)

[0987] The artificial intelligence uses a model trained based on a machine learning algorithm to identify objects in image data.

[0988] The system according to claim 1.

[0989] (Claim 3)

[0990] The aforementioned emotion analysis engine uses an emotion recognition algorithm to identify the user's emotions by analyzing facial expressions and biometric information contained in the image data.

[0991] The system according to claim 1.

[0992] "Application example 2 when combining with an emotional engine"

[0993] (Claim 1)

[0994] Receives image data captured by the user.

[0995] means and

[0996] Artificial intelligence is used to identify objects in the aforementioned image data.

[0997] means and

[0998] Generate a tag corresponding to the identified object.

[0999] means and

[1000] OCR technology is used to extract text information from the image data, and additional tags are generated based on this text information.

[1001] means and

[1002] Using an emotion engine to analyze the user's emotional state from the aforementioned image data, tags based on emotional information are generated.

[1003] means and

[1004] The generated tags are saved to the database.

[1005] means and

[1006] The system receives text-based search queries and searches for relevant image data based on the stored tags.

[1007] The means to provide users with recommendations based on these search results.

[1008] means and

[1009] Present search results to the user.

[1010] means and

[1011] A system that includes this.

[1012] (Claim 2)

[1013] The artificial intelligence uses a model trained based on a machine learning algorithm to identify objects in image data.

[1014] The system according to claim 1.

[1015] (Claim 3)

[1016] The aforementioned OCR technology uses an optical character recognition engine to identify and analyze the text contained in the image data.

[1017] The system according to claim 1. [Explanation of Symbols]

[1018] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of receiving image data captured by the user, Means for using artificial intelligence to identify objects in the aforementioned image data, Means for generating tags corresponding to identified objects, A means for extracting character information from the image data using OCR technology and generating additional tags based on this character information, Means for storing the generated tags in a database, A means for receiving a text-based search query and searching for related image data based on the stored tags, Means of presenting search results to users, A system that includes this.

2. The artificial intelligence uses a model trained based on a machine learning algorithm to identify objects in image data. The system according to claim 1.

3. The aforementioned OCR technology uses an optical character recognition engine to identify and analyze the text contained in the image data. The system according to claim 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A