System
The system addresses the challenge of searching diverse data formats by collecting, preprocessing, and analyzing data from multiple sources using a multimodal AI model, ensuring efficient access to relevant information and maintaining business continuity.
Patent Information
- Application Number
- JP2024122849
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-02-10
AI Technical Summary
Existing systems struggle to efficiently search for necessary information across vast amounts of internal documents and past communication logs, especially when documents are not properly handed over after employee transfers, and they have difficulty handling various data formats such as text and images, leading to reduced business efficiency.
A system that collects data from email servers, messaging platforms, and cloud storage, preprocesses it using NLP and OCR, builds a multimodal AI model for integrated analysis, and searches past communication logs to provide quick and efficient access to relevant materials.
Enables quick and efficient search for necessary information across various data formats, ensuring continuous business operations and knowledge accumulation even when documents are not properly transferred.
Smart Images

Figure 2026021167000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Creating an environment where necessary information can be quickly searched through vast amounts of internal documents and past communication logs is a challenge for many companies. In particular, when documents are not properly handed over after an employee leaves or is transferred, it becomes difficult to efficiently utilize past knowledge, which can significantly reduce business efficiency. Furthermore, conventional systems have difficulty searching for documents that contain not only text but also images and other data formats. This often results in a long time required to find the necessary documents. The present invention aims to solve these problems and improve business efficiency. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems with a system that includes: means for collecting data from email servers, messaging platforms, and cloud storage; means for analyzing and integrating text data, image data, and other data through preprocessing; means for building a multimodal AI model and analyzing data similarity; means for receiving search queries from users and identifying relevant data based on the search queries; means for searching past communication logs for sending relevant materials; and means for providing search results to user terminals. This system supports various data formats and provides an environment where necessary information can be quickly and efficiently searched for from past materials and communication logs. Furthermore, by making it possible to easily search and extract necessary materials even when materials are not properly handed over when employees resign or are transferred, it is possible to improve business efficiency and ensure the accumulation of knowledge.
[0006] A "mail server" is a server for sending and receiving e-mails, and is a system that manages users' mailboxes, stores e-mails, and implements transmission protocols.
[0007] A "messaging platform" is a communication tool for sending and receiving text messages in real time or asynchronously, such as Slack or Microsoft Teams.
[0008] "Cloud storage" refers to a remote server for storing data via the Internet, and examples include Google Drive and Dropbox.
[0009] "Preprocessing" is a step for converting collected data into a format that is easy to analyze, and includes data cleansing and formatting standardization.
[0010] "Text data" is information expressed as a string of characters, and is a data format obtained from the body of an email or a document file.
[0011] "Image data" refers to data that contains visual information, such as photographs and scanned images.
[0012] A "multimodal AI model" is an artificial intelligence model for integrated analysis of multiple different data formats (text, images, etc.).
[0013] "Natural language processing technology" is a computer program for analyzing, understanding, and generating human language, and is a technology for analyzing the meaning and grammatical structure of text data.
[0014] "Optical character recognition technology" is a technology that recognizes characters in images and converts them into digital text, and is used to extract text information from scanned documents and photographs.
[0015] A "search query" is a keyword or phrase that a user enters to search for information.
[0016] "Past communication logs" are records of communication, such as emails sent and received in the past and conversation history on messaging platforms.
[0017] A "user terminal" is a device operated by a user, such as a PC, smartphone, or tablet. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The present invention is a system for analyzing and searching data collected from an in-house mail server, a messaging platform, and cloud storage. The program processing of this system is explained below in natural language.
[0040] First, the server periodically connects to the company's internal mail server, messaging platforms (e.g., Slack and Microsoft Teams), and cloud storage (e.g., Google Drive and Dropbox) to collect all data, including email bodies and attachments, chat logs and uploaded files from messaging platforms, and various files from cloud storage.
[0041] Next, the server preprocesses the collected data. During this preprocessing step, the server analyzes the text data using natural language processing (NLP) technology, removes unnecessary information, and converts it into a format that is easier to analyze. For image data, the server also uses optical character recognition (OCR) technology to extract text information from the image. Furthermore, the server analyzes files such as PDFs and Word documents and converts them into text data.
[0042] The server then builds a multimodal AI model based on the preprocessed data (text data, image data, and other data formats). This AI model can comprehensively analyze text, images, and other data formats to extract data characteristics. This model is then trained to determine data similarity.
[0043] When a user searches for specific information on their device, they input a search query. For example, they input a query such as "Marketing strategy for new product Y." The device then sends this search query to a server. Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's internal database. The server then performs a search based on the query and ranks and extracts the most relevant materials.
[0044] The server also searches past related communication logs. For example, if the relevant document was sent by email, it searches the email body and attachments and extracts related Slack messages. This allows users to view the communication logs related to the document along with the related documents.
[0045] The server sends the search results to the terminal and displays them to the user. The user can check the displayed search results and select and view the necessary materials. For example, past reports on the marketing strategy for new product Y, presentation materials, related meeting notes, etc. can be displayed and viewed.
[0046] This system improves work efficiency by making it easy to search and extract necessary documents even when documents are not properly handed over after an employee leaves or is transferred. It also supports not only text but also images and other data formats, allowing for quick and efficient searches of various documents. This provides an environment that supports continuous business operations without compromising internal knowledge.
[0047] The processing flow will be explained below.
[0048] Step 1: Collect data
[0049] The server collects data from the company's email server, messaging platform, and cloud storage. From the email server, it collects email text, attachments, and metadata. From the messaging platform, it obtains chat logs and uploaded files. It also collects various document files (PDF, Word, image files, etc.) from cloud storage.
[0050] Step 2: Preprocessing the data
[0051] The server preprocesses the collected data by format. For text data, it analyzes it using natural language processing (NLP) technology and removes noise. For image data, it uses optical character recognition (OCR) technology to extract text information from the image. PDFs and Word documents are also converted into text data and their contents are analyzed.
[0052] Step 3: Building an AI model
[0053] The server then builds a multimodal AI model based on the preprocessed data. This model is trained to comprehensively analyze text, images, and other data formats and extract data characteristics, resulting in a model capable of similarity judgment.
[0054] Step 4: Receiving a search query
[0055] The terminal receives a search query from the user. For example, the user inputs a search query such as "Marketing strategy for new product Y." The terminal then sends this query to the server.
[0056] Step 5: Processing the search query
[0057] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's database. It performs a query-based search and ranks and extracts highly relevant materials. It performs an integrated search of materials, including not only text data but also different data formats such as images and PDFs.
[0058] Step 6: Search past communication logs
[0059] The server also searches past communication logs (emails and messages on messaging platforms) for related materials, extracting relevant email text and Slack messages, and searching for information from the communication logs.
[0060] Step 7: Serving search results
[0061] The server sends the search results and information extracted from past communication logs to the terminal. The terminal displays this data on a user interface, allowing the user to check each piece of information. For example, a report on the marketing strategy for new product Y and related communication records may be displayed.
[0062] Step 8: User validation and utilization
[0063] Users can check the search results displayed on their devices, select and view related materials as needed, and download or print any information they deem necessary for further use.
[0064] In this way, the system quickly searches for necessary information from vast amounts of internal company data, improving business efficiency.
[0065] Example 1
[0066] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0067] With conventional systems, it was difficult to efficiently collect and analyze data dispersed across internal emails, messaging platforms, and cloud storage, making it impossible to quickly search and retrieve the information needed. Furthermore, there was a lack of technology for comprehensively analyzing different data formats (text data, image data, PDF files, etc.), which led to problems that reduced operational efficiency. Furthermore, there was a need for a system that could include past communication logs in search results, allowing users to check the history of related communications.
[0068] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0069] In this invention, the server includes: means for collecting data from mail servers, messaging platforms, and cloud storage; means for analyzing text data, image data, and other data in preprocessing and integrating them; means for analyzing text data using natural language processing technology; means for extracting text information from image data using optical character recognition technology; means for building a multimodal AI model and analyzing data similarity; means for receiving search queries from users and identifying relevant data based on the search queries; means for searching past communication logs for sending relevant materials; and means for providing search results to user terminals. This enables centralized collection and analysis of distributed data within a company, allowing users to quickly and efficiently search for and obtain the information they need. It also enables integrated analysis of different data formats and provides comprehensive search results that include related communication history.
[0070] A "mail server" is a server that manages the sending and receiving of emails. Its role is to provide mailboxes, store sent emails, and forward received emails to users.
[0071] A "messaging platform" refers to a communication tool that provides functions such as text messaging, file sharing, voice calls, and video calls. Typical examples include online chat services and business communication tools.
[0072] "Cloud storage" refers to remote data storage services available over the Internet, allowing users to store data in the cloud and access, share, and manage it as needed.
[0073] "Data collection tools" refers to the technologies and processes used to extract and centralize data from different sources. This tool allows data from various platforms to be managed in a unified manner.
[0074] "Preprocessing" refers to a series of operations performed to convert raw data into a suitable format for data analysis and machine learning, including text segmentation, noise removal, and format conversion.
[0075] "Natural language processing technology" refers to technology that enables computers to understand, analyze, and generate human language, including text analysis, machine translation, and sentiment analysis.
[0076] Optical character recognition technology is a technology that analyzes characters in images captured by a scanner or camera and converts them into text data. It is used to digitize printed documents and handwritten characters.
[0077] A "multimodal AI model" is an artificial intelligence model that can integrate and analyze data in different formats (text, images, audio, etc.). Combining multiple data sources enables more accurate analysis.
[0078] A "search query" is a question or keyword that a user enters to search for specific information. It is used to refer to information in a search engine or database.
[0079] "Means for identifying relevant data" refers to techniques and processes for finding and extracting relevant data based on a search query from a user.
[0080] "Past communication logs" refers to records of past communications, such as email sending and receiving history and chat history on messaging platforms.
[0081] "Search Results" means a list of relevant data or information based on a user's search query, generated by a search engine or database and provided to the user.
[0082] "User terminal" refers to a device such as a computer, smartphone, or tablet that is directly operated by a user. It is a device that connects to a server via the Internet or an internal company network and allows users to use services.
[0083] The present invention is a system for collecting data from an in-house mail server, a messaging platform, and cloud storage, and analyzing and searching the data. Specific embodiments of this system are described below.
[0084] Hardware and Software Configuration
[0085] The server includes the following main software and hardware:
[0086] A scheduler (such as cron) to periodically run scripts (such as Python) for data collection
[0087] IMAP protocol library for connecting to mail servers
[0088] API client libraries for messaging platforms and cloud storage (e.g., Slack API, Google Drive API)
[0089] Libraries for using natural language processing technologies (NLTK and spaCy)
[0090] Library for using optical character recognition technology (Tesseract-OCR)
[0091] Libraries for building and training multimodal AI models (TensorFlow and PyTorch)
[0092] The terminals include computers, smartphones, tablets, etc. that are directly operated by the user and allow the user to access the system's search interface through a browser.
[0093] Data collection
[0094] The server periodically runs a data collection script that connects to the company's mail server, messaging platform, and cloud storage, collecting emails, chat logs, uploaded files, and various documents. The server then stores this data in local storage.
[0095] Data Preprocessing
[0096] The server preprocesses the collected data. Specifically, it uses natural language processing technology to analyze the text data and remove unnecessary information. Examples include removing stop words and stemming. For image data, it uses optical character recognition technology to extract text information, and converts PDFs and Word documents into text data. This integrates all data into an analyzable format.
[0097] Building a multimodal AI model
[0098] The server integrates preprocessed text data, image data, and other data formats into a single dataset. A multimodal AI model is built using TensorFlow and PyTorch to learn the characteristics of the data. This model can integrate and analyze different data formats, enabling more accurate searches.
[0099] User search and results display
[0100] A user accesses the system's search interface from a browser on their device and enters a search query, such as "marketing strategy for new product Y." The device then sends this search query to the server.
[0101] The server analyzes the received search query and uses a multimodal AI model to identify related materials. It also searches related past communication logs (including email and chat history) and displays them in an integrated manner, making it easier for users to grasp related information.
[0102] Providing search results
[0103] The server provides search results in JSON format to the terminal, which then displays them in HTML format, allowing the user to quickly and efficiently obtain the information they need. For example, past reports on marketing strategies for new product Y, related presentation materials, and meeting notes can be displayed.
[0104] Examples and prompts
[0105] For example, when searching using the prompt phrase "Marketing strategy for new product Y," the system quickly searches for relevant internal documents and past communication history and displays them to the user, allowing the user to carry out their work more efficiently.
[0106] In this way, this system provides specific methods and technologies for efficiently collecting and analyzing huge amounts of data and quickly responding to user search requests.
[0107] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0108] Program processing flow
[0109] Step 1: Collect data
[0110] The server periodically runs data collection scripts to connect to the company's mail server, messaging platform, and cloud storage.
[0111] Input: Connection and authentication information configured for the server
[0112] How it works: It connects to mail servers using the IMAP protocol, and to messaging platforms and cloud storage using the APIs of each service.
[0113] Output: Save retrieved emails, chat logs, and various documents to local storage.
[0114] Step 2: Preprocessing the data
[0115] The server pre-processes the collected data.
[0116] Input: Collected emails, chat logs, document files
[0117] How it works: It uses natural language processing (NLP) techniques to analyze text data and remove unnecessary information. Specifically, it performs text segmentation, stop word removal, and stemming. It also uses Tesseract-OCR to perform character recognition on image data, and PDF and Word documents are converted to text using PDFMiner and python-docx.
[0118] Output: Preprocessed text data, image data, and other data are integrated and saved.
[0119] Step 3: Building a multimodal AI model
[0120] The server uses the integrated data to build a multimodal AI model.
[0121] Input: Preprocessed dataset
[0122] How it works: We use TensorFlow and PyTorch to build models and train them to learn data features, specifically by integrating text, images, and other data formats and determining data similarity.
[0123] Output: A trained multimodal AI model
[0124] Step 4: Receiving and parsing the search query
[0125] A user accesses the search interface from a browser on the terminal and enters a search query.
[0126] Input: The search query entered by the user (e.g., "Marketing strategy for new product Y")
[0127] Operation: The terminal sends this query to the server, which analyzes the received query.
[0128] Output: Parsed search query
[0129] Step 5: Search and identify relevant data
[0130] The server uses the parsed search query to search for relevant data.
[0131] Input: Parsed search query, trained multimodal AI model
[0132] How it works: Uses a multimodal AI model to identify and rank relevant materials and data based on a search query.
[0133] Output: A list of related data
[0134] Step 6: Search past communication logs
[0135] The server searches past communication logs based on the associated data.
[0136] Input: List of related data
[0137] How it works: Uses email and messaging platform APIs to search and extract relevant past communication logs.
[0138] Output: Related past communication logs
[0139] Step 7: Serving search results
[0140] The server integrates the relevant data with past communication logs to generate search results.
[0141] Input: List of related data, past communication logs
[0142] What it does: Generates a unified search result in JSON format and sends it to the device.
[0143] Output: JSON data of search results
[0144] Step 8: Viewing search results
[0145] The terminal displays the received search results to the user.
[0146] Input: JSON data of search results sent from the server
[0147] What it does: Converts the results to HTML and displays them in the browser, allowing the user to view the results.
[0148] Output: Search results displayed to the user
[0149] This series of processes allows the user to efficiently search for and view the information they need.
[0150] (Application example 1)
[0151] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0152] While conventional in-house data search systems can collect and preprocess data from email servers and messaging platforms, they face the problem of making it difficult for store staff to quickly and efficiently obtain the information they need. In particular, there is a need for a way to quickly obtain information in real time, such as inventory information and past sales history, while serving customers in stores. Furthermore, there is a lack of systems that can handle multiple data formats and perform integrated analysis. Therefore, an effective solution is needed to improve the work efficiency of store staff.
[0153] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0154] In this invention, the server includes means for collecting data from mail servers, messaging platforms, and cloud storage, means for analyzing text data, image data, and other data in preprocessing and integrating them, means for building a multimodal AI model and analyzing data similarity, means for receiving search queries from users and identifying related data based on the search queries, means for searching past communication logs for sending related materials, means for providing search results to user terminals, and means for in-store staff to send queries by voice input and display search results through smart glasses, enabling store staff to obtain information in real time and respond to customers quickly and efficiently.
[0155] A "mail server" is a computer server that manages and stores the sending and receiving of email.
[0156] A "messaging platform" is an online communication tool for sending and receiving text messages and files in real time.
[0157] "Cloud storage" is an online storage service for storing, managing, and accessing data via the Internet.
[0158] "Preprocessing" refers to a data processing step to convert collected data into a format that is easy to analyze.
[0159] "Text data" is information expressed as characters or sentences.
[0160] "Image data" is information expressed in an image format.
[0161] "Integration" means combining multiple different data formats and information into one format and making them consistent.
[0162] A "multimodal AI model" is an artificial intelligence model that performs integrated analysis of multiple data formats (text, images, etc.).
[0163] "Data similarity" refers to common characteristics or patterns between different data.
[0164] A "search query" is a search question or keyword entered by a user to search for specific information.
[0165] "Related Materials" are useful information or documents identified based on the search query.
[0166] "Past communication logs" are records of messages and emails previously exchanged.
[0167] A "user terminal" is an electronic device that a user operates and receives information from, such as a smartphone or tablet.
[0168] A "brick and mortar store" is a sales facility that has a physical presence.
[0169] "Voice input" is a method of inputting characters and commands using voice.
[0170] "Smart glasses" are a wearable eyeglass-type device equipped with a display function.
[0171] The present invention is a system that enables in-store staff to quickly and efficiently obtain necessary information while serving customers. The system collects data from mail servers, messaging platforms, and cloud storage, and analyzes them in an integrated manner to provide relevant data based on user search queries.
[0172] 1. Program Overview
[0173] The server operates programs to achieve the following functions:
[0174] Data Collection: Periodically collect data from mail servers, messaging platforms, and cloud storage.
[0175] Pre-processing: Analyze collected data using NLP techniques (e.g., spaCy) and OCR techniques (e.g., Tesseract) to extract and convert text and image data.
[0176] Building multimodal AI models: Multimodal AI models can be trained based on text data, image data, and other data formats to determine data similarity.
[0177] Search query processing: Receives and analyzes search queries voice-entered by store staff through smart glasses.
[0178] Sending search results: Search results for related materials and past communication logs are displayed on the smart glasses.
[0179] 2. Hardware and Software
[0180] Hardware used: Server, smart glasses, user devices (e.g., smartphones and tablets).
[0181] Software used: NLP engine (spaCy), OCR engine (Tesseract), multimodal AI model (OpenAI's GPT-4).
[0182] 3. System operation explanation
[0183] The server first collects the necessary data from various data sources. The collected data undergoes a preprocessing step to convert it into a format that is easy to analyze. During this process, NLP and OCR technologies are used to extract text data, and image data is also handled. After preprocessing is complete, the data is analyzed comprehensively using a multimodal AI model to extract data characteristics and similarities.
[0184] The user, a store staff member, voice-inputs queries such as customer questions or inventory checks. For example, prompts such as "Check inventory, tell me the availability of product A" or "Search customer X's purchase history" are sent through the smart glasses. The server receives the query, searches relevant data using a multimodal AI model, and ranks and extracts the results. These results are then provided to the smart glasses, allowing the user to view the information in real time.
[0185] 4. Specific Examples
[0186] For example, if a store staff member sends a voice prompt such as "Tell me the purchase history of customer Yamada Taro for the past three months," the system will quickly collect relevant past communication logs, inventory information, and sales data and display them on the smart glasses' display. It also works in the same way with prompts such as "Show me the marketing materials for new product A that will be on sale next weekend," supporting efficient customer service.
[0187] The present invention allows staff in physical stores to obtain information in real time, enabling them to respond to customers quickly and accurately.
[0188] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0189] Step 1:
[0190] Data collection
[0191] The server periodically connects to mail servers, messaging platforms, and cloud storage to collect data, including email bodies and attachments, chat logs and uploaded files from messaging platforms, and various files from cloud storage.
[0192] Input: Each data source (mail server, messaging platform, cloud storage)
[0193] Output: Raw data collected (emails, chat logs, files)
[0194] Specifically, the server accesses these data sources using APIs or dedicated protocols, retrieves the latest data, and stores it.
[0195] Step 2:
[0196] Data Preprocessing
[0197] The server preprocesses the collected data. During this preprocessing step, the server analyzes the text data using natural language processing (NLP) technology, removes unnecessary information, and converts it into a format that is easier to analyze. For image data, the server also uses optical character recognition (OCR) technology to extract text information from the image. Furthermore, the server analyzes files such as PDFs and Word documents and converts them into text data.
[0198] Input: Raw data collected
[0199] Output: Preprocessed data (text format)
[0200] Specifically, the system uses an NLP engine (e.g., spaCy) to analyze text and remove unnecessary parts, and an OCR engine (e.g., Tesseract) to extract text information from the image.
[0201] Step 3:
[0202] Building a multimodal AI model
[0203] The server uses the preprocessed data to build a multimodal AI model that can comprehensively analyze text, images, and other data formats, and trains the model to determine data similarity.
[0204] Input: Preprocessed data
[0205] Output: A trained multimodal AI model
[0206] Specifically, it feeds data to AI models (such as OpenAI's GPT-4) to train them, and also continuously updates and improves them.
[0207] Step 4:
[0208] Receiving a search query
[0209] The user (store staff) sends a search query by voice input through the smart glasses, for example, "Check inventory, tell me the stock status of product X." The device converts this voice data into text format and sends it to the server.
[0210] Input: Voice query
[0211] Output: Text query
[0212] Specifically, it converts voice input into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API).
[0213] Step 5:
[0214] Finding related data
[0215] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's database, ranking and extracting the most relevant materials.
[0216] Input: Text query, trained multimodal AI model, in-house database
[0217] Output: Search results (list of related materials)
[0218] Specifically, the server inputs a query into the model, searches the database for materials that match the query, and retrieves the results.
[0219] Step 6:
[0220] Submitting and viewing search results
[0221] The server sends the search results to the user's device (smart glasses), and the user can view the displayed search results, such as "Report on marketing strategy for new product X" or "Related meeting notes."
[0222] Input: Search results
[0223] Output: Results displayed on smart glasses
[0224] Specifically, the server formats the search results and transmits them in a format suitable for the smart glasses' display.
[0225] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0226] This invention is a system that collects data from an in-house mail server, messaging platform, and cloud storage, and combines it with an emotion engine that recognizes the user's emotions to perform more effective document searches. Below, the program processing of this system is explained in natural language.
[0227] First, the server periodically connects to the company's internal mail server, messaging platforms (e.g., Slack and Microsoft Teams), and cloud storage (e.g., Google Drive and Dropbox) to collect all data, including email bodies and attachments, chat logs and uploaded files from messaging platforms, and various files from cloud storage.
[0228] Next, the server preprocesses the collected data. During this preprocessing step, the server analyzes the text data using natural language processing (NLP) technology, removes unnecessary information, and converts it into a format that is easier to analyze. For image data, the server also uses optical character recognition (OCR) technology to extract text information from the image. Furthermore, the server analyzes files such as PDFs and Word documents and converts them into text data.
[0229] The server then builds a multimodal AI model based on the preprocessed data (text data, image data, and other data formats). This AI model can comprehensively analyze text, images, and other data formats to extract data characteristics. This model is then trained to determine data similarity.
[0230] When a user searches for specific information on a device, they input a search query. For example, they input a query such as "Marketing strategy for new product Y." The device then sends this search query to the server. At this stage, the device can also use an emotion engine to detect the user's emotions. The emotion engine recognizes the user's facial expressions and voice and determines the user's current emotional state.
[0231] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's internal database. It then performs a search based on the query and ranks and extracts the most relevant materials. It performs an integrated search of materials that include not only text data but also different data formats such as images and PDFs. Furthermore, it adjusts the ranking of search results by taking into account the user's emotional state detected by the emotion engine. For example, if the user is feeling stressed, it will prioritize displaying materials that are more concise and easy to understand.
[0232] The server also searches past related communication logs. For example, if the relevant document was sent by email, it searches the email body and attachments and extracts related Slack messages. This allows users to view the communication logs related to the document along with the related documents.
[0233] The server sends the search results to the terminal and displays them to the user. The user can check the displayed search results and select and view the necessary materials. For example, past reports on the marketing strategy for new product Y, presentation materials, related meeting notes, etc. are displayed and can be reviewed. In addition, an emotion engine is used to provide information tailored to the user's emotions, improving the user experience.
[0234] This system improves work efficiency by making it easy to search and retrieve necessary documents even when documents are not properly handed over after an employee leaves or is transferred. It also supports not only text but also images and other data formats, allowing for quick and efficient searches of various documents. Furthermore, by utilizing an emotion engine, it is possible to provide search results that reflect the user's emotional state, improving user satisfaction.
[0235] The processing flow will be explained below.
[0236] Step 1: Collect data
[0237] The server periodically connects to the company's email server, messaging platform, and cloud storage to collect data. Specifically, it collects email text, attachments, and metadata from the email server, chat logs and uploaded files from the messaging platform, and various document files (PDF, Word, image files, etc.) from the cloud storage.
[0238] Step 2: Preprocessing the data
[0239] The server preprocesses the collected data by format. For text data, it analyzes it using natural language processing (NLP) technology and removes noise. For image data, it uses optical character recognition (OCR) technology to extract text information from the image. PDFs and Word documents are also converted into text data and their contents are analyzed.
[0240] Step 3: Building an AI model
[0241] The server then builds a multimodal AI model based on the preprocessed data. This model is trained to comprehensively analyze text, images, and other data formats and extract data characteristics, resulting in a model capable of similarity judgment.
[0242] Step 4: Receiving a search query
[0243] The device receives a search query from the user. For example, the user may enter a search query such as "Marketing strategy for new product Y." The device sends this query to the server and simultaneously activates an emotion engine to detect the user's emotional state.
[0244] Step 5: Detecting Emotional State
[0245] The device uses an emotion engine to analyze the user's emotional state. Specifically, it recognizes emotions from the user's facial expressions and voice and determines the user's current emotional state. For example, it detects whether the user is feeling stressed or relaxed.
[0246] Step 6: Processing the search query
[0247] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's database. It performs a query-based search and ranks and extracts highly relevant materials. It performs an integrated search of materials, including not only text data but also different data formats such as images and PDFs.
[0248] Step 7: Adjust your results based on your emotional state
[0249] The server adjusts the ranking of search results based on the user's emotional state as detected by the emotion engine. For example, if the user is feeling stressed, it will prioritize more concise and easy-to-understand materials. On the other hand, if the user is relaxed, it will provide more detailed and specialized materials.
[0250] Step 8: Search past communication logs
[0251] The server also searches past communication logs (emails and messages on messaging platforms) for related materials, extracting relevant email text, Slack messages, etc., and searching for information from the communication logs.
[0252] Step 9: Serving search results
[0253] The server sends the search results and information extracted from past communication logs to the terminal. The terminal displays this data on a user interface, allowing the user to check each piece of information. For example, a report on the marketing strategy for new product Y and related communication records may be displayed.
[0254] Step 10: User validation and utilization
[0255] Users can check the search results displayed on their devices, select and view related materials as needed, and download or print any information they deem necessary for further use.
[0256] In this way, this system aims to improve work efficiency and user satisfaction by quickly searching for necessary information from vast amounts of internal company data and providing information that corresponds to the user's emotional state.
[0257] Example 2
[0258] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0259] Conventional document search systems struggled to effectively collect and integrate data from internal email servers, messaging platforms, and cloud storage, and lacked a way to appropriately adjust search results based on the user's emotional state. This resulted in problems such as an inability to respond quickly and accurately to user requests, leading to reduced work efficiency. Furthermore, it was difficult to comprehensively search documents in different data formats, making it difficult to find the information needed when document handover was insufficient.
[0260] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0261] In this invention, the server includes: means for collecting data from mail servers, messaging platforms, and cloud storages; means for analyzing text data, image data, and other data using natural language processing technology and optical character recognition technology in preprocessing and integrating them; means for building a multimodal AI model including text, image, and other data formats and analyzing the similarity of the data; means for receiving a user's search query and identifying related data based on the search query using an emotion engine that detects the user's emotional state; means for searching past communication logs to send related materials; and means for providing the search results to the user terminal.
[0262] This makes it possible to centrally collect and integrate data from various data sources within the company and provide appropriate search results based on user sentiment. Users can quickly find the documents they need, improving work efficiency. In addition, by searching documents in different data formats in an integrated manner, related information can be easily obtained even if document handover is insufficient.
[0263] A "mail server" is a server that manages the sending and receiving of email within an organization or on the Internet.
[0264] A "messaging platform" is software or a service for sending and receiving messages in real time.
[0265] "Cloud storage" is an online service for storing and managing data over the Internet.
[0266] "Natural language processing technology" is a technology that analyzes text data and extracts linguistic features.
[0267] "Optical character recognition technology" is a technology that extracts text information from image data.
[0268] A "multimodal AI model" is an artificial intelligence model that comprehensively analyzes text, images, and other data formats to extract features.
[0269] "User emotional state" refers to the emotional state a user is experiencing when performing a search, and includes techniques for detecting this.
[0270] An "emotion engine" is a device or software that analyzes a user's facial expressions and voice to detect their emotional state.
[0271] A "search query" is text data indicating a search request that a user inputs into a search system.
[0272] "Communication logs" are data containing past communication history across email and messaging platforms.
[0273] The present invention is a system that collects data from an internal mail server, a messaging platform, and cloud storage, and combines it with an emotion engine that recognizes the user's emotional state to perform more effective document search. The program processing of this system is described below.
[0274] First, the server periodically connects to the company's email server, messaging platform, and cloud storage to collect all data. Specifically, Microsoft Exchange and Gmail are used as email servers, Slack and Microsoft Teams are used as messaging platforms, and Google Drive and Dropbox are used as cloud storage. The server uses APIs to collect email bodies and attachments, chat logs, uploaded files, and various files in cloud storage.
[0275] The server then preprocesses the collected data. This preprocessing step uses natural language processing (NLP) and optical character recognition (OCR) techniques. For text data, NLTK and spaCy are used to tokenize and remove stop words. For image data, Tesseract OCR is used to extract text information from images. For PDF and Word documents, Apache Tika is used to convert them to text data.
[0276] The server then builds a multimodal AI model based on the preprocessed data. This model is created using TensorFlow and PyTorch to comprehensively analyze text, images, and other data formats and extract data characteristics. Specifically, it uses a convolutional neural network (CNN) to extract image features, and a recurrent neural network (RNN) or BERT model to extract text features.
[0277] Next, the user inputs a search query from the terminal. For example, the query "Marketing strategy for new product Y" is input. This search query is sent as is from the terminal to the server as a character string.
[0278] The device is equipped with an emotion engine that detects the user's emotional state, for example by capturing and analyzing facial expressions with a camera. This analysis is performed using the Face API from Microsoft Azure's Cognitive Services. The emotion engine detects the user's emotional state and transmits that information along with a query to the server.
[0279] Based on the received search query and emotional information, the server uses a multimodal AI model to identify relevant materials from the company's database. It then assesses data similarity and ranks and extracts highly relevant materials. It also adjusts the ranking of results according to the user's emotional state. If the user is feeling stressed, the algorithm prioritizes displaying more concise and easy-to-understand materials.
[0280] The server also searches past communication logs for related materials. For example, if the material was sent by email, it extracts the email body, attachments, and chat logs from related messaging platforms. Finally, the server sends the search results and related communication logs to the device.
[0281] Users can select and view the necessary materials from the results displayed on their device. For example, they can check past reports, presentation materials, meeting notes, etc. related to "Marketing Strategy for New Product Y."
[0282] As a concrete example, here is a prompt example for a generative AI model when searching for "marketing strategy for new product Y":
[0283] "Search your internal database for materials related to the marketing strategy for new product Y. If the user is feeling stressed, prioritize materials that are more concise and easy to understand."
[0284] This system makes it possible to quickly and efficiently search and retrieve necessary documents even when documents are not fully handed over when an employee resigns or is transferred.In addition, by providing search results that respond to the user's emotions, it is expected that the user experience will improve and business efficiency will increase.
[0285] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0286] Step 1:
[0287] The server periodically connects to the company's mail server, messaging platform, and cloud storage to collect data. In this step, it retrieves email bodies and attachments from the mail server, chat logs and uploaded files from the messaging platform, and various files from the cloud storage. The specific input is a request using the API of each data source, and the output is a collection of collected data.
[0288] Step 2:
[0289] The server preprocesses the collected data. It uses natural language processing (NLP) and optical character recognition (OCR) techniques to tokenize text data, remove stop words, extract text information from image data, and convert PDFs and Word documents into text data. Specific operations include text analysis using NLTK and spaCy, image analysis using Tesseract OCR, and document analysis using Apache Tika. The input is the raw data collected in step 1, and the output is the preprocessed, clean data.
[0290] Step 3:
[0291] The server builds a multimodal AI model based on the preprocessed data. It uses TensorFlow and PyTorch to comprehensively analyze text, images, and other data formats and extract features. Specifically, it uses a Convolutional Neural Network (CNN) to extract image features, and a Recurrent Neural Network (RNN) or BERT model to extract text data features. The input is various preprocessed data, and the output is a trained AI model.
[0292] Step 4:
[0293] A user inputs a query to search for specific information from a terminal. For example, the query might be "Marketing strategy for new product Y." The query is sent directly from the terminal to the server in the form of a string. The input is the query entered by the user, and the output is a search request sent to the server.
[0294] Step 5:
[0295] The device uses an emotion engine to detect the user's emotional state. Specifically, it captures the user's facial expressions with a camera and analyzes them using the Face API of Microsoft Azure's Cognitive Services. The input is an image of the user's facial expression, and the output is the detected emotional state data.
[0296] Step 6:
[0297] The server uses a multimodal AI model to search for relevant materials from the company's database based on the received search query and emotional information. It ranks the search results and adjusts the ranking based on the emotion engine's evaluation. Specifically, if the user is feeling stressed, the server uses an algorithm that prioritizes displaying concise and easy-to-understand materials. The input is the search query and emotional information, and the output is ranked search results.
[0298] Step 7:
[0299] The server also searches past communication logs for related materials. For example, if the material was sent by email, it extracts the email body, attachments, and chat logs from related messaging platforms. The input is the search results, and the output is the related communication logs.
[0300] Step 8:
[0301] The server sends the search results and related communication logs to the terminal. The user can select and view the necessary materials from the results displayed on the terminal. The input is the ranked search results and related communication logs, and the output is the information displayed on the terminal in a user-accessible form.
[0302] Through these steps, users can quickly and efficiently search and extract the information they need.
[0303] (Application example 2)
[0304] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0305] Providing timely and appropriate information is essential to improving work efficiency in factories and reducing worker stress. However, conventional information search systems were unable to provide information tailored to the emotional state of workers, making it difficult to provide efficient work support. Furthermore, extracting information from a wide variety of data formats and past communication logs was time-consuming, making it difficult to provide practical information.
[0306] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0307] In this invention, the server includes: means for collecting data from mail servers, messaging platforms, and cloud storage; means for analyzing and integrating text data, image data, and other data in preprocessing; means for building a multimodal AI model and analyzing data similarity; means for receiving search queries from users and identifying relevant data based on the search queries; means for analyzing the emotional state of workers and adjusting search results based on the results; means for searching past communication logs for sending relevant materials; and means for providing search results to user terminals. This enables timely and appropriate information provision that takes emotional states into account. Furthermore, the system can efficiently extract information from a variety of data formats and past communication logs, thereby providing efficient support for workers.
[0308] A "mail server" is a server for sending, receiving, and storing e-mail.
[0309] A "messaging platform" is a system for exchanging messages between users using text, voice, video, etc.
[0310] "Cloud storage" refers to online storage services for storing and accessing data over the Internet.
[0311] "Means for collecting data" refers to the mechanisms for obtaining the necessary data from mail servers, messaging platforms, and cloud storage.
[0312] "Preprocessing" is the process of converting collected data into a format that is easy to analyze and removing unnecessary information.
[0313] "Text data" is a data format that includes character information.
[0314] "Image data" is a data format that contains visual information.
[0315] A "multimodal AI model" is an artificial intelligence model for integrated analysis of different data formats (text, images, etc.).
[0316] A "user search query" is a search request entered by a user to locate specific information.
[0317] The "emotional state" is a state that indicates the type and intensity of the emotion that the user is currently feeling.
[0318] "Related materials" are materials that contain information that is useful to the user and that are identified based on the search query.
[0319] A "communication log" is a record of messages and files sent and received in the past, as well as related information.
[0320] "Search Results" means a list of relevant materials identified based on a search query.
[0321] A "user terminal" is a device that a user operates directly to check information.
[0322] To realize this application example, the system program is configured as follows.
[0323] The server collects data from email servers, messaging platforms, and cloud storage. Specifically, it uses the Gmail API, Microsoft Teams API, Google Drive API, and other sources to obtain the necessary data. The obtained data is in the form of text, image data, and other data formats, and natural language processing (NLP) and optical character recognition (OCR) technologies are used to preprocess it. After preprocessing, the data is analyzed using Hugging Face's transformers library to build a multimodal AI model. The multimodal AI model comprehensively analyzes text, image data, and other data formats.
[0324] The server also receives search queries from users and identifies relevant data based on the search queries. A user can input a search query such as "How to troubleshoot error code E123 on machine X" using a device (e.g., a smartphone or smart glasses). The server also analyzes the user's emotional state in real time using a device equipped with an emotion engine (e.g., smart glasses with a built-in camera and microphone that analyzes facial expressions and voice). If the user is feeling stressed, the emotion engine will prioritize providing more concise and easy-to-understand materials.
[0325] The search engine searches the database based on the query and also searches past related communication logs. For example, if related documents were sent via email or messaging platforms, the contents of those documents will also be searched. This allows for comprehensive retrieval of past communications and related documents. Search results are ranked according to the user's emotional state, with the most relevant information being displayed first.
[0326] The terminal provides the search results to the user, who can then check and view the necessary materials. This improves work efficiency and reduces worker stress. This system makes it possible to efficiently extract information from a wide range of data formats, significantly improving the factory work environment.
[0327] As a concrete example, consider a factory worker troubleshooting a new machine. If the worker types the query "How to troubleshoot error code E123 on machine X" into the smart glasses and is stressed, the application will prioritize short, easy-to-understand instructions.
[0328] Example prompt sentence:
[0329] "Find past troubleshooting documents for error code E123 on machine X. Workers are stressed right now, so please prioritize concise documents."
[0330] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0331] Step 1:
[0332] The server collects data from email servers, messaging platforms, and cloud storage. As input, it uses the Gmail API, Microsoft Teams API, Google Drive API, etc. to obtain the target email, message, and file data. The output is the collected text data, image data, and other data formats.
[0333] Step 2:
[0334] The server preprocesses the collected data. Specifically, it uses natural language processing (NLP) technology to analyze the text data, remove unnecessary information, and convert it into a format that is easier to analyze. It also uses optical character recognition (OCR) technology to extract text information from image data. It also analyzes files such as PDFs and Word documents and converts them into text data. The input is the data collected in step 1, and the output is the preprocessed data.
[0335] Step 3:
[0336] The server builds a multimodal AI model based on various preprocessed data (text data, image data, and other data formats). Using Hugging Face's transformers library, it trains an AI model that can comprehensively analyze different data formats. The input is the preprocessed data output from Step 2, and the output is an AI model that can analyze data similarities.
[0337] Step 4:
[0338] A user uses a device (e.g., a smartphone, smart glasses) to input a search query, such as "how to troubleshoot error code E123 on machine X." The input is the user's query, and the output is the query data.
[0339] Step 5:
[0340] The device sends the search query to the server. At the same time, the device uses an emotion engine to analyze the user's emotional state. For example, it uses the camera and microphone built into the smart glasses to analyze the user's facial expressions and voice to determine their current emotional state. The input is the query from step 4 and the user's emotional data, and the output is composite data including the query and the emotional state.
[0341] Step 6:
[0342] The server searches the database based on the search query, using a multimodal AI model to identify relevant information and also searches past communication logs to retrieve relevant materials. The input is the composite data from step 5, and the output is a list of relevant materials.
[0343] Step 7:
[0344] The server adjusts the ranking of the retrieved search results taking into account the user's emotional state. If the emotion engine detects that the user is stressed, it prioritizes displaying materials that are more concise and easy to understand. The input is the list of related materials from step 6 and the emotion analysis data, and the output is search results that have been ranked based on emotion.
[0345] Step 8:
[0346] The server sends the final search results to the terminal, where the user can check the search results on the terminal and view the necessary materials. The input is the search results ranked based on the sentiment in step 7, and the output is the search results displayed on the terminal.
[0347] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0348] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0349] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0350] [Second embodiment]
[0351] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0352] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0353] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0354] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0355] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0356] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0357] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0358] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0359] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0360] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0361] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0362] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0363] The present invention is a system for analyzing and searching data collected from an in-house mail server, a messaging platform, and cloud storage. The program processing of this system is explained below in natural language.
[0364] First, the server periodically connects to the company's internal mail server, messaging platforms (e.g., Slack and Microsoft Teams), and cloud storage (e.g., Google Drive and Dropbox) to collect all data, including email bodies and attachments, chat logs and uploaded files from messaging platforms, and various files from cloud storage.
[0365] Next, the server preprocesses the collected data. During this preprocessing step, the server analyzes the text data using natural language processing (NLP) technology, removes unnecessary information, and converts it into a format that is easier to analyze. For image data, the server also uses optical character recognition (OCR) technology to extract text information from the image. Furthermore, the server analyzes files such as PDFs and Word documents and converts them into text data.
[0366] The server then builds a multimodal AI model based on the preprocessed data (text data, image data, and other data formats). This AI model can comprehensively analyze text, images, and other data formats to extract data characteristics. This model is then trained to determine data similarity.
[0367] When a user searches for specific information on their device, they input a search query. For example, they input a query such as "Marketing strategy for new product Y." The device then sends this search query to a server. Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's internal database. The server then performs a search based on the query and ranks and extracts the most relevant materials.
[0368] The server also searches past related communication logs. For example, if the relevant document was sent by email, it searches the email body and attachments and extracts related Slack messages. This allows users to view the communication logs related to the document along with the related documents.
[0369] The server sends the search results to the terminal and displays them to the user. The user can check the displayed search results and select and view the necessary materials. For example, past reports on the marketing strategy for new product Y, presentation materials, related meeting notes, etc. can be displayed and viewed.
[0370] This system improves work efficiency by making it easy to search and extract necessary documents even when documents are not properly handed over after an employee leaves or is transferred. It also supports not only text but also images and other data formats, allowing for quick and efficient searches of various documents. This provides an environment that supports continuous business operations without compromising internal knowledge.
[0371] The processing flow will be explained below.
[0372] Step 1: Collect data
[0373] The server collects data from the company's email server, messaging platform, and cloud storage. From the email server, it collects email text, attachments, and metadata. From the messaging platform, it obtains chat logs and uploaded files. It also collects various document files (PDF, Word, image files, etc.) from cloud storage.
[0374] Step 2: Preprocessing the data
[0375] The server preprocesses the collected data by format. For text data, it analyzes it using natural language processing (NLP) technology and removes noise. For image data, it uses optical character recognition (OCR) technology to extract text information from the image. PDFs and Word documents are also converted into text data and their contents are analyzed.
[0376] Step 3: Building an AI model
[0377] The server then builds a multimodal AI model based on the preprocessed data. This model is trained to comprehensively analyze text, images, and other data formats and extract data characteristics, resulting in a model capable of similarity judgment.
[0378] Step 4: Receiving a search query
[0379] The terminal receives a search query from the user. For example, the user inputs a search query such as "Marketing strategy for new product Y." The terminal then sends this query to the server.
[0380] Step 5: Processing the search query
[0381] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's database. It performs a query-based search and ranks and extracts highly relevant materials. It performs an integrated search of materials, including not only text data but also different data formats such as images and PDFs.
[0382] Step 6: Search past communication logs
[0383] The server also searches past communication logs (emails and messages on messaging platforms) for related materials, extracting relevant email text and Slack messages, and searching for information from the communication logs.
[0384] Step 7: Serving search results
[0385] The server sends the search results and information extracted from past communication logs to the terminal. The terminal displays this data on a user interface, allowing the user to check each piece of information. For example, a report on the marketing strategy for new product Y and related communication records may be displayed.
[0386] Step 8: User validation and utilization
[0387] Users can check the search results displayed on their devices, select and view related materials as needed, and download or print any information they deem necessary for further use.
[0388] In this way, the system quickly searches for necessary information from vast amounts of internal company data, improving business efficiency.
[0389] Example 1
[0390] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0391] With conventional systems, it was difficult to efficiently collect and analyze data dispersed across internal emails, messaging platforms, and cloud storage, making it impossible to quickly search and retrieve the information needed. Furthermore, there was a lack of technology for comprehensively analyzing different data formats (text data, image data, PDF files, etc.), which led to problems that reduced operational efficiency. Furthermore, there was a need for a system that could include past communication logs in search results, allowing users to check the history of related communications.
[0392] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0393] In this invention, the server includes: means for collecting data from mail servers, messaging platforms, and cloud storage; means for analyzing text data, image data, and other data in preprocessing and integrating them; means for analyzing text data using natural language processing technology; means for extracting text information from image data using optical character recognition technology; means for building a multimodal AI model and analyzing data similarity; means for receiving search queries from users and identifying relevant data based on the search queries; means for searching past communication logs for sending relevant materials; and means for providing search results to user terminals. This enables centralized collection and analysis of distributed data within a company, allowing users to quickly and efficiently search for and obtain the information they need. It also enables integrated analysis of different data formats and provides comprehensive search results that include related communication history.
[0394] A "mail server" is a server that manages the sending and receiving of emails. Its role is to provide mailboxes, store sent emails, and forward received emails to users.
[0395] A "messaging platform" refers to a communication tool that provides functions such as text messaging, file sharing, voice calls, and video calls. Typical examples include online chat services and business communication tools.
[0396] "Cloud storage" refers to remote data storage services available over the Internet, allowing users to store data in the cloud and access, share, and manage it as needed.
[0397] "Data collection tools" refers to the technologies and processes used to extract and centralize data from different sources. This tool allows data from various platforms to be managed in a unified manner.
[0398] "Preprocessing" refers to a series of operations performed to convert raw data into a suitable format for data analysis and machine learning, including text segmentation, noise removal, and format conversion.
[0399] "Natural language processing technology" refers to technology that enables computers to understand, analyze, and generate human language, including text analysis, machine translation, and sentiment analysis.
[0400] Optical character recognition technology is a technology that analyzes characters in images captured by a scanner or camera and converts them into text data. It is used to digitize printed documents and handwritten characters.
[0401] A "multimodal AI model" is an artificial intelligence model that can integrate and analyze data in different formats (text, images, audio, etc.). Combining multiple data sources enables more accurate analysis.
[0402] A "search query" is a question or keyword that a user enters to search for specific information. It is used to refer to information in a search engine or database.
[0403] "Means for identifying relevant data" refers to techniques and processes for finding and extracting relevant data based on a search query from a user.
[0404] "Past communication logs" refers to records of past communications, such as email sending and receiving history and chat history on messaging platforms.
[0405] "Search Results" means a list of relevant data or information based on a user's search query, generated by a search engine or database and provided to the user.
[0406] "User terminal" refers to a device such as a computer, smartphone, or tablet that is directly operated by a user. It is a device that connects to a server via the Internet or an internal company network and allows users to use services.
[0407] The present invention is a system for collecting data from an in-house mail server, a messaging platform, and cloud storage, and analyzing and searching the data. Specific embodiments of this system are described below.
[0408] Hardware and Software Configuration
[0409] The server includes the following main software and hardware:
[0410] A scheduler (such as cron) to periodically run scripts (such as Python) for data collection
[0411] IMAP protocol library for connecting to mail servers
[0412] API client libraries for messaging platforms and cloud storage (e.g., Slack API, Google Drive API)
[0413] Libraries for using natural language processing technologies (NLTK and spaCy)
[0414] Library for using optical character recognition technology (Tesseract-OCR)
[0415] Libraries for building and training multimodal AI models (TensorFlow and PyTorch)
[0416] The terminals include computers, smartphones, tablets, etc. that are directly operated by the user and allow the user to access the system's search interface through a browser.
[0417] Data collection
[0418] The server periodically runs a data collection script that connects to the company's mail server, messaging platform, and cloud storage, collecting emails, chat logs, uploaded files, and various documents. The server then stores this data in local storage.
[0419] Data Preprocessing
[0420] The server preprocesses the collected data. Specifically, it uses natural language processing technology to analyze the text data and remove unnecessary information. Examples include removing stop words and stemming. For image data, it uses optical character recognition technology to extract text information, and converts PDFs and Word documents into text data. This integrates all data into an analyzable format.
[0421] Building a multimodal AI model
[0422] The server integrates preprocessed text data, image data, and other data formats into a single dataset. A multimodal AI model is built using TensorFlow and PyTorch to learn the characteristics of the data. This model can integrate and analyze different data formats, enabling more accurate searches.
[0423] User search and results display
[0424] A user accesses the system's search interface from a browser on their device and enters a search query, such as "marketing strategy for new product Y." The device then sends this search query to the server.
[0425] The server analyzes the received search query and uses a multimodal AI model to identify related materials. It also searches related past communication logs (including email and chat history) and displays them in an integrated manner, making it easier for users to grasp related information.
[0426] Providing search results
[0427] The server provides search results in JSON format to the terminal, which then displays them in HTML format, allowing the user to quickly and efficiently obtain the information they need. For example, past reports on marketing strategies for new product Y, related presentation materials, and meeting notes can be displayed.
[0428] Examples and prompts
[0429] For example, when searching using the prompt phrase "Marketing strategy for new product Y," the system quickly searches for relevant internal documents and past communication history and displays them to the user, allowing the user to carry out their work more efficiently.
[0430] In this way, this system provides specific methods and technologies for efficiently collecting and analyzing huge amounts of data and quickly responding to user search requests.
[0431] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0432] Program processing flow
[0433] Step 1: Collect data
[0434] The server periodically runs data collection scripts to connect to the company's mail server, messaging platform, and cloud storage.
[0435] Input: Connection and authentication information configured for the server
[0436] How it works: It connects to mail servers using the IMAP protocol, and to messaging platforms and cloud storage using the APIs of each service.
[0437] Output: Save retrieved emails, chat logs, and various documents to local storage.
[0438] Step 2: Preprocessing the data
[0439] The server pre-processes the collected data.
[0440] Input: Collected emails, chat logs, document files
[0441] How it works: It uses natural language processing (NLP) techniques to analyze text data and remove unnecessary information. Specifically, it performs text segmentation, stop word removal, and stemming. It also uses Tesseract-OCR to perform character recognition on image data, and PDF and Word documents are converted to text using PDFMiner and python-docx.
[0442] Output: Preprocessed text data, image data, and other data are integrated and saved.
[0443] Step 3: Building a multimodal AI model
[0444] The server uses the integrated data to build a multimodal AI model.
[0445] Input: Preprocessed dataset
[0446] How it works: We use TensorFlow and PyTorch to build models and train them to learn data features, specifically by integrating text, images, and other data formats and determining data similarity.
[0447] Output: A trained multimodal AI model
[0448] Step 4: Receiving and parsing the search query
[0449] A user accesses the search interface from a browser on the terminal and enters a search query.
[0450] Input: The search query entered by the user (e.g., "Marketing strategy for new product Y")
[0451] Operation: The terminal sends this query to the server, which analyzes the received query.
[0452] Output: Parsed search query
[0453] Step 5: Search and identify relevant data
[0454] The server uses the parsed search query to search for relevant data.
[0455] Input: Parsed search query, trained multimodal AI model
[0456] How it works: Uses a multimodal AI model to identify and rank relevant materials and data based on a search query.
[0457] Output: A list of related data
[0458] Step 6: Search past communication logs
[0459] The server searches past communication logs based on the associated data.
[0460] Input: List of related data
[0461] How it works: Uses email and messaging platform APIs to search and extract relevant past communication logs.
[0462] Output: Related past communication logs
[0463] Step 7: Serving search results
[0464] The server integrates the relevant data with past communication logs to generate search results.
[0465] Input: List of related data, past communication logs
[0466] What it does: Generates a unified search result in JSON format and sends it to the device.
[0467] Output: JSON data of search results
[0468] Step 8: Viewing search results
[0469] The terminal displays the received search results to the user.
[0470] Input: JSON data of search results sent from the server
[0471] What it does: Converts the results to HTML and displays them in the browser, allowing the user to view the results.
[0472] Output: Search results displayed to the user
[0473] This series of processes allows the user to efficiently search for and view the information they need.
[0474] (Application example 1)
[0475] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0476] While conventional in-house data search systems can collect and preprocess data from email servers and messaging platforms, they face the problem of making it difficult for store staff to quickly and efficiently obtain the information they need. In particular, there is a need for a way to quickly obtain information in real time, such as inventory information and past sales history, while serving customers in stores. Furthermore, there is a lack of systems that can handle multiple data formats and perform integrated analysis. Therefore, an effective solution is needed to improve the work efficiency of store staff.
[0477] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0478] In this invention, the server includes means for collecting data from mail servers, messaging platforms, and cloud storage, means for analyzing text data, image data, and other data in preprocessing and integrating them, means for building a multimodal AI model and analyzing data similarity, means for receiving search queries from users and identifying related data based on the search queries, means for searching past communication logs for sending related materials, means for providing search results to user terminals, and means for in-store staff to send queries by voice input and display search results through smart glasses, enabling store staff to obtain information in real time and respond to customers quickly and efficiently.
[0479] A "mail server" is a computer server that manages and stores the sending and receiving of email.
[0480] A "messaging platform" is an online communication tool for sending and receiving text messages and files in real time.
[0481] "Cloud storage" is an online storage service for storing, managing, and accessing data via the Internet.
[0482] "Preprocessing" refers to a data processing step to convert collected data into a format that is easy to analyze.
[0483] "Text data" is information expressed as characters or sentences.
[0484] "Image data" is information expressed in an image format.
[0485] "Integration" means combining multiple different data formats and information into one format and making them consistent.
[0486] A "multimodal AI model" is an artificial intelligence model that performs integrated analysis of multiple data formats (text, images, etc.).
[0487] "Data similarity" refers to common characteristics or patterns between different data.
[0488] A "search query" is a search question or keyword entered by a user to search for specific information.
[0489] "Related Materials" are useful information or documents identified based on the search query.
[0490] "Past communication logs" are records of messages and emails previously exchanged.
[0491] A "user terminal" is an electronic device that a user operates and receives information from, such as a smartphone or tablet.
[0492] A "brick and mortar store" is a sales facility that has a physical presence.
[0493] "Voice input" is a method of inputting characters and commands using voice.
[0494] "Smart glasses" are a wearable eyeglass-type device equipped with a display function.
[0495] The present invention is a system that enables in-store staff to quickly and efficiently obtain necessary information while serving customers. The system collects data from mail servers, messaging platforms, and cloud storage, and analyzes them in an integrated manner to provide relevant data based on user search queries.
[0496] 1. Program Overview
[0497] The server operates programs to achieve the following functions:
[0498] Data Collection: Periodically collect data from mail servers, messaging platforms, and cloud storage.
[0499] Pre-processing: Analyze collected data using NLP techniques (e.g., spaCy) and OCR techniques (e.g., Tesseract) to extract and convert text and image data.
[0500] Building multimodal AI models: Multimodal AI models can be trained based on text data, image data, and other data formats to determine data similarity.
[0501] Search query processing: Receives and analyzes search queries voice-entered by store staff through smart glasses.
[0502] Sending search results: Search results for related materials and past communication logs are displayed on the smart glasses.
[0503] 2. Hardware and Software
[0504] Hardware used: Server, smart glasses, user devices (e.g., smartphones and tablets).
[0505] Software used: NLP engine (spaCy), OCR engine (Tesseract), multimodal AI model (OpenAI's GPT-4).
[0506] 3. System operation explanation
[0507] The server first collects the necessary data from various data sources. The collected data undergoes a preprocessing step to convert it into a format that is easy to analyze. During this process, NLP and OCR technologies are used to extract text data, and image data is also handled. After preprocessing is complete, the data is analyzed comprehensively using a multimodal AI model to extract data characteristics and similarities.
[0508] The user, a store staff member, voice-inputs queries such as customer questions or inventory checks. For example, prompts such as "Check inventory, tell me the availability of product A" or "Search customer X's purchase history" are sent through the smart glasses. The server receives the query, searches relevant data using a multimodal AI model, and ranks and extracts the results. These results are then provided to the smart glasses, allowing the user to view the information in real time.
[0509] 4. Specific Examples
[0510] For example, if a store staff member sends a voice prompt such as "Tell me the purchase history of customer Yamada Taro for the past three months," the system will quickly collect relevant past communication logs, inventory information, and sales data and display them on the smart glasses' display. It also works in the same way with prompts such as "Show me the marketing materials for new product A that will be on sale next weekend," supporting efficient customer service.
[0511] The present invention allows staff in physical stores to obtain information in real time, enabling them to respond to customers quickly and accurately.
[0512] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0513] Step 1:
[0514] Data collection
[0515] The server periodically connects to mail servers, messaging platforms, and cloud storage to collect data, including email bodies and attachments, chat logs and uploaded files from messaging platforms, and various files from cloud storage.
[0516] Input: Each data source (mail server, messaging platform, cloud storage)
[0517] Output: Raw data collected (emails, chat logs, files)
[0518] Specifically, the server accesses these data sources using APIs or dedicated protocols, retrieves the latest data, and stores it.
[0519] Step 2:
[0520] Data Preprocessing
[0521] The server preprocesses the collected data. During this preprocessing step, the server analyzes the text data using natural language processing (NLP) technology, removes unnecessary information, and converts it into a format that is easier to analyze. For image data, the server also uses optical character recognition (OCR) technology to extract text information from the image. Furthermore, the server analyzes files such as PDFs and Word documents and converts them into text data.
[0522] Input: Raw data collected
[0523] Output: Preprocessed data (text format)
[0524] Specifically, the system uses an NLP engine (e.g., spaCy) to analyze text and remove unnecessary parts, and an OCR engine (e.g., Tesseract) to extract text information from the image.
[0525] Step 3:
[0526] Building a multimodal AI model
[0527] The server uses the preprocessed data to build a multimodal AI model that can comprehensively analyze text, images, and other data formats, and trains the model to determine data similarity.
[0528] Input: Preprocessed data
[0529] Output: A trained multimodal AI model
[0530] Specifically, it feeds data to AI models (such as OpenAI's GPT-4) to train them, and also continuously updates and improves them.
[0531] Step 4:
[0532] Receiving a search query
[0533] The user (store staff) sends a search query by voice input through the smart glasses, for example, "Check inventory, tell me the stock status of product X." The device converts this voice data into text format and sends it to the server.
[0534] Input: Voice query
[0535] Output: Text query
[0536] Specifically, it converts voice input into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API).
[0537] Step 5:
[0538] Finding related data
[0539] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's database, ranking and extracting the most relevant materials.
[0540] Input: Text query, trained multimodal AI model, in-house database
[0541] Output: Search results (list of related materials)
[0542] Specifically, the server inputs a query into the model, searches the database for materials that match the query, and retrieves the results.
[0543] Step 6:
[0544] Submitting and viewing search results
[0545] The server sends the search results to the user's device (smart glasses), and the user can view the displayed search results, such as "Report on marketing strategy for new product X" or "Related meeting notes."
[0546] Input: Search results
[0547] Output: Results displayed on smart glasses
[0548] Specifically, the server formats the search results and transmits them in a format suitable for the smart glasses' display.
[0549] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0550] This invention is a system that collects data from an in-house mail server, messaging platform, and cloud storage, and combines it with an emotion engine that recognizes the user's emotions to perform more effective document searches. Below, the program processing of this system is explained in natural language.
[0551] First, the server periodically connects to the company's internal mail server, messaging platforms (e.g., Slack and Microsoft Teams), and cloud storage (e.g., Google Drive and Dropbox) to collect all data, including email bodies and attachments, chat logs and uploaded files from messaging platforms, and various files from cloud storage.
[0552] Next, the server preprocesses the collected data. During this preprocessing step, the server analyzes the text data using natural language processing (NLP) technology, removes unnecessary information, and converts it into a format that is easier to analyze. For image data, the server also uses optical character recognition (OCR) technology to extract text information from the image. Furthermore, the server analyzes files such as PDFs and Word documents and converts them into text data.
[0553] The server then builds a multimodal AI model based on the preprocessed data (text data, image data, and other data formats). This AI model can comprehensively analyze text, images, and other data formats to extract data characteristics. This model is then trained to determine data similarity.
[0554] When a user searches for specific information on a device, they input a search query. For example, they input a query such as "Marketing strategy for new product Y." The device then sends this search query to the server. At this stage, the device can also use an emotion engine to detect the user's emotions. The emotion engine recognizes the user's facial expressions and voice and determines the user's current emotional state.
[0555] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's internal database. It then performs a search based on the query and ranks and extracts the most relevant materials. It performs an integrated search of materials that include not only text data but also different data formats such as images and PDFs. Furthermore, it adjusts the ranking of search results by taking into account the user's emotional state detected by the emotion engine. For example, if the user is feeling stressed, it will prioritize displaying materials that are more concise and easy to understand.
[0556] The server also searches past related communication logs. For example, if the relevant document was sent by email, it searches the email body and attachments and extracts related Slack messages. This allows users to view the communication logs related to the document along with the related documents.
[0557] The server sends the search results to the terminal and displays them to the user. The user can check the displayed search results and select and view the necessary materials. For example, past reports on the marketing strategy for new product Y, presentation materials, related meeting notes, etc. are displayed and can be reviewed. In addition, an emotion engine is used to provide information tailored to the user's emotions, improving the user experience.
[0558] This system improves work efficiency by making it easy to search and retrieve necessary documents even when documents are not properly handed over after an employee leaves or is transferred. It also supports not only text but also images and other data formats, allowing for quick and efficient searches of various documents. Furthermore, by utilizing an emotion engine, it is possible to provide search results that reflect the user's emotional state, improving user satisfaction.
[0559] The processing flow will be explained below.
[0560] Step 1: Collect data
[0561] The server periodically connects to the company's email server, messaging platform, and cloud storage to collect data. Specifically, it collects email text, attachments, and metadata from the email server, chat logs and uploaded files from the messaging platform, and various document files (PDF, Word, image files, etc.) from the cloud storage.
[0562] Step 2: Preprocessing the data
[0563] The server preprocesses the collected data by format. For text data, it analyzes it using natural language processing (NLP) technology and removes noise. For image data, it uses optical character recognition (OCR) technology to extract text information from the image. PDFs and Word documents are also converted into text data and their contents are analyzed.
[0564] Step 3: Building an AI model
[0565] The server then builds a multimodal AI model based on the preprocessed data. This model is trained to comprehensively analyze text, images, and other data formats and extract data characteristics, resulting in a model capable of similarity judgment.
[0566] Step 4: Receiving a search query
[0567] The device receives a search query from the user. For example, the user may enter a search query such as "Marketing strategy for new product Y." The device sends this query to the server and simultaneously activates an emotion engine to detect the user's emotional state.
[0568] Step 5: Detecting Emotional State
[0569] The device uses an emotion engine to analyze the user's emotional state. Specifically, it recognizes emotions from the user's facial expressions and voice and determines the user's current emotional state. For example, it detects whether the user is feeling stressed or relaxed.
[0570] Step 6: Processing the search query
[0571] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's database. It performs a query-based search and ranks and extracts highly relevant materials. It performs an integrated search of materials, including not only text data but also different data formats such as images and PDFs.
[0572] Step 7: Adjust your results based on your emotional state
[0573] The server adjusts the ranking of search results based on the user's emotional state as detected by the emotion engine. For example, if the user is feeling stressed, it will prioritize more concise and easy-to-understand materials. On the other hand, if the user is relaxed, it will provide more detailed and specialized materials.
[0574] Step 8: Search past communication logs
[0575] The server also searches past communication logs (emails and messages on messaging platforms) for related materials, extracting relevant email text, Slack messages, etc., and searching for information from the communication logs.
[0576] Step 9: Serving search results
[0577] The server sends the search results and information extracted from past communication logs to the terminal. The terminal displays this data on a user interface, allowing the user to check each piece of information. For example, a report on the marketing strategy for new product Y and related communication records may be displayed.
[0578] Step 10: User validation and utilization
[0579] Users can check the search results displayed on their devices, select and view related materials as needed, and download or print any information they deem necessary for further use.
[0580] In this way, this system aims to improve work efficiency and user satisfaction by quickly searching for necessary information from vast amounts of internal company data and providing information that corresponds to the user's emotional state.
[0581] Example 2
[0582] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0583] Conventional document search systems struggled to effectively collect and integrate data from internal email servers, messaging platforms, and cloud storage, and lacked a way to appropriately adjust search results based on the user's emotional state. This resulted in problems such as an inability to respond quickly and accurately to user requests, leading to reduced work efficiency. Furthermore, it was difficult to comprehensively search documents in different data formats, making it difficult to find the information needed when document handover was insufficient.
[0584] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0585] In this invention, the server includes: means for collecting data from mail servers, messaging platforms, and cloud storages; means for analyzing text data, image data, and other data using natural language processing technology and optical character recognition technology in preprocessing and integrating them; means for building a multimodal AI model including text, image, and other data formats and analyzing the similarity of the data; means for receiving a user's search query and identifying related data based on the search query using an emotion engine that detects the user's emotional state; means for searching past communication logs to send related materials; and means for providing the search results to the user terminal.
[0586] This makes it possible to centrally collect and integrate data from various data sources within the company and provide appropriate search results based on user sentiment. Users can quickly find the documents they need, improving work efficiency. In addition, by searching documents in different data formats in an integrated manner, related information can be easily obtained even if document handover is insufficient.
[0587] A "mail server" is a server that manages the sending and receiving of email within an organization or on the Internet.
[0588] A "messaging platform" is software or a service for sending and receiving messages in real time.
[0589] "Cloud storage" is an online service for storing and managing data over the Internet.
[0590] "Natural language processing technology" is a technology that analyzes text data and extracts linguistic features.
[0591] "Optical character recognition technology" is a technology that extracts text information from image data.
[0592] A "multimodal AI model" is an artificial intelligence model that comprehensively analyzes text, images, and other data formats to extract features.
[0593] "User emotional state" refers to the emotional state a user is experiencing when performing a search, and includes techniques for detecting this.
[0594] An "emotion engine" is a device or software that analyzes a user's facial expressions and voice to detect their emotional state.
[0595] A "search query" is text data indicating a search request that a user inputs into a search system.
[0596] "Communication logs" are data containing past communication history across email and messaging platforms.
[0597] The present invention is a system that collects data from an internal mail server, a messaging platform, and cloud storage, and combines it with an emotion engine that recognizes the user's emotional state to perform more effective document search. The program processing of this system is described below.
[0598] First, the server periodically connects to the company's email server, messaging platform, and cloud storage to collect all data. Specifically, Microsoft Exchange and Gmail are used as email servers, Slack and Microsoft Teams are used as messaging platforms, and Google Drive and Dropbox are used as cloud storage. The server uses APIs to collect email bodies and attachments, chat logs, uploaded files, and various files in cloud storage.
[0599] The server then preprocesses the collected data. This preprocessing step uses natural language processing (NLP) and optical character recognition (OCR) techniques. For text data, NLTK and spaCy are used to tokenize and remove stop words. For image data, Tesseract OCR is used to extract text information from images. For PDF and Word documents, Apache Tika is used to convert them to text data.
[0600] The server then builds a multimodal AI model based on the preprocessed data. This model is created using TensorFlow and PyTorch to comprehensively analyze text, images, and other data formats and extract data characteristics. Specifically, it uses a convolutional neural network (CNN) to extract image features, and a recurrent neural network (RNN) or BERT model to extract text features.
[0601] Next, the user inputs a search query from the terminal. For example, the query "Marketing strategy for new product Y" is input. This search query is sent as is from the terminal to the server as a character string.
[0602] The device is equipped with an emotion engine that detects the user's emotional state, for example by capturing and analyzing facial expressions with a camera. This analysis is performed using the Face API from Microsoft Azure's Cognitive Services. The emotion engine detects the user's emotional state and transmits that information along with a query to the server.
[0603] Based on the received search query and emotional information, the server uses a multimodal AI model to identify relevant materials from the company's database. It then assesses data similarity and ranks and extracts highly relevant materials. It also adjusts the ranking of results according to the user's emotional state. If the user is feeling stressed, the algorithm prioritizes displaying more concise and easy-to-understand materials.
[0604] The server also searches past communication logs for related materials. For example, if the material was sent by email, it extracts the email body, attachments, and chat logs from related messaging platforms. Finally, the server sends the search results and related communication logs to the device.
[0605] Users can select and view the necessary materials from the results displayed on their device. For example, they can check past reports, presentation materials, meeting notes, etc. related to "Marketing Strategy for New Product Y."
[0606] As a concrete example, here is a prompt example for a generative AI model when searching for "marketing strategy for new product Y":
[0607] "Search your internal database for materials related to the marketing strategy for new product Y. If the user is feeling stressed, prioritize materials that are more concise and easy to understand."
[0608] This system makes it possible to quickly and efficiently search and retrieve necessary documents even when documents are not fully handed over when an employee resigns or is transferred.In addition, by providing search results that respond to the user's emotions, it is expected that the user experience will improve and business efficiency will increase.
[0609] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0610] Step 1:
[0611] The server periodically connects to the company's mail server, messaging platform, and cloud storage to collect data. In this step, it retrieves email bodies and attachments from the mail server, chat logs and uploaded files from the messaging platform, and various files from the cloud storage. The specific input is a request using the API of each data source, and the output is a collection of collected data.
[0612] Step 2:
[0613] The server preprocesses the collected data. It uses natural language processing (NLP) and optical character recognition (OCR) techniques to tokenize text data, remove stop words, extract text information from image data, and convert PDFs and Word documents into text data. Specific operations include text analysis using NLTK and spaCy, image analysis using Tesseract OCR, and document analysis using Apache Tika. The input is the raw data collected in step 1, and the output is the preprocessed, clean data.
[0614] Step 3:
[0615] The server builds a multimodal AI model based on the preprocessed data. It uses TensorFlow and PyTorch to comprehensively analyze text, images, and other data formats and extract features. Specifically, it uses a Convolutional Neural Network (CNN) to extract image features, and a Recurrent Neural Network (RNN) or BERT model to extract text data features. The input is various preprocessed data, and the output is a trained AI model.
[0616] Step 4:
[0617] A user inputs a query to search for specific information from a terminal. For example, the query might be "Marketing strategy for new product Y." The query is sent directly from the terminal to the server in the form of a string. The input is the query entered by the user, and the output is a search request sent to the server.
[0618] Step 5:
[0619] The device uses an emotion engine to detect the user's emotional state. Specifically, it captures the user's facial expressions with a camera and analyzes them using the Face API of Microsoft Azure's Cognitive Services. The input is an image of the user's facial expression, and the output is the detected emotional state data.
[0620] Step 6:
[0621] The server uses a multimodal AI model to search for relevant materials from the company's database based on the received search query and emotional information. It ranks the search results and adjusts the ranking based on the emotion engine's evaluation. Specifically, if the user is feeling stressed, the server uses an algorithm that prioritizes displaying concise and easy-to-understand materials. The input is the search query and emotional information, and the output is ranked search results.
[0622] Step 7:
[0623] The server also searches past communication logs for related materials. For example, if the material was sent by email, it extracts the email body, attachments, and chat logs from related messaging platforms. The input is the search results, and the output is the related communication logs.
[0624] Step 8:
[0625] The server sends the search results and related communication logs to the terminal. The user can select and view the necessary materials from the results displayed on the terminal. The input is the ranked search results and related communication logs, and the output is the information displayed on the terminal in a user-accessible form.
[0626] Through these steps, users can quickly and efficiently search and extract the information they need.
[0627] (Application example 2)
[0628] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0629] Providing timely and appropriate information is essential to improving work efficiency in factories and reducing worker stress. However, conventional information search systems were unable to provide information tailored to the emotional state of workers, making it difficult to provide efficient work support. Furthermore, extracting information from a wide variety of data formats and past communication logs was time-consuming, making it difficult to provide practical information.
[0630] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0631] In this invention, the server includes: means for collecting data from mail servers, messaging platforms, and cloud storage; means for analyzing and integrating text data, image data, and other data in preprocessing; means for building a multimodal AI model and analyzing data similarity; means for receiving search queries from users and identifying relevant data based on the search queries; means for analyzing the emotional state of workers and adjusting search results based on the results; means for searching past communication logs for sending relevant materials; and means for providing search results to user terminals. This enables timely and appropriate information provision that takes emotional states into account. Furthermore, the system can efficiently extract information from a variety of data formats and past communication logs, thereby providing efficient support for workers.
[0632] A "mail server" is a server for sending, receiving, and storing e-mail.
[0633] A "messaging platform" is a system for exchanging messages between users using text, voice, video, etc.
[0634] "Cloud storage" refers to online storage services for storing and accessing data over the Internet.
[0635] "Means for collecting data" refers to the mechanisms for obtaining the necessary data from mail servers, messaging platforms, and cloud storage.
[0636] "Preprocessing" is the process of converting collected data into a format that is easy to analyze and removing unnecessary information.
[0637] "Text data" is a data format that includes character information.
[0638] "Image data" is a data format that contains visual information.
[0639] A "multimodal AI model" is an artificial intelligence model for integrated analysis of different data formats (text, images, etc.).
[0640] A "user search query" is a search request entered by a user to locate specific information.
[0641] The "emotional state" is a state that indicates the type and intensity of the emotion that the user is currently feeling.
[0642] "Related materials" are materials that contain information that is useful to the user and that are identified based on the search query.
[0643] A "communication log" is a record of messages and files sent and received in the past, as well as related information.
[0644] "Search Results" means a list of relevant materials identified based on a search query.
[0645] A "user terminal" is a device that a user operates directly to check information.
[0646] To realize this application example, the system program is configured as follows.
[0647] The server collects data from email servers, messaging platforms, and cloud storage. Specifically, it uses the Gmail API, Microsoft Teams API, Google Drive API, and other sources to obtain the necessary data. The obtained data is in the form of text, image data, and other data formats, and natural language processing (NLP) and optical character recognition (OCR) technologies are used to preprocess it. After preprocessing, the data is analyzed using Hugging Face's transformers library to build a multimodal AI model. The multimodal AI model comprehensively analyzes text, image data, and other data formats.
[0648] The server also receives search queries from users and identifies relevant data based on the search queries. A user can input a search query such as "How to troubleshoot error code E123 on machine X" using a device (e.g., a smartphone or smart glasses). The server also analyzes the user's emotional state in real time using a device equipped with an emotion engine (e.g., smart glasses with a built-in camera and microphone that analyzes facial expressions and voice). If the user is feeling stressed, the emotion engine will prioritize providing more concise and easy-to-understand materials.
[0649] The search engine searches the database based on the query and also searches past related communication logs. For example, if related documents were sent via email or messaging platforms, the contents of those documents will also be searched. This allows for comprehensive retrieval of past communications and related documents. Search results are ranked according to the user's emotional state, with the most relevant information being displayed first.
[0650] The terminal provides the search results to the user, who can then check and view the necessary materials. This improves work efficiency and reduces worker stress. This system makes it possible to efficiently extract information from a wide range of data formats, significantly improving the factory work environment.
[0651] As a concrete example, consider a factory worker troubleshooting a new machine. If the worker types the query "How to troubleshoot error code E123 on machine X" into the smart glasses and is stressed, the application will prioritize short, easy-to-understand instructions.
[0652] Example prompt sentence:
[0653] "Find past troubleshooting documents for error code E123 on machine X. Workers are stressed right now, so please prioritize concise documents."
[0654] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0655] Step 1:
[0656] The server collects data from email servers, messaging platforms, and cloud storage. As input, it uses the Gmail API, Microsoft Teams API, Google Drive API, etc. to obtain the target email, message, and file data. The output is the collected text data, image data, and other data formats.
[0657] Step 2:
[0658] The server preprocesses the collected data. Specifically, it uses natural language processing (NLP) technology to analyze the text data, remove unnecessary information, and convert it into a format that is easier to analyze. It also uses optical character recognition (OCR) technology to extract text information from image data. It also analyzes files such as PDFs and Word documents and converts them into text data. The input is the data collected in step 1, and the output is the preprocessed data.
[0659] Step 3:
[0660] The server builds a multimodal AI model based on various preprocessed data (text data, image data, and other data formats). Using Hugging Face's transformers library, it trains an AI model that can comprehensively analyze different data formats. The input is the preprocessed data output from Step 2, and the output is an AI model that can analyze data similarities.
[0661] Step 4:
[0662] A user uses a device (e.g., a smartphone, smart glasses) to input a search query, such as "how to troubleshoot error code E123 on machine X." The input is the user's query, and the output is the query data.
[0663] Step 5:
[0664] The device sends the search query to the server. At the same time, the device uses an emotion engine to analyze the user's emotional state. For example, it uses the camera and microphone built into the smart glasses to analyze the user's facial expressions and voice to determine their current emotional state. The input is the query from step 4 and the user's emotional data, and the output is composite data including the query and the emotional state.
[0665] Step 6:
[0666] The server searches the database based on the search query, using a multimodal AI model to identify relevant information and also searches past communication logs to retrieve relevant materials. The input is the composite data from step 5, and the output is a list of relevant materials.
[0667] Step 7:
[0668] The server adjusts the ranking of the retrieved search results taking into account the user's emotional state. If the emotion engine detects that the user is stressed, it prioritizes displaying materials that are more concise and easy to understand. The input is the list of related materials from step 6 and the emotion analysis data, and the output is search results that have been ranked based on emotion.
[0669] Step 8:
[0670] The server sends the final search results to the terminal, where the user can check the search results on the terminal and view the necessary materials. The input is the search results ranked based on the sentiment in step 7, and the output is the search results displayed on the terminal.
[0671] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0672] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0673] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0674] [Third embodiment]
[0675] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0676] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0677] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0678] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0679] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0680] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0681] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0682] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0683] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0684] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0685] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0686] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0687] The present invention is a system for analyzing and searching data collected from an in-house mail server, a messaging platform, and cloud storage. The program processing of this system is explained below in natural language.
[0688] First, the server periodically connects to the company's internal mail server, messaging platforms (e.g., Slack and Microsoft Teams), and cloud storage (e.g., Google Drive and Dropbox) to collect all data, including email bodies and attachments, chat logs and uploaded files from messaging platforms, and various files from cloud storage.
[0689] Next, the server preprocesses the collected data. During this preprocessing step, the server analyzes the text data using natural language processing (NLP) technology, removes unnecessary information, and converts it into a format that is easier to analyze. For image data, the server also uses optical character recognition (OCR) technology to extract text information from the image. Furthermore, the server analyzes files such as PDFs and Word documents and converts them into text data.
[0690] The server then builds a multimodal AI model based on the preprocessed data (text data, image data, and other data formats). This AI model can comprehensively analyze text, images, and other data formats to extract data characteristics. This model is then trained to determine data similarity.
[0691] When a user searches for specific information on their device, they input a search query. For example, they input a query such as "Marketing strategy for new product Y." The device then sends this search query to a server. Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's internal database. The server then performs a search based on the query and ranks and extracts the most relevant materials.
[0692] The server also searches past related communication logs. For example, if the relevant document was sent by email, it searches the email body and attachments and extracts related Slack messages. This allows users to view the communication logs related to the document along with the related documents.
[0693] The server sends the search results to the terminal and displays them to the user. The user can check the displayed search results and select and view the necessary materials. For example, past reports on the marketing strategy for new product Y, presentation materials, related meeting notes, etc. can be displayed and viewed.
[0694] This system improves work efficiency by making it easy to search and extract necessary documents even when documents are not properly handed over after an employee leaves or is transferred. It also supports not only text but also images and other data formats, allowing for quick and efficient searches of various documents. This provides an environment that supports continuous business operations without compromising internal knowledge.
[0695] The processing flow will be explained below.
[0696] Step 1: Collect data
[0697] The server collects data from the company's email server, messaging platform, and cloud storage. From the email server, it collects email text, attachments, and metadata. From the messaging platform, it obtains chat logs and uploaded files. It also collects various document files (PDF, Word, image files, etc.) from cloud storage.
[0698] Step 2: Preprocessing the data
[0699] The server preprocesses the collected data by format. For text data, it analyzes it using natural language processing (NLP) technology and removes noise. For image data, it uses optical character recognition (OCR) technology to extract text information from the image. PDFs and Word documents are also converted into text data and their contents are analyzed.
[0700] Step 3: Building an AI model
[0701] The server then builds a multimodal AI model based on the preprocessed data. This model is trained to comprehensively analyze text, images, and other data formats and extract data characteristics, resulting in a model capable of similarity judgment.
[0702] Step 4: Receiving a search query
[0703] The terminal receives a search query from the user. For example, the user inputs a search query such as "Marketing strategy for new product Y." The terminal then sends this query to the server.
[0704] Step 5: Processing the search query
[0705] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's database. It performs a query-based search and ranks and extracts highly relevant materials. It performs an integrated search of materials, including not only text data but also different data formats such as images and PDFs.
[0706] Step 6: Search past communication logs
[0707] The server also searches past communication logs (emails and messages on messaging platforms) for related materials, extracting relevant email text and Slack messages, and searching for information from the communication logs.
[0708] Step 7: Serving search results
[0709] The server sends the search results and information extracted from past communication logs to the terminal. The terminal displays this data on a user interface, allowing the user to check each piece of information. For example, a report on the marketing strategy for new product Y and related communication records may be displayed.
[0710] Step 8: User validation and utilization
[0711] Users can check the search results displayed on their devices, select and view related materials as needed, and download or print any information they deem necessary for further use.
[0712] In this way, the system quickly searches for necessary information from vast amounts of internal company data, improving business efficiency.
[0713] Example 1
[0714] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0715] With conventional systems, it was difficult to efficiently collect and analyze data dispersed across internal emails, messaging platforms, and cloud storage, making it impossible to quickly search and retrieve the information needed. Furthermore, there was a lack of technology for comprehensively analyzing different data formats (text data, image data, PDF files, etc.), which led to problems that reduced operational efficiency. Furthermore, there was a need for a system that could include past communication logs in search results, allowing users to check the history of related communications.
[0716] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0717] In this invention, the server includes: means for collecting data from mail servers, messaging platforms, and cloud storage; means for analyzing text data, image data, and other data in preprocessing and integrating them; means for analyzing text data using natural language processing technology; means for extracting text information from image data using optical character recognition technology; means for building a multimodal AI model and analyzing data similarity; means for receiving search queries from users and identifying relevant data based on the search queries; means for searching past communication logs for sending relevant materials; and means for providing search results to user terminals. This enables centralized collection and analysis of distributed data within a company, allowing users to quickly and efficiently search for and obtain the information they need. It also enables integrated analysis of different data formats and provides comprehensive search results that include related communication history.
[0718] A "mail server" is a server that manages the sending and receiving of emails. Its role is to provide mailboxes, store sent emails, and forward received emails to users.
[0719] A "messaging platform" refers to a communication tool that provides functions such as text messaging, file sharing, voice calls, and video calls. Typical examples include online chat services and business communication tools.
[0720] "Cloud storage" refers to remote data storage services available over the Internet, allowing users to store data in the cloud and access, share, and manage it as needed.
[0721] "Data collection tools" refers to the technologies and processes used to extract and centralize data from different sources. This tool allows data from various platforms to be managed in a unified manner.
[0722] "Preprocessing" refers to a series of operations performed to convert raw data into a suitable format for data analysis and machine learning, including text segmentation, noise removal, and format conversion.
[0723] "Natural language processing technology" refers to technology that enables computers to understand, analyze, and generate human language, including text analysis, machine translation, and sentiment analysis.
[0724] Optical character recognition technology is a technology that analyzes characters in images captured by a scanner or camera and converts them into text data. It is used to digitize printed documents and handwritten characters.
[0725] A "multimodal AI model" is an artificial intelligence model that can integrate and analyze data in different formats (text, images, audio, etc.). Combining multiple data sources enables more accurate analysis.
[0726] A "search query" is a question or keyword that a user enters to search for specific information. It is used to refer to information in a search engine or database.
[0727] "Means for identifying relevant data" refers to techniques and processes for finding and extracting relevant data based on a search query from a user.
[0728] "Past communication logs" refers to records of past communications, such as email sending and receiving history and chat history on messaging platforms.
[0729] "Search Results" means a list of relevant data or information based on a user's search query, generated by a search engine or database and provided to the user.
[0730] "User terminal" refers to a device such as a computer, smartphone, or tablet that is directly operated by a user. It is a device that connects to a server via the Internet or an internal company network and allows users to use services.
[0731] The present invention is a system for collecting data from an in-house mail server, a messaging platform, and cloud storage, and analyzing and searching the data. Specific embodiments of this system are described below.
[0732] Hardware and Software Configuration
[0733] The server includes the following main software and hardware:
[0734] A scheduler (such as cron) to periodically run scripts (such as Python) for data collection
[0735] IMAP protocol library for connecting to mail servers
[0736] API client libraries for messaging platforms and cloud storage (e.g., Slack API, Google Drive API)
[0737] Libraries for using natural language processing technologies (NLTK and spaCy)
[0738] Library for using optical character recognition technology (Tesseract-OCR)
[0739] Libraries for building and training multimodal AI models (TensorFlow and PyTorch)
[0740] The terminals include computers, smartphones, tablets, etc. that are directly operated by the user and allow the user to access the system's search interface through a browser.
[0741] Data collection
[0742] The server periodically runs a data collection script that connects to the company's mail server, messaging platform, and cloud storage, collecting emails, chat logs, uploaded files, and various documents. The server then stores this data in local storage.
[0743] Data Preprocessing
[0744] The server preprocesses the collected data. Specifically, it uses natural language processing technology to analyze the text data and remove unnecessary information. Examples include removing stop words and stemming. For image data, it uses optical character recognition technology to extract text information, and converts PDFs and Word documents into text data. This integrates all data into an analyzable format.
[0745] Building a multimodal AI model
[0746] The server integrates preprocessed text data, image data, and other data formats into a single dataset. A multimodal AI model is built using TensorFlow and PyTorch to learn the characteristics of the data. This model can integrate and analyze different data formats, enabling more accurate searches.
[0747] User search and results display
[0748] A user accesses the system's search interface from a browser on their device and enters a search query, such as "marketing strategy for new product Y." The device then sends this search query to the server.
[0749] The server analyzes the received search query and uses a multimodal AI model to identify related materials. It also searches related past communication logs (including email and chat history) and displays them in an integrated manner, making it easier for users to grasp related information.
[0750] Providing search results
[0751] The server provides search results in JSON format to the terminal, which then displays them in HTML format, allowing the user to quickly and efficiently obtain the information they need. For example, past reports on marketing strategies for new product Y, related presentation materials, and meeting notes can be displayed.
[0752] Examples and prompts
[0753] For example, when searching using the prompt phrase "Marketing strategy for new product Y," the system quickly searches for relevant internal documents and past communication history and displays them to the user, allowing the user to carry out their work more efficiently.
[0754] In this way, this system provides specific methods and technologies for efficiently collecting and analyzing huge amounts of data and quickly responding to user search requests.
[0755] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0756] Program processing flow
[0757] Step 1: Collect data
[0758] The server periodically runs data collection scripts to connect to the company's mail server, messaging platform, and cloud storage.
[0759] Input: Connection and authentication information configured for the server
[0760] How it works: It connects to mail servers using the IMAP protocol, and to messaging platforms and cloud storage using the APIs of each service.
[0761] Output: Save retrieved emails, chat logs, and various documents to local storage.
[0762] Step 2: Preprocessing the data
[0763] The server pre-processes the collected data.
[0764] Input: Collected emails, chat logs, document files
[0765] How it works: It uses natural language processing (NLP) techniques to analyze text data and remove unnecessary information. Specifically, it performs text segmentation, stop word removal, and stemming. It also uses Tesseract-OCR to perform character recognition on image data, and PDF and Word documents are converted to text using PDFMiner and python-docx.
[0766] Output: Preprocessed text data, image data, and other data are integrated and saved.
[0767] Step 3: Building a multimodal AI model
[0768] The server uses the integrated data to build a multimodal AI model.
[0769] Input: Preprocessed dataset
[0770] How it works: We use TensorFlow and PyTorch to build models and train them to learn data features, specifically by integrating text, images, and other data formats and determining data similarity.
[0771] Output: A trained multimodal AI model
[0772] Step 4: Receiving and parsing the search query
[0773] A user accesses the search interface from a browser on the terminal and enters a search query.
[0774] Input: The search query entered by the user (e.g., "Marketing strategy for new product Y")
[0775] Operation: The terminal sends this query to the server, which analyzes the received query.
[0776] Output: Parsed search query
[0777] Step 5: Search and identify relevant data
[0778] The server uses the parsed search query to search for relevant data.
[0779] Input: Parsed search query, trained multimodal AI model
[0780] How it works: Uses a multimodal AI model to identify and rank relevant materials and data based on a search query.
[0781] Output: A list of related data
[0782] Step 6: Search past communication logs
[0783] The server searches past communication logs based on the associated data.
[0784] Input: List of related data
[0785] How it works: Uses email and messaging platform APIs to search and extract relevant past communication logs.
[0786] Output: Related past communication logs
[0787] Step 7: Serving search results
[0788] The server integrates the relevant data with past communication logs to generate search results.
[0789] Input: List of related data, past communication logs
[0790] What it does: Generates a unified search result in JSON format and sends it to the device.
[0791] Output: JSON data of search results
[0792] Step 8: Viewing search results
[0793] The terminal displays the received search results to the user.
[0794] Input: JSON data of search results sent from the server
[0795] What it does: Converts the results to HTML and displays them in the browser, allowing the user to view the results.
[0796] Output: Search results displayed to the user
[0797] This series of processes allows the user to efficiently search for and view the information they need.
[0798] (Application example 1)
[0799] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0800] While conventional in-house data search systems can collect and preprocess data from email servers and messaging platforms, they face the problem of making it difficult for store staff to quickly and efficiently obtain the information they need. In particular, there is a need for a way to quickly obtain information in real time, such as inventory information and past sales history, while serving customers in stores. Furthermore, there is a lack of systems that can handle multiple data formats and perform integrated analysis. Therefore, an effective solution is needed to improve the work efficiency of store staff.
[0801] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0802] In this invention, the server includes means for collecting data from mail servers, messaging platforms, and cloud storage, means for analyzing text data, image data, and other data in preprocessing and integrating them, means for building a multimodal AI model and analyzing data similarity, means for receiving search queries from users and identifying related data based on the search queries, means for searching past communication logs for sending related materials, means for providing search results to user terminals, and means for in-store staff to send queries by voice input and display search results through smart glasses, enabling store staff to obtain information in real time and respond to customers quickly and efficiently.
[0803] A "mail server" is a computer server that manages and stores the sending and receiving of email.
[0804] A "messaging platform" is an online communication tool for sending and receiving text messages and files in real time.
[0805] "Cloud storage" is an online storage service for storing, managing, and accessing data via the Internet.
[0806] "Preprocessing" refers to a data processing step to convert collected data into a format that is easy to analyze.
[0807] "Text data" is information expressed as characters or sentences.
[0808] "Image data" is information expressed in an image format.
[0809] "Integration" means combining multiple different data formats and information into one format and making them consistent.
[0810] A "multimodal AI model" is an artificial intelligence model that performs integrated analysis of multiple data formats (text, images, etc.).
[0811] "Data similarity" refers to common characteristics or patterns between different data.
[0812] A "search query" is a search question or keyword entered by a user to search for specific information.
[0813] "Related Materials" are useful information or documents identified based on the search query.
[0814] "Past communication logs" are records of messages and emails previously exchanged.
[0815] A "user terminal" is an electronic device that a user operates and receives information from, such as a smartphone or tablet.
[0816] A "brick and mortar store" is a sales facility that has a physical presence.
[0817] "Voice input" is a method of inputting characters and commands using voice.
[0818] "Smart glasses" are a wearable eyeglass-type device equipped with a display function.
[0819] The present invention is a system that enables in-store staff to quickly and efficiently obtain necessary information while serving customers. The system collects data from mail servers, messaging platforms, and cloud storage, and analyzes them in an integrated manner to provide relevant data based on user search queries.
[0820] 1. Program Overview
[0821] The server operates programs to achieve the following functions:
[0822] Data Collection: Periodically collect data from mail servers, messaging platforms, and cloud storage.
[0823] Pre-processing: Analyze collected data using NLP techniques (e.g., spaCy) and OCR techniques (e.g., Tesseract) to extract and convert text and image data.
[0824] Building multimodal AI models: Multimodal AI models can be trained based on text data, image data, and other data formats to determine data similarity.
[0825] Search query processing: Receives and analyzes search queries voice-entered by store staff through smart glasses.
[0826] Sending search results: Search results for related materials and past communication logs are displayed on the smart glasses.
[0827] 2. Hardware and Software
[0828] Hardware used: Server, smart glasses, user devices (e.g., smartphones and tablets).
[0829] Software used: NLP engine (spaCy), OCR engine (Tesseract), multimodal AI model (OpenAI's GPT-4).
[0830] 3. System operation explanation
[0831] The server first collects the necessary data from various data sources. The collected data undergoes a preprocessing step to convert it into a format that is easy to analyze. During this process, NLP and OCR technologies are used to extract text data, and image data is also handled. After preprocessing is complete, the data is analyzed comprehensively using a multimodal AI model to extract data characteristics and similarities.
[0832] The user, a store staff member, voice-inputs queries such as customer questions or inventory checks. For example, prompts such as "Check inventory, tell me the availability of product A" or "Search customer X's purchase history" are sent through the smart glasses. The server receives the query, searches relevant data using a multimodal AI model, and ranks and extracts the results. These results are then provided to the smart glasses, allowing the user to view the information in real time.
[0833] 4. Specific Examples
[0834] For example, if a store staff member sends a voice prompt such as "Tell me the purchase history of customer Yamada Taro for the past three months," the system will quickly collect relevant past communication logs, inventory information, and sales data and display them on the smart glasses' display. It also works in the same way with prompts such as "Show me the marketing materials for new product A that will be on sale next weekend," supporting efficient customer service.
[0835] The present invention allows staff in physical stores to obtain information in real time, enabling them to respond to customers quickly and accurately.
[0836] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0837] Step 1:
[0838] Data collection
[0839] The server periodically connects to mail servers, messaging platforms, and cloud storage to collect data, including email bodies and attachments, chat logs and uploaded files from messaging platforms, and various files from cloud storage.
[0840] Input: Each data source (mail server, messaging platform, cloud storage)
[0841] Output: Raw data collected (emails, chat logs, files)
[0842] Specifically, the server accesses these data sources using APIs or dedicated protocols, retrieves the latest data, and stores it.
[0843] Step 2:
[0844] Data Preprocessing
[0845] The server preprocesses the collected data. During this preprocessing step, the server analyzes the text data using natural language processing (NLP) technology, removes unnecessary information, and converts it into a format that is easier to analyze. For image data, the server also uses optical character recognition (OCR) technology to extract text information from the image. Furthermore, the server analyzes files such as PDFs and Word documents and converts them into text data.
[0846] Input: Raw data collected
[0847] Output: Preprocessed data (text format)
[0848] Specifically, the system uses an NLP engine (e.g., spaCy) to analyze text and remove unnecessary parts, and an OCR engine (e.g., Tesseract) to extract text information from the image.
[0849] Step 3:
[0850] Building a multimodal AI model
[0851] The server uses the preprocessed data to build a multimodal AI model that can comprehensively analyze text, images, and other data formats, and trains the model to determine data similarity.
[0852] Input: Preprocessed data
[0853] Output: A trained multimodal AI model
[0854] Specifically, it feeds data to AI models (such as OpenAI's GPT-4) to train them, and also continuously updates and improves them.
[0855] Step 4:
[0856] Receiving a search query
[0857] The user (store staff) sends a search query by voice input through the smart glasses, for example, "Check inventory, tell me the stock status of product X." The device converts this voice data into text format and sends it to the server.
[0858] Input: Voice query
[0859] Output: Text query
[0860] Specifically, it converts voice input into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API).
[0861] Step 5:
[0862] Finding related data
[0863] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's database, ranking and extracting the most relevant materials.
[0864] Input: Text query, trained multimodal AI model, in-house database
[0865] Output: Search results (list of related materials)
[0866] Specifically, the server inputs a query into the model, searches the database for materials that match the query, and retrieves the results.
[0867] Step 6:
[0868] Submitting and viewing search results
[0869] The server sends the search results to the user's device (smart glasses), and the user can view the displayed search results, such as "Report on marketing strategy for new product X" or "Related meeting notes."
[0870] Input: Search results
[0871] Output: Results displayed on smart glasses
[0872] Specifically, the server formats the search results and transmits them in a format suitable for the smart glasses' display.
[0873] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0874] This invention is a system that collects data from an in-house mail server, messaging platform, and cloud storage, and combines it with an emotion engine that recognizes the user's emotions to perform more effective document searches. Below, the program processing of this system is explained in natural language.
[0875] First, the server periodically connects to the company's internal mail server, messaging platforms (e.g., Slack and Microsoft Teams), and cloud storage (e.g., Google Drive and Dropbox) to collect all data, including email bodies and attachments, chat logs and uploaded files from messaging platforms, and various files from cloud storage.
[0876] Next, the server preprocesses the collected data. During this preprocessing step, the server analyzes the text data using natural language processing (NLP) technology, removes unnecessary information, and converts it into a format that is easier to analyze. For image data, the server also uses optical character recognition (OCR) technology to extract text information from the image. Furthermore, the server analyzes files such as PDFs and Word documents and converts them into text data.
[0877] The server then builds a multimodal AI model based on the preprocessed data (text data, image data, and other data formats). This AI model can comprehensively analyze text, images, and other data formats to extract data characteristics. This model is then trained to determine data similarity.
[0878] When a user searches for specific information on a device, they input a search query. For example, they input a query such as "Marketing strategy for new product Y." The device then sends this search query to the server. At this stage, the device can also use an emotion engine to detect the user's emotions. The emotion engine recognizes the user's facial expressions and voice and determines the user's current emotional state.
[0879] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's internal database. It then performs a search based on the query and ranks and extracts the most relevant materials. It performs an integrated search of materials that include not only text data but also different data formats such as images and PDFs. Furthermore, it adjusts the ranking of search results by taking into account the user's emotional state detected by the emotion engine. For example, if the user is feeling stressed, it will prioritize displaying materials that are more concise and easy to understand.
[0880] The server also searches past related communication logs. For example, if the relevant document was sent by email, it searches the email body and attachments and extracts related Slack messages. This allows users to view the communication logs related to the document along with the related documents.
[0881] The server sends the search results to the terminal and displays them to the user. The user can check the displayed search results and select and view the necessary materials. For example, past reports on the marketing strategy for new product Y, presentation materials, related meeting notes, etc. are displayed and can be reviewed. In addition, an emotion engine is used to provide information tailored to the user's emotions, improving the user experience.
[0882] This system improves work efficiency by making it easy to search and retrieve necessary documents even when documents are not properly handed over after an employee leaves or is transferred. It also supports not only text but also images and other data formats, allowing for quick and efficient searches of various documents. Furthermore, by utilizing an emotion engine, it is possible to provide search results that reflect the user's emotional state, improving user satisfaction.
[0883] The processing flow will be explained below.
[0884] Step 1: Collect data
[0885] The server periodically connects to the company's email server, messaging platform, and cloud storage to collect data. Specifically, it collects email text, attachments, and metadata from the email server, chat logs and uploaded files from the messaging platform, and various document files (PDF, Word, image files, etc.) from the cloud storage.
[0886] Step 2: Preprocessing the data
[0887] The server preprocesses the collected data by format. For text data, it analyzes it using natural language processing (NLP) technology and removes noise. For image data, it uses optical character recognition (OCR) technology to extract text information from the image. PDFs and Word documents are also converted into text data and their contents are analyzed.
[0888] Step 3: Building an AI model
[0889] The server then builds a multimodal AI model based on the preprocessed data. This model is trained to comprehensively analyze text, images, and other data formats and extract data characteristics, resulting in a model capable of similarity judgment.
[0890] Step 4: Receiving a search query
[0891] The device receives a search query from the user. For example, the user may enter a search query such as "Marketing strategy for new product Y." The device sends this query to the server and simultaneously activates an emotion engine to detect the user's emotional state.
[0892] Step 5: Detecting Emotional State
[0893] The device uses an emotion engine to analyze the user's emotional state. Specifically, it recognizes emotions from the user's facial expressions and voice and determines the user's current emotional state. For example, it detects whether the user is feeling stressed or relaxed.
[0894] Step 6: Processing the search query
[0895] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's database. It performs a query-based search and ranks and extracts highly relevant materials. It performs an integrated search of materials, including not only text data but also different data formats such as images and PDFs.
[0896] Step 7: Adjust your results based on your emotional state
[0897] The server adjusts the ranking of search results based on the user's emotional state as detected by the emotion engine. For example, if the user is feeling stressed, it will prioritize more concise and easy-to-understand materials. On the other hand, if the user is relaxed, it will provide more detailed and specialized materials.
[0898] Step 8: Search past communication logs
[0899] The server also searches past communication logs (emails and messages on messaging platforms) for related materials, extracting relevant email text, Slack messages, etc., and searching for information from the communication logs.
[0900] Step 9: Serving search results
[0901] The server sends the search results and information extracted from past communication logs to the terminal. The terminal displays this data on a user interface, allowing the user to check each piece of information. For example, a report on the marketing strategy for new product Y and related communication records may be displayed.
[0902] Step 10: User validation and utilization
[0903] Users can check the search results displayed on their devices, select and view related materials as needed, and download or print any information they deem necessary for further use.
[0904] In this way, this system aims to improve work efficiency and user satisfaction by quickly searching for necessary information from vast amounts of internal company data and providing information that corresponds to the user's emotional state.
[0905] Example 2
[0906] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0907] Conventional document search systems struggled to effectively collect and integrate data from internal email servers, messaging platforms, and cloud storage, and lacked a way to appropriately adjust search results based on the user's emotional state. This resulted in problems such as an inability to respond quickly and accurately to user requests, leading to reduced work efficiency. Furthermore, it was difficult to comprehensively search documents in different data formats, making it difficult to find the information needed when document handover was insufficient.
[0908] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0909] In this invention, the server includes: means for collecting data from mail servers, messaging platforms, and cloud storages; means for analyzing text data, image data, and other data using natural language processing technology and optical character recognition technology in preprocessing and integrating them; means for building a multimodal AI model including text, image, and other data formats and analyzing the similarity of the data; means for receiving a user's search query and identifying related data based on the search query using an emotion engine that detects the user's emotional state; means for searching past communication logs to send related materials; and means for providing the search results to the user terminal.
[0910] This makes it possible to centrally collect and integrate data from various data sources within the company and provide appropriate search results based on user sentiment. Users can quickly find the documents they need, improving work efficiency. In addition, by searching documents in different data formats in an integrated manner, related information can be easily obtained even if document handover is insufficient.
[0911] A "mail server" is a server that manages the sending and receiving of email within an organization or on the Internet.
[0912] A "messaging platform" is software or a service for sending and receiving messages in real time.
[0913] "Cloud storage" is an online service for storing and managing data over the Internet.
[0914] "Natural language processing technology" is a technology that analyzes text data and extracts linguistic features.
[0915] "Optical character recognition technology" is a technology that extracts text information from image data.
[0916] A "multimodal AI model" is an artificial intelligence model that comprehensively analyzes text, images, and other data formats to extract features.
[0917] "User emotional state" refers to the emotional state a user is experiencing when performing a search, and includes techniques for detecting this.
[0918] An "emotion engine" is a device or software that analyzes a user's facial expressions and voice to detect their emotional state.
[0919] A "search query" is text data indicating a search request that a user inputs into a search system.
[0920] "Communication logs" are data containing past communication history across email and messaging platforms.
[0921] The present invention is a system that collects data from an internal mail server, a messaging platform, and cloud storage, and combines it with an emotion engine that recognizes the user's emotional state to perform more effective document search. The program processing of this system is described below.
[0922] First, the server periodically connects to the company's email server, messaging platform, and cloud storage to collect all data. Specifically, Microsoft Exchange and Gmail are used as email servers, Slack and Microsoft Teams are used as messaging platforms, and Google Drive and Dropbox are used as cloud storage. The server uses APIs to collect email bodies and attachments, chat logs, uploaded files, and various files in cloud storage.
[0923] The server then preprocesses the collected data. This preprocessing step uses natural language processing (NLP) and optical character recognition (OCR) techniques. For text data, NLTK and spaCy are used to tokenize and remove stop words. For image data, Tesseract OCR is used to extract text information from images. For PDF and Word documents, Apache Tika is used to convert them to text data.
[0924] The server then builds a multimodal AI model based on the preprocessed data. This model is created using TensorFlow and PyTorch to comprehensively analyze text, images, and other data formats and extract data characteristics. Specifically, it uses a convolutional neural network (CNN) to extract image features, and a recurrent neural network (RNN) or BERT model to extract text features.
[0925] Next, the user inputs a search query from the terminal. For example, the query "Marketing strategy for new product Y" is input. This search query is sent as is from the terminal to the server as a character string.
[0926] The device is equipped with an emotion engine that detects the user's emotional state, for example by capturing and analyzing facial expressions with a camera. This analysis is performed using the Face API from Microsoft Azure's Cognitive Services. The emotion engine detects the user's emotional state and transmits that information along with a query to the server.
[0927] Based on the received search query and emotional information, the server uses a multimodal AI model to identify relevant materials from the company's database. It then assesses data similarity and ranks and extracts highly relevant materials. It also adjusts the ranking of results according to the user's emotional state. If the user is feeling stressed, the algorithm prioritizes displaying more concise and easy-to-understand materials.
[0928] The server also searches past communication logs for related materials. For example, if the material was sent by email, it extracts the email body, attachments, and chat logs from related messaging platforms. Finally, the server sends the search results and related communication logs to the device.
[0929] Users can select and view the necessary materials from the results displayed on their device. For example, they can check past reports, presentation materials, meeting notes, etc. related to "Marketing Strategy for New Product Y."
[0930] As a concrete example, here is a prompt example for a generative AI model when searching for "marketing strategy for new product Y":
[0931] "Search your internal database for materials related to the marketing strategy for new product Y. If the user is feeling stressed, prioritize materials that are more concise and easy to understand."
[0932] This system makes it possible to quickly and efficiently search and retrieve necessary documents even when documents are not fully handed over when an employee resigns or is transferred.In addition, by providing search results that respond to the user's emotions, it is expected that the user experience will improve and business efficiency will increase.
[0933] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0934] Step 1:
[0935] The server periodically connects to the company's mail server, messaging platform, and cloud storage to collect data. In this step, it retrieves email bodies and attachments from the mail server, chat logs and uploaded files from the messaging platform, and various files from the cloud storage. The specific input is a request using the API of each data source, and the output is a collection of collected data.
[0936] Step 2:
[0937] The server preprocesses the collected data. It uses natural language processing (NLP) and optical character recognition (OCR) techniques to tokenize text data, remove stop words, extract text information from image data, and convert PDFs and Word documents into text data. Specific operations include text analysis using NLTK and spaCy, image analysis using Tesseract OCR, and document analysis using Apache Tika. The input is the raw data collected in step 1, and the output is the preprocessed, clean data.
[0938] Step 3:
[0939] The server builds a multimodal AI model based on the preprocessed data. It uses TensorFlow and PyTorch to comprehensively analyze text, images, and other data formats and extract features. Specifically, it uses a Convolutional Neural Network (CNN) to extract image features, and a Recurrent Neural Network (RNN) or BERT model to extract text data features. The input is various preprocessed data, and the output is a trained AI model.
[0940] Step 4:
[0941] A user inputs a query to search for specific information from a terminal. For example, the query might be "Marketing strategy for new product Y." The query is sent directly from the terminal to the server in the form of a string. The input is the query entered by the user, and the output is a search request sent to the server.
[0942] Step 5:
[0943] The device uses an emotion engine to detect the user's emotional state. Specifically, it captures the user's facial expressions with a camera and analyzes them using the Face API of Microsoft Azure's Cognitive Services. The input is an image of the user's facial expression, and the output is the detected emotional state data.
[0944] Step 6:
[0945] The server uses a multimodal AI model to search for relevant materials from the company's database based on the received search query and emotional information. It ranks the search results and adjusts the ranking based on the emotion engine's evaluation. Specifically, if the user is feeling stressed, the server uses an algorithm that prioritizes displaying concise and easy-to-understand materials. The input is the search query and emotional information, and the output is ranked search results.
[0946] Step 7:
[0947] The server also searches past communication logs for related materials. For example, if the material was sent by email, it extracts the email body, attachments, and chat logs from related messaging platforms. The input is the search results, and the output is the related communication logs.
[0948] Step 8:
[0949] The server sends the search results and related communication logs to the terminal. The user can select and view the necessary materials from the results displayed on the terminal. The input is the ranked search results and related communication logs, and the output is the information displayed on the terminal in a user-accessible form.
[0950] Through these steps, users can quickly and efficiently search and extract the information they need.
[0951] (Application example 2)
[0952] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0953] Providing timely and appropriate information is essential to improving work efficiency in factories and reducing worker stress. However, conventional information search systems were unable to provide information tailored to the emotional state of workers, making it difficult to provide efficient work support. Furthermore, extracting information from a wide variety of data formats and past communication logs was time-consuming, making it difficult to provide practical information.
[0954] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0955] In this invention, the server includes: means for collecting data from mail servers, messaging platforms, and cloud storage; means for analyzing and integrating text data, image data, and other data in preprocessing; means for building a multimodal AI model and analyzing data similarity; means for receiving search queries from users and identifying relevant data based on the search queries; means for analyzing the emotional state of workers and adjusting search results based on the results; means for searching past communication logs for sending relevant materials; and means for providing search results to user terminals. This enables timely and appropriate information provision that takes emotional states into account. Furthermore, the system can efficiently extract information from a variety of data formats and past communication logs, thereby providing efficient support for workers.
[0956] A "mail server" is a server for sending, receiving, and storing e-mail.
[0957] A "messaging platform" is a system for exchanging messages between users using text, voice, video, etc.
[0958] "Cloud storage" refers to online storage services for storing and accessing data over the Internet.
[0959] "Means for collecting data" refers to the mechanisms for obtaining the necessary data from mail servers, messaging platforms, and cloud storage.
[0960] "Preprocessing" is the process of converting collected data into a format that is easy to analyze and removing unnecessary information.
[0961] "Text data" is a data format that includes character information.
[0962] "Image data" is a data format that contains visual information.
[0963] A "multimodal AI model" is an artificial intelligence model for integrated analysis of different data formats (text, images, etc.).
[0964] A "user search query" is a search request entered by a user to locate specific information.
[0965] The "emotional state" is a state that indicates the type and intensity of the emotion that the user is currently feeling.
[0966] "Related materials" are materials that contain information that is useful to the user and that are identified based on the search query.
[0967] A "communication log" is a record of messages and files sent and received in the past, as well as related information.
[0968] "Search Results" means a list of relevant materials identified based on a search query.
[0969] A "user terminal" is a device that a user operates directly to check information.
[0970] To realize this application example, the system program is configured as follows.
[0971] The server collects data from email servers, messaging platforms, and cloud storage. Specifically, it uses the Gmail API, Microsoft Teams API, Google Drive API, and other sources to obtain the necessary data. The obtained data is in the form of text, image data, and other data formats, and natural language processing (NLP) and optical character recognition (OCR) technologies are used to preprocess it. After preprocessing, the data is analyzed using Hugging Face's transformers library to build a multimodal AI model. The multimodal AI model comprehensively analyzes text, image data, and other data formats.
[0972] The server also receives search queries from users and identifies relevant data based on the search queries. A user can input a search query such as "How to troubleshoot error code E123 on machine X" using a device (e.g., a smartphone or smart glasses). The server also analyzes the user's emotional state in real time using a device equipped with an emotion engine (e.g., smart glasses with a built-in camera and microphone that analyzes facial expressions and voice). If the user is feeling stressed, the emotion engine will prioritize providing more concise and easy-to-understand materials.
[0973] The search engine searches the database based on the query and also searches past related communication logs. For example, if related documents were sent via email or messaging platforms, the contents of those documents will also be searched. This allows for comprehensive retrieval of past communications and related documents. Search results are ranked according to the user's emotional state, with the most relevant information being displayed first.
[0974] The terminal provides the search results to the user, who can then check and view the necessary materials. This improves work efficiency and reduces worker stress. This system makes it possible to efficiently extract information from a wide range of data formats, significantly improving the factory work environment.
[0975] As a concrete example, consider a factory worker troubleshooting a new machine. If the worker types the query "How to troubleshoot error code E123 on machine X" into the smart glasses and is stressed, the application will prioritize short, easy-to-understand instructions.
[0976] Example prompt sentence:
[0977] "Find past troubleshooting documents for error code E123 on machine X. Workers are stressed right now, so please prioritize concise documents."
[0978] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0979] Step 1:
[0980] The server collects data from email servers, messaging platforms, and cloud storage. As input, it uses the Gmail API, Microsoft Teams API, Google Drive API, etc. to obtain the target email, message, and file data. The output is the collected text data, image data, and other data formats.
[0981] Step 2:
[0982] The server preprocesses the collected data. Specifically, it uses natural language processing (NLP) technology to analyze the text data, remove unnecessary information, and convert it into a format that is easier to analyze. It also uses optical character recognition (OCR) technology to extract text information from image data. It also analyzes files such as PDFs and Word documents and converts them into text data. The input is the data collected in step 1, and the output is the preprocessed data.
[0983] Step 3:
[0984] The server builds a multimodal AI model based on various preprocessed data (text data, image data, and other data formats). Using Hugging Face's transformers library, it trains an AI model that can comprehensively analyze different data formats. The input is the preprocessed data output from Step 2, and the output is an AI model that can analyze data similarities.
[0985] Step 4:
[0986] A user uses a device (e.g., a smartphone, smart glasses) to input a search query, such as "how to troubleshoot error code E123 on machine X." The input is the user's query, and the output is the query data.
[0987] Step 5:
[0988] The device sends the search query to the server. At the same time, the device uses an emotion engine to analyze the user's emotional state. For example, it uses the camera and microphone built into the smart glasses to analyze the user's facial expressions and voice to determine their current emotional state. The input is the query from step 4 and the user's emotional data, and the output is composite data including the query and the emotional state.
[0989] Step 6:
[0990] The server searches the database based on the search query, using a multimodal AI model to identify relevant information and also searches past communication logs to retrieve relevant materials. The input is the composite data from step 5, and the output is a list of relevant materials.
[0991] Step 7:
[0992] The server adjusts the ranking of the retrieved search results taking into account the user's emotional state. If the emotion engine detects that the user is stressed, it prioritizes displaying materials that are more concise and easy to understand. The input is the list of related materials from step 6 and the emotion analysis data, and the output is search results that have been ranked based on emotion.
[0993] Step 8:
[0994] The server sends the final search results to the terminal, where the user can check the search results on the terminal and view the necessary materials. The input is the search results ranked based on the sentiment in step 7, and the output is the search results displayed on the terminal.
[0995] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0996] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0997] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0998] [Fourth embodiment]
[0999] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1000] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1001] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1002] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1003] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1004] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1005] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1006] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1007] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1008] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1009] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1010] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1011] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1012] The present invention is a system for analyzing and searching data collected from an in-house mail server, a messaging platform, and cloud storage. The program processing of this system is explained below in natural language.
[1013] First, the server periodically connects to the company's internal mail server, messaging platforms (e.g., Slack and Microsoft Teams), and cloud storage (e.g., Google Drive and Dropbox) to collect all data, including email bodies and attachments, chat logs and uploaded files from messaging platforms, and various files from cloud storage.
[1014] Next, the server preprocesses the collected data. During this preprocessing step, the server analyzes the text data using natural language processing (NLP) technology, removes unnecessary information, and converts it into a format that is easier to analyze. For image data, the server also uses optical character recognition (OCR) technology to extract text information from the image. Furthermore, the server analyzes files such as PDFs and Word documents and converts them into text data.
[1015] The server then builds a multimodal AI model based on the preprocessed data (text data, image data, and other data formats). This AI model can comprehensively analyze text, images, and other data formats to extract data characteristics. This model is then trained to determine data similarity.
[1016] When a user searches for specific information on their device, they input a search query. For example, they input a query such as "Marketing strategy for new product Y." The device then sends this search query to a server. Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's internal database. The server then performs a search based on the query and ranks and extracts the most relevant materials.
[1017] The server also searches past related communication logs. For example, if the relevant document was sent by email, it searches the email body and attachments and extracts related Slack messages. This allows users to view the communication logs related to the document along with the related documents.
[1018] The server sends the search results to the terminal and displays them to the user. The user can check the displayed search results and select and view the necessary materials. For example, past reports on the marketing strategy for new product Y, presentation materials, related meeting notes, etc. can be displayed and viewed.
[1019] This system improves work efficiency by making it easy to search and extract necessary documents even when documents are not properly handed over after an employee leaves or is transferred. It also supports not only text but also images and other data formats, allowing for quick and efficient searches of various documents. This provides an environment that supports continuous business operations without compromising internal knowledge.
[1020] The processing flow will be explained below.
[1021] Step 1: Collect data
[1022] The server collects data from the company's email server, messaging platform, and cloud storage. From the email server, it collects email text, attachments, and metadata. From the messaging platform, it obtains chat logs and uploaded files. It also collects various document files (PDF, Word, image files, etc.) from cloud storage.
[1023] Step 2: Preprocessing the data
[1024] The server preprocesses the collected data by format. For text data, it analyzes it using natural language processing (NLP) technology and removes noise. For image data, it uses optical character recognition (OCR) technology to extract text information from the image. PDFs and Word documents are also converted into text data and their contents are analyzed.
[1025] Step 3: Building an AI model
[1026] The server then builds a multimodal AI model based on the preprocessed data. This model is trained to comprehensively analyze text, images, and other data formats and extract data characteristics, resulting in a model capable of similarity judgment.
[1027] Step 4: Receiving a search query
[1028] The terminal receives a search query from the user. For example, the user inputs a search query such as "Marketing strategy for new product Y." The terminal then sends this query to the server.
[1029] Step 5: Processing the search query
[1030] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's database. It performs a query-based search and ranks and extracts highly relevant materials. It performs an integrated search of materials, including not only text data but also different data formats such as images and PDFs.
[1031] Step 6: Search past communication logs
[1032] The server also searches past communication logs (emails and messages on messaging platforms) for related materials, extracting relevant email text and Slack messages, and searching for information from the communication logs.
[1033] Step 7: Serving search results
[1034] The server sends the search results and information extracted from past communication logs to the terminal. The terminal displays this data on a user interface, allowing the user to check each piece of information. For example, a report on the marketing strategy for new product Y and related communication records may be displayed.
[1035] Step 8: User validation and utilization
[1036] Users can check the search results displayed on their devices, select and view related materials as needed, and download or print any information they deem necessary for further use.
[1037] In this way, the system quickly searches for necessary information from vast amounts of internal company data, improving business efficiency.
[1038] Example 1
[1039] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1040] With conventional systems, it was difficult to efficiently collect and analyze data dispersed across internal emails, messaging platforms, and cloud storage, making it impossible to quickly search and retrieve the information needed. Furthermore, there was a lack of technology for comprehensively analyzing different data formats (text data, image data, PDF files, etc.), which led to problems that reduced operational efficiency. Furthermore, there was a need for a system that could include past communication logs in search results, allowing users to check the history of related communications.
[1041] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1042] In this invention, the server includes: means for collecting data from mail servers, messaging platforms, and cloud storage; means for analyzing text data, image data, and other data in preprocessing and integrating them; means for analyzing text data using natural language processing technology; means for extracting text information from image data using optical character recognition technology; means for building a multimodal AI model and analyzing data similarity; means for receiving search queries from users and identifying relevant data based on the search queries; means for searching past communication logs for sending relevant materials; and means for providing search results to user terminals. This enables centralized collection and analysis of distributed data within a company, allowing users to quickly and efficiently search for and obtain the information they need. It also enables integrated analysis of different data formats and provides comprehensive search results that include related communication history.
[1043] A "mail server" is a server that manages the sending and receiving of emails. Its role is to provide mailboxes, store sent emails, and forward received emails to users.
[1044] A "messaging platform" refers to a communication tool that provides functions such as text messaging, file sharing, voice calls, and video calls. Typical examples include online chat services and business communication tools.
[1045] "Cloud storage" refers to remote data storage services available over the Internet, allowing users to store data in the cloud and access, share, and manage it as needed.
[1046] "Data collection tools" refers to the technologies and processes used to extract and centralize data from different sources. This tool allows data from various platforms to be managed in a unified manner.
[1047] "Preprocessing" refers to a series of operations performed to convert raw data into a suitable format for data analysis and machine learning, including text segmentation, noise removal, and format conversion.
[1048] "Natural language processing technology" refers to technology that enables computers to understand, analyze, and generate human language, including text analysis, machine translation, and sentiment analysis.
[1049] Optical character recognition technology is a technology that analyzes characters in images captured by a scanner or camera and converts them into text data. It is used to digitize printed documents and handwritten characters.
[1050] A "multimodal AI model" is an artificial intelligence model that can integrate and analyze data in different formats (text, images, audio, etc.). Combining multiple data sources enables more accurate analysis.
[1051] A "search query" is a question or keyword that a user enters to search for specific information. It is used to refer to information in a search engine or database.
[1052] "Means for identifying relevant data" refers to techniques and processes for finding and extracting relevant data based on a search query from a user.
[1053] "Past communication logs" refers to records of past communications, such as email sending and receiving history and chat history on messaging platforms.
[1054] "Search Results" means a list of relevant data or information based on a user's search query, generated by a search engine or database and provided to the user.
[1055] "User terminal" refers to a device such as a computer, smartphone, or tablet that is directly operated by a user. It is a device that connects to a server via the Internet or an internal company network and allows users to use services.
[1056] The present invention is a system for collecting data from an in-house mail server, a messaging platform, and cloud storage, and analyzing and searching the data. Specific embodiments of this system are described below.
[1057] Hardware and Software Configuration
[1058] The server includes the following main software and hardware:
[1059] A scheduler (such as cron) to periodically run scripts (such as Python) for data collection
[1060] IMAP protocol library for connecting to mail servers
[1061] API client libraries for messaging platforms and cloud storage (e.g., Slack API, Google Drive API)
[1062] Libraries for using natural language processing technologies (NLTK and spaCy)
[1063] Library for using optical character recognition technology (Tesseract-OCR)
[1064] Libraries for building and training multimodal AI models (TensorFlow and PyTorch)
[1065] The terminals include computers, smartphones, tablets, etc. that are directly operated by the user and allow the user to access the system's search interface through a browser.
[1066] Data collection
[1067] The server periodically runs a data collection script that connects to the company's mail server, messaging platform, and cloud storage, collecting emails, chat logs, uploaded files, and various documents. The server then stores this data in local storage.
[1068] Data Preprocessing
[1069] The server preprocesses the collected data. Specifically, it uses natural language processing technology to analyze the text data and remove unnecessary information. Examples include removing stop words and stemming. For image data, it uses optical character recognition technology to extract text information, and converts PDFs and Word documents into text data. This integrates all data into an analyzable format.
[1070] Building a multimodal AI model
[1071] The server integrates preprocessed text data, image data, and other data formats into a single dataset. A multimodal AI model is built using TensorFlow and PyTorch to learn the characteristics of the data. This model can integrate and analyze different data formats, enabling more accurate searches.
[1072] User search and results display
[1073] A user accesses the system's search interface from a browser on their device and enters a search query, such as "marketing strategy for new product Y." The device then sends this search query to the server.
[1074] The server analyzes the received search query and uses a multimodal AI model to identify related materials. It also searches related past communication logs (including email and chat history) and displays them in an integrated manner, making it easier for users to grasp related information.
[1075] Providing search results
[1076] The server provides search results in JSON format to the terminal, which then displays them in HTML format, allowing the user to quickly and efficiently obtain the information they need. For example, past reports on marketing strategies for new product Y, related presentation materials, and meeting notes can be displayed.
[1077] Examples and prompts
[1078] For example, when searching using the prompt phrase "Marketing strategy for new product Y," the system quickly searches for relevant internal documents and past communication history and displays them to the user, allowing the user to carry out their work more efficiently.
[1079] In this way, this system provides specific methods and technologies for efficiently collecting and analyzing huge amounts of data and quickly responding to user search requests.
[1080] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1081] Program processing flow
[1082] Step 1: Collect data
[1083] The server periodically runs data collection scripts to connect to the company's mail server, messaging platform, and cloud storage.
[1084] Input: Connection and authentication information configured for the server
[1085] How it works: It connects to mail servers using the IMAP protocol, and to messaging platforms and cloud storage using the APIs of each service.
[1086] Output: Save retrieved emails, chat logs, and various documents to local storage.
[1087] Step 2: Preprocessing the data
[1088] The server pre-processes the collected data.
[1089] Input: Collected emails, chat logs, document files
[1090] How it works: It uses natural language processing (NLP) techniques to analyze text data and remove unnecessary information. Specifically, it performs text segmentation, stop word removal, and stemming. It also uses Tesseract-OCR to perform character recognition on image data, and PDF and Word documents are converted to text using PDFMiner and python-docx.
[1091] Output: Preprocessed text data, image data, and other data are integrated and saved.
[1092] Step 3: Building a multimodal AI model
[1093] The server uses the integrated data to build a multimodal AI model.
[1094] Input: Preprocessed dataset
[1095] How it works: We use TensorFlow and PyTorch to build models and train them to learn data features, specifically by integrating text, images, and other data formats and determining data similarity.
[1096] Output: A trained multimodal AI model
[1097] Step 4: Receiving and parsing the search query
[1098] A user accesses the search interface from a browser on the terminal and enters a search query.
[1099] Input: The search query entered by the user (e.g., "Marketing strategy for new product Y")
[1100] Operation: The terminal sends this query to the server, which analyzes the received query.
[1101] Output: Parsed search query
[1102] Step 5: Search and identify relevant data
[1103] The server uses the parsed search query to search for relevant data.
[1104] Input: Parsed search query, trained multimodal AI model
[1105] How it works: Uses a multimodal AI model to identify and rank relevant materials and data based on a search query.
[1106] Output: A list of related data
[1107] Step 6: Search past communication logs
[1108] The server searches past communication logs based on the associated data.
[1109] Input: List of related data
[1110] How it works: Uses email and messaging platform APIs to search and extract relevant past communication logs.
[1111] Output: Related past communication logs
[1112] Step 7: Serving search results
[1113] The server integrates the relevant data with past communication logs to generate search results.
[1114] Input: List of related data, past communication logs
[1115] What it does: Generates a unified search result in JSON format and sends it to the device.
[1116] Output: JSON data of search results
[1117] Step 8: Viewing search results
[1118] The terminal displays the received search results to the user.
[1119] Input: JSON data of search results sent from the server
[1120] What it does: Converts the results to HTML and displays them in the browser, allowing the user to view the results.
[1121] Output: Search results displayed to the user
[1122] This series of processes allows the user to efficiently search for and view the information they need.
[1123] (Application example 1)
[1124] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1125] While conventional in-house data search systems can collect and preprocess data from email servers and messaging platforms, they face the problem of making it difficult for store staff to quickly and efficiently obtain the information they need. In particular, there is a need for a way to quickly obtain information in real time, such as inventory information and past sales history, while serving customers in stores. Furthermore, there is a lack of systems that can handle multiple data formats and perform integrated analysis. Therefore, an effective solution is needed to improve the work efficiency of store staff.
[1126] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1127] In this invention, the server includes means for collecting data from mail servers, messaging platforms, and cloud storage, means for analyzing text data, image data, and other data in preprocessing and integrating them, means for building a multimodal AI model and analyzing data similarity, means for receiving search queries from users and identifying related data based on the search queries, means for searching past communication logs for sending related materials, means for providing search results to user terminals, and means for in-store staff to send queries by voice input and display search results through smart glasses, enabling store staff to obtain information in real time and respond to customers quickly and efficiently.
[1128] A "mail server" is a computer server that manages and stores the sending and receiving of email.
[1129] A "messaging platform" is an online communication tool for sending and receiving text messages and files in real time.
[1130] "Cloud storage" is an online storage service for storing, managing, and accessing data via the Internet.
[1131] "Preprocessing" refers to a data processing step to convert collected data into a format that is easy to analyze.
[1132] "Text data" is information expressed as characters or sentences.
[1133] "Image data" is information expressed in an image format.
[1134] "Integration" means combining multiple different data formats and information into one format and making them consistent.
[1135] A "multimodal AI model" is an artificial intelligence model that performs integrated analysis of multiple data formats (text, images, etc.).
[1136] "Data similarity" refers to common characteristics or patterns between different data.
[1137] A "search query" is a search question or keyword entered by a user to search for specific information.
[1138] "Related Materials" are useful information or documents identified based on the search query.
[1139] "Past communication logs" are records of messages and emails previously exchanged.
[1140] A "user terminal" is an electronic device that a user operates and receives information from, such as a smartphone or tablet.
[1141] A "brick and mortar store" is a sales facility that has a physical presence.
[1142] "Voice input" is a method of inputting characters and commands using voice.
[1143] "Smart glasses" are a wearable eyeglass-type device equipped with a display function.
[1144] The present invention is a system that enables in-store staff to quickly and efficiently obtain necessary information while serving customers. The system collects data from mail servers, messaging platforms, and cloud storage, and analyzes them in an integrated manner to provide relevant data based on user search queries.
[1145] 1. Program Overview
[1146] The server operates programs to achieve the following functions:
[1147] Data Collection: Periodically collect data from mail servers, messaging platforms, and cloud storage.
[1148] Pre-processing: Analyze collected data using NLP techniques (e.g., spaCy) and OCR techniques (e.g., Tesseract) to extract and convert text and image data.
[1149] Building multimodal AI models: Multimodal AI models can be trained based on text data, image data, and other data formats to determine data similarity.
[1150] Search query processing: Receives and analyzes search queries voice-entered by store staff through smart glasses.
[1151] Sending search results: Search results for related materials and past communication logs are displayed on the smart glasses.
[1152] 2. Hardware and Software
[1153] Hardware used: Server, smart glasses, user devices (e.g., smartphones and tablets).
[1154] Software used: NLP engine (spaCy), OCR engine (Tesseract), multimodal AI model (OpenAI's GPT-4).
[1155] 3. System operation explanation
[1156] The server first collects the necessary data from various data sources. The collected data undergoes a preprocessing step to convert it into a format that is easy to analyze. During this process, NLP and OCR technologies are used to extract text data, and image data is also handled. After preprocessing is complete, the data is analyzed comprehensively using a multimodal AI model to extract data characteristics and similarities.
[1157] The user, a store staff member, voice-inputs queries such as customer questions or inventory checks. For example, prompts such as "Check inventory, tell me the availability of product A" or "Search customer X's purchase history" are sent through the smart glasses. The server receives the query, searches relevant data using a multimodal AI model, and ranks and extracts the results. These results are then provided to the smart glasses, allowing the user to view the information in real time.
[1158] 4. Specific Examples
[1159] For example, if a store staff member sends a voice prompt such as "Tell me the purchase history of customer Yamada Taro for the past three months," the system will quickly collect relevant past communication logs, inventory information, and sales data and display them on the smart glasses' display. It also works in the same way with prompts such as "Show me the marketing materials for new product A that will be on sale next weekend," supporting efficient customer service.
[1160] The present invention allows staff in physical stores to obtain information in real time, enabling them to respond to customers quickly and accurately.
[1161] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1162] Step 1:
[1163] Data collection
[1164] The server periodically connects to mail servers, messaging platforms, and cloud storage to collect data, including email bodies and attachments, chat logs and uploaded files from messaging platforms, and various files from cloud storage.
[1165] Input: Each data source (mail server, messaging platform, cloud storage)
[1166] Output: Raw data collected (emails, chat logs, files)
[1167] Specifically, the server accesses these data sources using APIs or dedicated protocols, retrieves the latest data, and stores it.
[1168] Step 2:
[1169] Data Preprocessing
[1170] The server preprocesses the collected data. During this preprocessing step, the server analyzes the text data using natural language processing (NLP) technology, removes unnecessary information, and converts it into a format that is easier to analyze. For image data, the server also uses optical character recognition (OCR) technology to extract text information from the image. Furthermore, the server analyzes files such as PDFs and Word documents and converts them into text data.
[1171] Input: Raw data collected
[1172] Output: Preprocessed data (text format)
[1173] Specifically, the system uses an NLP engine (e.g., spaCy) to analyze text and remove unnecessary parts, and an OCR engine (e.g., Tesseract) to extract text information from the image.
[1174] Step 3:
[1175] Building a multimodal AI model
[1176] The server uses the preprocessed data to build a multimodal AI model that can comprehensively analyze text, images, and other data formats, and trains the model to determine data similarity.
[1177] Input: Preprocessed data
[1178] Output: A trained multimodal AI model
[1179] Specifically, it feeds data to AI models (such as OpenAI's GPT-4) to train them, and also continuously updates and improves them.
[1180] Step 4:
[1181] Receiving a search query
[1182] The user (store staff) sends a search query by voice input through the smart glasses, for example, "Check inventory, tell me the stock status of product X." The device converts this voice data into text format and sends it to the server.
[1183] Input: Voice query
[1184] Output: Text query
[1185] Specifically, it converts voice input into text using speech recognition technology (e.g., Google Cloud Speech-to-Text API).
[1186] Step 5:
[1187] Finding related data
[1188] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's database, ranking and extracting the most relevant materials.
[1189] Input: Text query, trained multimodal AI model, in-house database
[1190] Output: Search results (list of related materials)
[1191] Specifically, the server inputs a query into the model, searches the database for materials that match the query, and retrieves the results.
[1192] Step 6:
[1193] Submitting and viewing search results
[1194] The server sends the search results to the user's device (smart glasses), and the user can view the displayed search results, such as "Report on marketing strategy for new product X" or "Related meeting notes."
[1195] Input: Search results
[1196] Output: Results displayed on smart glasses
[1197] Specifically, the server formats the search results and transmits them in a format suitable for the smart glasses' display.
[1198] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1199] This invention is a system that collects data from an in-house mail server, messaging platform, and cloud storage, and combines it with an emotion engine that recognizes the user's emotions to perform more effective document searches. Below, the program processing of this system is explained in natural language.
[1200] First, the server periodically connects to the company's internal mail server, messaging platforms (e.g., Slack and Microsoft Teams), and cloud storage (e.g., Google Drive and Dropbox) to collect all data, including email bodies and attachments, chat logs and uploaded files from messaging platforms, and various files from cloud storage.
[1201] Next, the server preprocesses the collected data. During this preprocessing step, the server analyzes the text data using natural language processing (NLP) technology, removes unnecessary information, and converts it into a format that is easier to analyze. For image data, the server also uses optical character recognition (OCR) technology to extract text information from the image. Furthermore, the server analyzes files such as PDFs and Word documents and converts them into text data.
[1202] The server then builds a multimodal AI model based on the preprocessed data (text data, image data, and other data formats). This AI model can comprehensively analyze text, images, and other data formats to extract data characteristics. This model is then trained to determine data similarity.
[1203] When a user searches for specific information on a device, they input a search query. For example, they input a query such as "Marketing strategy for new product Y." The device then sends this search query to the server. At this stage, the device can also use an emotion engine to detect the user's emotions. The emotion engine recognizes the user's facial expressions and voice and determines the user's current emotional state.
[1204] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's internal database. It then performs a search based on the query and ranks and extracts the most relevant materials. It performs an integrated search of materials that include not only text data but also different data formats such as images and PDFs. Furthermore, it adjusts the ranking of search results by taking into account the user's emotional state detected by the emotion engine. For example, if the user is feeling stressed, it will prioritize displaying materials that are more concise and easy to understand.
[1205] The server also searches past related communication logs. For example, if the relevant document was sent by email, it searches the email body and attachments and extracts related Slack messages. This allows users to view the communication logs related to the document along with the related documents.
[1206] The server sends the search results to the terminal and displays them to the user. The user can check the displayed search results and select and view the necessary materials. For example, past reports on the marketing strategy for new product Y, presentation materials, related meeting notes, etc. are displayed and can be reviewed. In addition, an emotion engine is used to provide information tailored to the user's emotions, improving the user experience.
[1207] This system improves work efficiency by making it easy to search and retrieve necessary documents even when documents are not properly handed over after an employee leaves or is transferred. It also supports not only text but also images and other data formats, allowing for quick and efficient searches of various documents. Furthermore, by utilizing an emotion engine, it is possible to provide search results that reflect the user's emotional state, improving user satisfaction.
[1208] The processing flow will be explained below.
[1209] Step 1: Collect data
[1210] The server periodically connects to the company's email server, messaging platform, and cloud storage to collect data. Specifically, it collects email text, attachments, and metadata from the email server, chat logs and uploaded files from the messaging platform, and various document files (PDF, Word, image files, etc.) from the cloud storage.
[1211] Step 2: Preprocessing the data
[1212] The server preprocesses the collected data by format. For text data, it analyzes it using natural language processing (NLP) technology and removes noise. For image data, it uses optical character recognition (OCR) technology to extract text information from the image. PDFs and Word documents are also converted into text data and their contents are analyzed.
[1213] Step 3: Building an AI model
[1214] The server then builds a multimodal AI model based on the preprocessed data. This model is trained to comprehensively analyze text, images, and other data formats and extract data characteristics, resulting in a model capable of similarity judgment.
[1215] Step 4: Receiving a search query
[1216] The device receives a search query from the user. For example, the user may enter a search query such as "Marketing strategy for new product Y." The device sends this query to the server and simultaneously activates an emotion engine to detect the user's emotional state.
[1217] Step 5: Detecting Emotional State
[1218] The device uses an emotion engine to analyze the user's emotional state. Specifically, it recognizes emotions from the user's facial expressions and voice and determines the user's current emotional state. For example, it detects whether the user is feeling stressed or relaxed.
[1219] Step 6: Processing the search query
[1220] Based on the received search query, the server uses a multimodal AI model to identify relevant materials from the company's database. It performs a query-based search and ranks and extracts highly relevant materials. It performs an integrated search of materials, including not only text data but also different data formats such as images and PDFs.
[1221] Step 7: Adjust your results based on your emotional state
[1222] The server adjusts the ranking of search results based on the user's emotional state as detected by the emotion engine. For example, if the user is feeling stressed, it will prioritize more concise and easy-to-understand materials. On the other hand, if the user is relaxed, it will provide more detailed and specialized materials.
[1223] Step 8: Search past communication logs
[1224] The server also searches past communication logs (emails and messages on messaging platforms) for related materials, extracting relevant email text, Slack messages, etc., and searching for information from the communication logs.
[1225] Step 9: Serving search results
[1226] The server sends the search results and information extracted from past communication logs to the terminal. The terminal displays this data on a user interface, allowing the user to check each piece of information. For example, a report on the marketing strategy for new product Y and related communication records may be displayed.
[1227] Step 10: User validation and utilization
[1228] Users can check the search results displayed on their devices, select and view related materials as needed, and download or print any information they deem necessary for further use.
[1229] In this way, this system aims to improve work efficiency and user satisfaction by quickly searching for necessary information from vast amounts of internal company data and providing information that corresponds to the user's emotional state.
[1230] Example 2
[1231] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1232] Conventional document search systems struggled to effectively collect and integrate data from internal email servers, messaging platforms, and cloud storage, and lacked a way to appropriately adjust search results based on the user's emotional state. This resulted in problems such as an inability to respond quickly and accurately to user requests, leading to reduced work efficiency. Furthermore, it was difficult to comprehensively search documents in different data formats, making it difficult to find the information needed when document handover was insufficient.
[1233] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1234] In this invention, the server includes: means for collecting data from mail servers, messaging platforms, and cloud storages; means for analyzing text data, image data, and other data using natural language processing technology and optical character recognition technology in preprocessing and integrating them; means for building a multimodal AI model including text, image, and other data formats and analyzing the similarity of the data; means for receiving a user's search query and identifying related data based on the search query using an emotion engine that detects the user's emotional state; means for searching past communication logs to send related materials; and means for providing the search results to the user terminal.
[1235] This makes it possible to centrally collect and integrate data from various data sources within the company and provide appropriate search results based on user sentiment. Users can quickly find the documents they need, improving work efficiency. In addition, by searching documents in different data formats in an integrated manner, related information can be easily obtained even if document handover is insufficient.
[1236] A "mail server" is a server that manages the sending and receiving of email within an organization or on the Internet.
[1237] A "messaging platform" is software or a service for sending and receiving messages in real time.
[1238] "Cloud storage" is an online service for storing and managing data over the Internet.
[1239] "Natural language processing technology" is a technology that analyzes text data and extracts linguistic features.
[1240] "Optical character recognition technology" is a technology that extracts text information from image data.
[1241] A "multimodal AI model" is an artificial intelligence model that comprehensively analyzes text, images, and other data formats to extract features.
[1242] "User emotional state" refers to the emotional state a user is experiencing when performing a search, and includes techniques for detecting this.
[1243] An "emotion engine" is a device or software that analyzes a user's facial expressions and voice to detect their emotional state.
[1244] A "search query" is text data indicating a search request that a user inputs into a search system.
[1245] "Communication logs" are data containing past communication history across email and messaging platforms.
[1246] The present invention is a system that collects data from an internal mail server, a messaging platform, and cloud storage, and combines it with an emotion engine that recognizes the user's emotional state to perform more effective document search. The program processing of this system is described below.
[1247] First, the server periodically connects to the company's email server, messaging platform, and cloud storage to collect all data. Specifically, Microsoft Exchange and Gmail are used as email servers, Slack and Microsoft Teams are used as messaging platforms, and Google Drive and Dropbox are used as cloud storage. The server uses APIs to collect email bodies and attachments, chat logs, uploaded files, and various files in cloud storage.
[1248] The server then preprocesses the collected data. This preprocessing step uses natural language processing (NLP) and optical character recognition (OCR) techniques. For text data, NLTK and spaCy are used to tokenize and remove stop words. For image data, Tesseract OCR is used to extract text information from images. For PDF and Word documents, Apache Tika is used to convert them to text data.
[1249] The server then builds a multimodal AI model based on the preprocessed data. This model is created using TensorFlow and PyTorch to comprehensively analyze text, images, and other data formats and extract data characteristics. Specifically, it uses a convolutional neural network (CNN) to extract image features, and a recurrent neural network (RNN) or BERT model to extract text features.
[1250] Next, the user inputs a search query from the terminal. For example, the query "Marketing strategy for new product Y" is input. This search query is sent as is from the terminal to the server as a character string.
[1251] The device is equipped with an emotion engine that detects the user's emotional state, for example by capturing and analyzing facial expressions with a camera. This analysis is performed using the Face API from Microsoft Azure's Cognitive Services. The emotion engine detects the user's emotional state and transmits that information along with a query to the server.
[1252] Based on the received search query and emotional information, the server uses a multimodal AI model to identify relevant materials from the company's database. It then assesses data similarity and ranks and extracts highly relevant materials. It also adjusts the ranking of results according to the user's emotional state. If the user is feeling stressed, the algorithm prioritizes displaying more concise and easy-to-understand materials.
[1253] The server also searches past communication logs for related materials. For example, if the material was sent by email, it extracts the email body, attachments, and chat logs from related messaging platforms. Finally, the server sends the search results and related communication logs to the device.
[1254] Users can select and view the necessary materials from the results displayed on their device. For example, they can check past reports, presentation materials, meeting notes, etc. related to "Marketing Strategy for New Product Y."
[1255] As a concrete example, here is a prompt example for a generative AI model when searching for "marketing strategy for new product Y":
[1256] "Search your internal database for materials related to the marketing strategy for new product Y. If the user is feeling stressed, prioritize materials that are more concise and easy to understand."
[1257] This system makes it possible to quickly and efficiently search and retrieve necessary documents even when documents are not fully handed over when an employee resigns or is transferred.In addition, by providing search results that respond to the user's emotions, it is expected that the user experience will improve and business efficiency will increase.
[1258] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1259] Step 1:
[1260] The server periodically connects to the company's mail server, messaging platform, and cloud storage to collect data. In this step, it retrieves email bodies and attachments from the mail server, chat logs and uploaded files from the messaging platform, and various files from the cloud storage. The specific input is a request using the API of each data source, and the output is a collection of collected data.
[1261] Step 2:
[1262] The server preprocesses the collected data. It uses natural language processing (NLP) and optical character recognition (OCR) techniques to tokenize text data, remove stop words, extract text information from image data, and convert PDFs and Word documents into text data. Specific operations include text analysis using NLTK and spaCy, image analysis using Tesseract OCR, and document analysis using Apache Tika. The input is the raw data collected in step 1, and the output is the preprocessed, clean data.
[1263] Step 3:
[1264] The server builds a multimodal AI model based on the preprocessed data. It uses TensorFlow and PyTorch to comprehensively analyze text, images, and other data formats and extract features. Specifically, it uses a Convolutional Neural Network (CNN) to extract image features, and a Recurrent Neural Network (RNN) or BERT model to extract text data features. The input is various preprocessed data, and the output is a trained AI model.
[1265] Step 4:
[1266] A user inputs a query to search for specific information from a terminal. For example, the query might be "Marketing strategy for new product Y." The query is sent directly from the terminal to the server in the form of a string. The input is the query entered by the user, and the output is a search request sent to the server.
[1267] Step 5:
[1268] The device uses an emotion engine to detect the user's emotional state. Specifically, it captures the user's facial expressions with a camera and analyzes them using the Face API of Microsoft Azure's Cognitive Services. The input is an image of the user's facial expression, and the output is the detected emotional state data.
[1269] Step 6:
[1270] The server uses a multimodal AI model to search for relevant materials from the company's database based on the received search query and emotional information. It ranks the search results and adjusts the ranking based on the emotion engine's evaluation. Specifically, if the user is feeling stressed, the server uses an algorithm that prioritizes displaying concise and easy-to-understand materials. The input is the search query and emotional information, and the output is ranked search results.
[1271] Step 7:
[1272] The server also searches past communication logs for related materials. For example, if the material was sent by email, it extracts the email body, attachments, and chat logs from related messaging platforms. The input is the search results, and the output is the related communication logs.
[1273] Step 8:
[1274] The server sends the search results and related communication logs to the terminal. The user can select and view the necessary materials from the results displayed on the terminal. The input is the ranked search results and related communication logs, and the output is the information displayed on the terminal in a user-accessible form.
[1275] Through these steps, users can quickly and efficiently search and extract the information they need.
[1276] (Application example 2)
[1277] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1278] Providing timely and appropriate information is essential to improving work efficiency in factories and reducing worker stress. However, conventional information search systems were unable to provide information tailored to the emotional state of workers, making it difficult to provide efficient work support. Furthermore, extracting information from a wide variety of data formats and past communication logs was time-consuming, making it difficult to provide practical information.
[1279] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1280] In this invention, the server includes: means for collecting data from mail servers, messaging platforms, and cloud storage; means for analyzing and integrating text data, image data, and other data in preprocessing; means for building a multimodal AI model and analyzing data similarity; means for receiving search queries from users and identifying relevant data based on the search queries; means for analyzing the emotional state of workers and adjusting search results based on the results; means for searching past communication logs for sending relevant materials; and means for providing search results to user terminals. This enables timely and appropriate information provision that takes emotional states into account. Furthermore, the system can efficiently extract information from a variety of data formats and past communication logs, thereby providing efficient support for workers.
[1281] A "mail server" is a server for sending, receiving, and storing e-mail.
[1282] A "messaging platform" is a system for exchanging messages between users using text, voice, video, etc.
[1283] "Cloud storage" refers to online storage services for storing and accessing data over the Internet.
[1284] "Means for collecting data" refers to the mechanisms for obtaining the necessary data from mail servers, messaging platforms, and cloud storage.
[1285] "Preprocessing" is the process of converting collected data into a format that is easy to analyze and removing unnecessary information.
[1286] "Text data" is a data format that includes character information.
[1287] "Image data" is a data format that contains visual information.
[1288] A "multimodal AI model" is an artificial intelligence model for integrated analysis of different data formats (text, images, etc.).
[1289] A "user search query" is a search request entered by a user to locate specific information.
[1290] The "emotional state" is a state that indicates the type and intensity of the emotion that the user is currently feeling.
[1291] "Related materials" are materials that contain information that is useful to the user and that are identified based on the search query.
[1292] A "communication log" is a record of messages and files sent and received in the past, as well as related information.
[1293] "Search Results" means a list of relevant materials identified based on a search query.
[1294] A "user terminal" is a device that a user operates directly to check information.
[1295] To realize this application example, the system program is configured as follows.
[1296] The server collects data from email servers, messaging platforms, and cloud storage. Specifically, it uses the Gmail API, Microsoft Teams API, Google Drive API, and other sources to obtain the necessary data. The obtained data is in the form of text, image data, and other data formats, and natural language processing (NLP) and optical character recognition (OCR) technologies are used to preprocess it. After preprocessing, the data is analyzed using Hugging Face's transformers library to build a multimodal AI model. The multimodal AI model comprehensively analyzes text, image data, and other data formats.
[1297] The server also receives search queries from users and identifies relevant data based on the search queries. A user can input a search query such as "How to troubleshoot error code E123 on machine X" using a device (e.g., a smartphone or smart glasses). The server also analyzes the user's emotional state in real time using a device equipped with an emotion engine (e.g., smart glasses with a built-in camera and microphone that analyzes facial expressions and voice). If the user is feeling stressed, the emotion engine will prioritize providing more concise and easy-to-understand materials.
[1298] The search engine searches the database based on the query and also searches past related communication logs. For example, if related documents were sent via email or messaging platforms, the contents of those documents will also be searched. This allows for comprehensive retrieval of past communications and related documents. Search results are ranked according to the user's emotional state, with the most relevant information being displayed first.
[1299] The terminal provides the search results to the user, who can then check and view the necessary materials. This improves work efficiency and reduces worker stress. This system makes it possible to efficiently extract information from a wide range of data formats, significantly improving the factory work environment.
[1300] As a concrete example, consider a factory worker troubleshooting a new machine. If the worker types the query "How to troubleshoot error code E123 on machine X" into the smart glasses and is stressed, the application will prioritize short, easy-to-understand instructions.
[1301] Example prompt sentence:
[1302] "Find past troubleshooting documents for error code E123 on machine X. Workers are stressed right now, so please prioritize concise documents."
[1303] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1304] Step 1:
[1305] The server collects data from email servers, messaging platforms, and cloud storage. As input, it uses the Gmail API, Microsoft Teams API, Google Drive API, etc. to obtain the target email, message, and file data. The output is the collected text data, image data, and other data formats.
[1306] Step 2:
[1307] The server preprocesses the collected data. Specifically, it uses natural language processing (NLP) technology to analyze the text data, remove unnecessary information, and convert it into a format that is easier to analyze. It also uses optical character recognition (OCR) technology to extract text information from image data. It also analyzes files such as PDFs and Word documents and converts them into text data. The input is the data collected in step 1, and the output is the preprocessed data.
[1308] Step 3:
[1309] The server builds a multimodal AI model based on various preprocessed data (text data, image data, and other data formats). Using Hugging Face's transformers library, it trains an AI model that can comprehensively analyze different data formats. The input is the preprocessed data output from Step 2, and the output is an AI model that can analyze data similarities.
[1310] Step 4:
[1311] A user uses a device (e.g., a smartphone, smart glasses) to input a search query, such as "how to troubleshoot error code E123 on machine X." The input is the user's query, and the output is the query data.
[1312] Step 5:
[1313] The device sends the search query to the server. At the same time, the device uses an emotion engine to analyze the user's emotional state. For example, it uses the camera and microphone built into the smart glasses to analyze the user's facial expressions and voice to determine their current emotional state. The input is the query from step 4 and the user's emotional data, and the output is composite data including the query and the emotional state.
[1314] Step 6:
[1315] The server searches the database based on the search query, using a multimodal AI model to identify relevant information and also searches past communication logs to retrieve relevant materials. The input is the composite data from step 5, and the output is a list of relevant materials.
[1316] Step 7:
[1317] The server adjusts the ranking of the retrieved search results taking into account the user's emotional state. If the emotion engine detects that the user is stressed, it prioritizes displaying materials that are more concise and easy to understand. The input is the list of related materials from step 6 and the emotion analysis data, and the output is search results that have been ranked based on emotion.
[1318] Step 8:
[1319] The server sends the final search results to the terminal, where the user can check the search results on the terminal and view the necessary materials. The input is the search results ranked based on the sentiment in step 7, and the output is the search results displayed on the terminal.
[1320] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1321] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1322] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1323] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1324] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1325] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1326] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1327] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1328] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1329] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1330] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1331] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1332] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1333] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1334] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1335] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1336] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1337] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1338] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1339] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1340] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1341] The following is further disclosed regarding the above embodiment.
[1342] (Claim 1)
[1343] a means for collecting data from mail servers, messaging platforms, and cloud storage;
[1344] A means of analyzing and integrating text data, image data, and other data in preprocessing;
[1345] A means to build multimodal AI models and analyze data similarities;
[1346] means for receiving a search query from a user and identifying relevant data based on the search query;
[1347] a means of searching past communication logs to transmit relevant materials;
[1348] means for providing search results to a user terminal;
[1349] A system including:
[1350] (Claim 2)
[1351] 2. The system according to claim 1, wherein the preprocessing means uses natural language processing technology and optical character recognition technology.
[1352] (Claim 3)
[1353] 10. The system of claim 1, wherein the multimodal AI model combines and analyzes text data, image data, and other data formats.
[1354] "Example 1"
[1355] (Claim 1)
[1356] a means for collecting data from mail servers, messaging platforms, and cloud storage;
[1357] A means of analyzing and integrating text data, image data, and other data in preprocessing;
[1358] A means for analyzing text data using natural language processing technology;
[1359] means for extracting text information from the image data using optical character recognition techniques;
[1360] A means to build multimodal AI models and analyze data similarities;
[1361] means for receiving a search query from a user and identifying relevant data based on the search query;
[1362] a means of searching past communication logs to transmit relevant materials;
[1363] means for providing search results to a user terminal;
[1364] A system including:
[1365] (Claim 2)
[1366] 2. The system according to claim 1, wherein the preprocessing means uses natural language processing technology and optical character recognition technology.
[1367] (Claim 3)
[1368] 10. The system of claim 1, wherein the multimodal AI model combines and analyzes text data, image data, and other data formats.
[1369] "Application Example 1"
[1370] (Claim 1)
[1371] a means for collecting data from mail servers, messaging platforms, and cloud storage;
[1372] A means of analyzing and integrating text data, image data, and other data in preprocessing;
[1373] A means to build multimodal AI models and analyze data similarities;
[1374] means for receiving a search query from a user and identifying relevant data based on the search query;
[1375] a means of searching past communication logs to transmit relevant materials;
[1376] means for providing search results to a user terminal;
[1377] A way for in-store staff to submit queries via voice input and display search results through smart glasses;
[1378] A system including:
[1379] (Claim 2)
[1380] 2. The system according to claim 1, wherein the preprocessing means uses natural language processing technology and optical character recognition technology.
[1381] (Claim 3)
[1382] 10. The system of claim 1, wherein the multimodal AI model combines and analyzes text data, image data, and other data formats.
[1383] "Example 2: Combining Emotion Engines"
[1384] (Claim 1)
[1385] a means for collecting data from mail servers, messaging platforms, and cloud storage;
[1386] A means for analyzing text data, image data, and other data using natural language processing and optical character recognition technologies in preprocessing and integrating them;
[1387] A means to build multimodal AI models that include text, images, and other data formats and analyze data similarities;
[1388] means for receiving a user's search query and identifying relevant data based on the search query using an emotion engine that detects the user's emotional state;
[1389] a means of searching past communication logs to transmit relevant materials;
[1390] means for providing search results to a user terminal;
[1391] A system including:
[1392] (Claim 2)
[1393] 2. The system according to claim 1, wherein the preprocessing means uses natural language processing technology and optical character recognition technology.
[1394] (Claim 3)
[1395] 10. The system of claim 1, wherein the multimodal AI model combines and analyzes text data, image data, and other data formats.
[1396] (Claim 4)
[1397] 10. The system of claim 1, further comprising: detecting a user's emotional state and adjusting the ranking of search results based thereon.
[1398] "Application example 2 when combining emotion engines"
[1399] (Claim 1)
[1400] a means for collecting data from mail servers, messaging platforms, and cloud storage;
[1401] A means of analyzing and integrating text data, image data, and other data in preprocessing;
[1402] A means to build multimodal AI models and analyze data similarities;
[1403] means for receiving a search query from a user and identifying relevant data based on the search query;
[1404] a means for analyzing the emotional state of the worker and adjusting search results accordingly;
[1405] a means of searching past communication logs to transmit relevant materials;
[1406] means for providing search results to a user terminal;
[1407] A system including:
[1408] (Claim 2)
[1409] 2. The system according to claim 1, wherein the preprocessing means uses natural language processing technology and optical character recognition technology.
[1410] (Claim 3)
[1411] 10. The system of claim 1, wherein the multimodal AI model combines and analyzes text data, image data, and other data formats. [Explanation of symbols]
[1412] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for collecting data from mail servers, messaging platforms, and cloud storage; A means of analyzing and integrating text data, image data, and other data in preprocessing; A means to build multimodal AI models and analyze data similarities; means for receiving a search query from a user and identifying relevant data based on the search query; a means of searching past communication logs to transmit relevant materials; means for providing search results to a user terminal; A system including:
2. 2. The system according to claim 1, wherein the preprocessing means uses natural language processing technology and optical character recognition technology.
3. 10. The system of claim 1, wherein the multimodal AI model combines and analyzes text data, image data, and other data formats.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A