System

A system addresses inconsistent business terminology by monitoring documents, extracting terms, standardizing spellings, and providing real-time annotations, enabling efficient adaptation and understanding for new employees.

JP2026019735APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024121483
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Newly assigned employees or employees transferred to a new department face challenges in understanding department-specific business terminology due to inconsistent spelling and lack of maintenance, leading to reduced work efficiency.

Method used

A system that monitors documents and files for new uploads, extracts business terms using NLP, stores them in a database for standardization, generates and updates Wiki pages, detects spelling variations, and provides real-time annotations during meetings to ensure consistent terminology understanding.

Benefits of technology

The system facilitates quick adaptation to new work environments by standardizing business terminology, reducing inconsistencies, and enhancing work efficiency by providing immediate explanations and corrections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019735000001_ABST
    Figure 2026019735000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: This system is provided with a means for monitoring materials and documents and detecting newly uploaded materials and documents, a means for extracting business words from the materials and documents by using a natural language processing model, a means for storing the extracted business words in a database, confirming coincidence with existing words and unifying notation, and a means for automatically generating and updating a Wiki page based on the business words. A system comprising: means for adding a description of a term; means for periodically monitoring a file sharing system, detecting spelling variations from new documents, and sending a notification to prompt correction; and means for receiving audio input during a meeting, detecting business terms from the audio data, and generating and displaying annotations in real-time.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In order for newly assigned employees or employees who have transferred to a new department to efficiently adapt to their work, they need to understand the terminology and abbreviations specific to their department. However, the current business terminology summary suffers from a lack of maintenance and inconsistent spelling, such as the existence of multiple terms that refer to the same phenomenon, making it difficult for new employees to understand. This can result in a lot of time and effort being required before employees can begin work, and can lead to reduced work efficiency. A system that can solve these problems and standardize business terminology and speed up understanding is needed. [Means for solving the problem]

[0005] The present invention solves this problem by providing a system that includes: a means for monitoring materials and documents and detecting newly uploaded ones; a means for extracting business terms from materials and documents using a natural language processing model; a means for storing the extracted business terms in a database, checking for matches with existing terms, and standardizing spelling; a means for automatically generating and updating Wiki pages based on the business terms and adding explanations of the terms; a means for periodically monitoring a file sharing system, detecting spelling variations in new documents, and sending notifications prompting corrections; and a means for receiving voice input during meetings, detecting business terms from the voice data, and generating and displaying annotations in real time. This system improves work efficiency by making it easier for new employees to quickly understand business terms and reducing spelling variations.

[0006] "Means for monitoring materials and documents and detecting newly uploaded materials and documents" refers to the processes and functions for automatically detecting newly added files in document management systems and file sharing systems.

[0007] "Means for extracting business terms from materials and documents using natural language processing models" refers to algorithms and methods that use natural language processing (NLP) technology to automatically identify and list specific business terms within text.

[0008] "Means of storing extracted business terms in a database, checking for matches with existing terms, and standardizing notation" refers to a function that stores extracted business terms in a database, compares them with existing data to check for matching terms, and adjusts different notations to make them uniform.

[0009] "Means for automatically generating and updating Wiki pages based on business terms and adding explanations of terms" refers to the process or function for automatically creating and updating Wiki-style pages based on business terms registered in a database and adding detailed explanations of each term.

[0010] "Means for periodically monitoring the file sharing system, detecting spelling variations in new documents, and sending notifications urging users to make corrections" refers to a function that monitors the file sharing system at regular intervals, detects different spellings with the same meaning that exist in newly uploaded documents, and notifies users to correct them to the appropriate unified spelling.

[0011] "Means for receiving voice input during a meeting, detecting business terms from the voice data, and generating and displaying annotations in real time" refers to the process or technology of collecting voice input during a meeting, converting it into text data, identifying business terms, and generating explanations and annotations for those terms in real time and displaying them on the screen. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0013] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0014] First, the terms used in the following description will be explained.

[0015] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0016] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0017] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0018] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0020] [First embodiment]

[0021] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0022] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0025] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0028] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0032] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0033] MODE FOR CARRYING OUT THE INVENTION

[0034] The present invention is a system that uses AI to automatically create a Wiki that summarizes business terminology, eliminating problems such as lack of maintenance and inconsistent spelling. An embodiment of this system is described in detail below.

[0035] System Overview

[0036] The system has the following main functions:

[0037] 1. Monitoring and detection of materials and documents

[0038] 2. Extracting business terms using NLP models

[0039] 3. Storing terms in the database and standardizing notation

[0040] 4. Automatic generation and updating of Wiki pages

[0041] 5. Spelling Variation Detection and Notification

[0042] 6. Real-time annotation generation during meetings

[0043] Monitoring and detection of materials and documents

[0044] The server monitors the file sharing system and automatically detects newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[0045] Extracting business terms using NLP models

[0046] The server uses natural language processing (NLP) models to analyze materials and documents and extract business terms, abbreviations, and technical terms, which then lists specific business terms and allows for efficient analysis.

[0047] Terminology storage in database and standardization of notation

[0048] The server stores the extracted business terms in a database, checks whether they match existing terms, and corrects any inconsistencies in spelling to create a consistent term list.

[0049] Automatic generation and updating of Wiki pages

[0050] The server automatically generates and updates Wiki pages based on business terms stored in the database, adding detailed explanations and related links for each term to make them easily accessible to employees.

[0051] Spelling variation detection and notification

[0052] The server periodically monitors the file sharing system and analyzes newly uploaded documents. If it detects a spelling variation that does not match the existing terminology list, it notifies the user and encourages them to correct the spelling.

[0053] Real-time annotation during meetings

[0054] The device receives voice input during the meeting and converts the voice data into text. The server detects business terms from the converted data and generates annotations in real time, which are sent to the device. The device then displays the annotations on the screen of the meeting participants to support understanding.

[0055] Specific examples

[0056] Consider the process flow when a newly assigned employee downloads "Project Meeting Materials." This material contains specific technical terms and abbreviations.

[0057] 1. Material detection and analysis:

[0058] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[0059] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[0060] 2. Terminology database storage and standardization:

[0061] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[0062] 3. Generate Wiki pages:

[0063] The server automatically generates Wiki pages based on the extracted terms, adding detailed descriptions and related links.

[0064] 4. Notification of spelling variations:

[0065] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[0066] 5. Real-time annotation during meetings:

[0067] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this as "ROI (Return on Investment)" and displays it in real time.

[0068] In this way, the system helps new employees quickly adapt to their work and contributes to standardizing and promoting understanding of work terminology.

[0069] The processing flow will be explained below.

[0070] Program processing

[0071] 1. Monitoring and detection of materials and documents

[0072] Step 1:

[0073] The server accesses the file sharing system and monitors newly uploaded materials and documents.

[0074] Step 2:

[0075] The server scans the directory of the file sharing system at regular time intervals to check whether new files exist.

[0076] Step 3:

[0077] The server detects the newly uploaded file and retrieves the file's metadata (such as name, upload date and time, file type, etc.).

[0078] 2. Extracting business terms using NLP models

[0079] Step 4:

[0080] The server inputs the detected new files into a natural language processing (NLP) model.

[0081] Step 5:

[0082] The server uses NLP models to analyze the content of the files and extract business terms, abbreviations, and technical terms.

[0083] Step 6:

[0084] The server organizes the extracted business terms in list format and temporarily stores the analysis results.

[0085] 3. Storing terms in the database and standardizing notation

[0086] Step 7:

[0087] The server stores the extracted business terms in a database.

[0088] Step 8:

[0089] The server checks the newly extracted terms against existing terms in the database to see if there is a match.

[0090] Step 9:

[0091] If the term is spelled differently, the server will correct it to a unified spelling for consistency.

[0092] 4. Automatic generation and updating of Wiki pages

[0093] Step 10:

[0094] The server runs a script that automatically generates Wiki pages based on business terms in the database.

[0095] Step 11:

[0096] The server adds detailed descriptions and related links to each term to the Wiki page.

[0097] Step 12:

[0098] The server uploads the generated Wiki page to the company's internal Wiki system so that users can view it.

[0099] 5. Spelling Variation Detection and Notification

[0100] Step 13:

[0101] The server periodically monitors the file sharing system and analyzes newly uploaded documents.

[0102] Step 14:

[0103] The server detects spelling variations when a term in the parsed document does not match an existing list of terms.

[0104] Step 15:

[0105] When the server detects a spelling variation, it sends a notification email to the relevant user, prompting them to correct the spelling to the recommended spelling.

[0106] 6. Real-time annotation generation during meetings

[0107] Step 16:

[0108] The terminal receives voice input during the conference and acquires the voice data.

[0109] Step 17:

[0110] The server passes the received voice data to a voice recognition service and converts it into text.

[0111] Step 18:

[0112] The server detects business terms from the text data and generates annotations in real time.

[0113] Step 19:

[0114] The terminal displays the generated annotations on the screens of the conference participants to support their understanding of the terminology.

[0115] Example 1

[0116] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0117] Inconsistent notation and lack of maintenance in materials and documents are often problems in business activities. New employees and those transferred from other departments often have difficulty understanding business terms and abbreviations, which can lead to reduced work efficiency. Also, the use of technical terms during meetings often reduces participants' understanding. It is necessary to solve these problems and provide information quickly and clearly while maintaining consistency in business terminology.

[0118] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0119] In this invention, the server includes a means for monitoring documents and automatically detecting newly uploaded documents, a means for extracting business terms from documents using a natural language processing model, and a means for storing the extracted business terms in a database, checking for matches with existing terms, and standardizing their spellings to a consistent format. This allows for the consolidation and standardization of business terms. The server also includes a means for automatically generating and updating Wiki pages based on business terms and adding term explanations, a means for periodically monitoring the file system, detecting spelling variations in new documents, and sending notifications prompting corrections. The server also includes a means for receiving voice input during meetings, detecting business terms from the voice data, and generating and displaying annotations in real time. This allows for the rapid sharing and understanding of new information, improving work efficiency and accuracy.

[0120] "Materials" refers to information media such as documents, files, and reports used in business activities.

[0121] A "natural language processing model" refers to a model that makes full use of artificial intelligence technology to analyze and understand human language.

[0122] "Business jargon" refers to specialized words and abbreviations used in specific industries or jobs.

[0123] A "database" refers to a system for organizing and storing extracted information and data so that it can be used efficiently.

[0124] "Orthographic variation" refers to differences between words or terms that have the same meaning but are expressed in different forms.

[0125] A "Wiki page" refers to a web page that can be collaboratively edited on the Internet or an intranet.

[0126] A "file system" refers to the method and implementation for managing and storing data and files.

[0127] "Annotation" refers to explanations or explanatory text added to aid understanding.

[0128] "Voice input" refers to a method of inputting voice information into a computer via a microphone or the like.

[0129] MODE FOR CARRYING OUT THE INVENTION

[0130] The present invention is a system that uses AI to automatically create summaries of business terms and resolves issues such as lack of maintenance and inconsistent spelling. Specific embodiments of this system are described in detail below.

[0131] System configuration

[0132] This system is mainly composed of three elements: a server, a terminal, and a user. The server processes and monitors data, the terminal functions as an input and display device, and the user provides and uses information.

[0133] Key Features

[0134] Material monitoring and detection

[0135] The server monitors the file sharing system and automatically detects newly uploaded files. Specifically, the server uses a Python script to check the file system for changes at regular intervals and obtain a list of new or updated files. This operation utilizes the OS's file monitoring function and the cloud service's API.

[0136] Extracting business terms using NLP models

[0137] The server analyzes the documents using a natural language processing (NLP) model to extract business terms. This operation uses a pre-trained NLP model (e.g., BERT or spaCy). Depending on the file format, the server converts the file (e.g., PDF or Word) into text format, and then uses the NLP model to analyze and extract key terms.

[0138] Terminology storage in database and standardization of notation

[0139] The server stores the extracted business terms in a dedicated database. Here, it checks whether they match existing terms and standardizes them into consistent notations. MySQL, PostgreSQL, or other database systems are used. The server executes queries to check and update the contents of the database.

[0140] Automatic generation and updating of Wiki pages

[0141] The server automatically generates and updates Wiki pages based on business terms stored in the database. Wiki pages are generated using Markdown format or HTML templates. Specifically, the server retrieves terminology information from the database and applies it to templates to generate Wiki pages.

[0142] Spelling variation detection and notification

[0143] The server periodically monitors the file sharing system and compares new material with the existing database. If a spelling variation is detected, the server notifies the user via email or instant messaging services (e.g., Slack).

[0144] Real-time annotation during meetings

[0145] The device receives voice input during the meeting and converts the voice data into text. The Google Speech-to-Text API is used for speech recognition. For example, if someone says "ROI" during a meeting, the server analyzes it in real time and generates the annotation "ROI (Return on Investment)." The generated annotation is then displayed on the device of each meeting participant.

[0146] Specific examples

[0147] The following shows the process flow when a newly assigned employee downloads "Project Meeting Materials."

[0148] 1. Material detection and analysis:

[0149] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[0150] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[0151] 2. Terminology database storage and standardization:

[0152] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[0153] 3. Generate Wiki pages:

[0154] The server automatically generates Wiki pages based on the extracted terms, adding detailed descriptions and related links.

[0155] 4. Notification of spelling variations:

[0156] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[0157] 5. Real-time annotation during meetings:

[0158] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this as "ROI (Return on Investment)" and displays it in real time.

[0159] Prompt Sentence Examples

[0160] 1. Material detection and analysis:

[0161] "Detect newly uploaded conference materials and extract the technical terms they contain."

[0162] 2. Terminology database storage and standardization:

[0163] "Store the extracted terms in a database and check if they match existing terms."

[0164] 3. Generate Wiki pages:

[0165] "Generate wiki pages based on terms stored in your database and add detailed descriptions and related links."

[0166] 4. Notification of spelling variations:

[0167] "If you detect spelling variations in newly uploaded materials, please notify us of the recommended spelling."

[0168] 5. Real-time annotation during meetings:

[0169] "Extract technical terms from audio during meetings in real time and generate and display annotations."

[0170] The above describes a specific embodiment of the present invention. This system makes it possible to standardize business terminology and promote understanding, thereby realizing efficient business operations.

[0171] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0172] Step 1: Monitoring and detecting material

[0173] The server monitors the file sharing system. Specifically, the server scans the contents of the file system at regular intervals to detect newly added or updated files. The input is the file system path and the monitoring interval. The output is a list of new or updated files. The server also records file metadata (file name, creation date and time, and update date and time).

[0174] Step 2: Prepare the file for analysis

[0175] The server obtains the paths of the detected files and prepares them for analysis. First, it checks the file format (PDF, Word, etc.) and converts it to text data using an appropriate text conversion tool (e.g., PyPDF2 library for PDF, python-docx library for Word). The input is a list of detected files, and the output is text data.

[0176] Step 3: Extracting business terms using an NLP model

[0177] The server inputs the converted text data into an NLP model to extract business terms, abbreviations, and technical terms. This process uses a pre-trained NLP model (e.g., BERT or spaCy). The input is the text data, and the output is a list of extracted business terms. Specifically, terms such as "ROI" and "critical path" are extracted.

[0178] Step 4: Store business terms in a database

[0179] The server stores the extracted business terms in a database. During this process, it checks for matches with existing terms and standardizes the notation. The database used is MySQL or PostgreSQL. The input is the extracted business term list, and the output is the updated database. For example, if "ROI" and "return on investment" match, they are unified to "ROI."

[0180] Step 5: Automatically generate and update Wiki pages

[0181] The server automatically generates and updates Wiki pages based on business terms stored in the database. Markdown format and HTML templates are used for this process. The input is the business term information in the database, and the output is the generated or updated Wiki page. For example, the "ROI" page contains the description "Return on Investment."

[0182] Step 6: Detecting and reporting spelling variations

[0183] The server compares newly uploaded materials with the existing database to detect variations in notation. If any discrepancies are found during this process, a notification is sent to the user. Notification methods include email and instant messaging services (e.g., Slack). The input is the new material and database information, and the output is a notification to the user. For example, a notification recommending that "return on investment" be changed to "ROI" is sent.

[0184] Step 7: Real-time annotation during the meeting

[0185] The device receives voice input during the meeting and converts the voice data into text. This process uses the Google Speech-to-Text API. The input is the voice data from the meeting, and the output is text data. The converted text data is sent to the server, and if a business term is detected, an annotation is generated in real time and displayed on the device. For example, if "ROI" is spoken during a meeting, the annotation "ROI (Return on Investment)" is generated and displayed.

[0186] (Application example 1)

[0187] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0188] In today's business environment, multiple documents and materials are frequently uploaded, often containing many technical terms and abbreviations. This makes it difficult for newly assigned employees and external participants to quickly adapt to the work. Furthermore, inconsistencies in the spelling of business terms occur, reducing the consistency and efficiency of information. There is a need to solve these problems and improve business efficiency.

[0189] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0190] In this invention, the server includes: means for monitoring materials and documents and detecting newly uploaded materials and documents; means for extracting business terms from materials and documents using a natural language processing model; means for storing the extracted business terms in a database, checking for matches with existing terms, and standardizing spellings; means for automatically generating and updating information pages based on the business terms and adding term explanations; means for periodically monitoring the file sharing system, detecting spelling variations in new documents, and sending notifications prompting corrections; means for receiving voice input during meetings, detecting business terms from the voice data, and generating and displaying annotations in real time; means for analyzing the camera feed of a visual device and extracting text from documents; means for displaying business terms as annotations from the extracted text; and means for analyzing the voice input of the visual device, performing speech recognition, and providing detailed explanations of the relevant business terms by voice. This enables newly assigned employees to quickly adapt to their work, promotes standardization of business terminology, and promotes understanding.

[0191] "Materials" refers to documents and other information related to business.

[0192] "Document" refers to a medium containing information stored as printed material or digital data.

[0193] "Monitoring" means that a system periodically or continuously checks a particular file sharing system or database to detect newly added information.

[0194] A "natural language processing model" refers to a computer program or algorithm that analyzes human language and extracts useful information from text data.

[0195] "Business jargon" refers to technical terms and abbreviations frequently used in a particular business or industry.

[0196] A "database" refers to a system for efficiently storing, retrieving, and managing information.

[0197] "Unification of spelling" refers to the standardization of words with the same meaning that are expressed in different spellings or writing styles into a consistent format.

[0198] An "information page" is a web page or document that provides detailed explanations and related information about a particular term or concept.

[0199] A "file sharing system" refers to a system that allows multiple users within a specific network to share and access files.

[0200] "Spelling variation" refers to the situation where the same term or concept is written in different ways.

[0201] "Notification" refers to a system-generated message that informs the user of a particular event or situation.

[0202] "Voice input during a meeting" refers to the process by which voice data spoken during a meeting is received by the system.

[0203] "Audio Data" refers to a digital recording of sound collected by a microphone or the like.

[0204] "Generating and displaying annotations in real time" refers to the process of instantly analyzing voice or text input and displaying additional information based on the results.

[0205] "Visual devices" refer to devices worn by users that visually present information, such as smart glasses and head-mounted displays.

[0206] "Camera feed" refers to video data captured by a camera in real time.

[0207] "Extracting text from a document" refers to the process of analyzing character information from camera images or scanned data and extracting that text data.

[0208] "Speech recognition" refers to the technology and algorithms that analyze audio signals and identify meaningful words and sentences from them.

[0209] This invention is a system that uses AI to consolidate business terminology, particularly to improve efficiency in factory work environments. This system is composed of a server, terminals, and users. The specific configuration and functions of this system are described below.

[0210] System configuration and functions

[0211] 1. Monitoring and detection of materials and documents

[0212] The server monitors the factory's file sharing system and database, automatically detecting newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[0213] 2. Extracting business terms using natural language processing models

[0214] The server uses natural language processing (NLP) models to analyze documents and extract business terms, abbreviations, and technical terms, using NLP libraries such as SpaCy.

[0215] 3. Storing terms in the database and standardizing notation

[0216] The server stores the extracted business terms in a database and checks whether they match existing terms. If there are any variations in spelling, they are unified to create a consistent term list. For this purpose, an SQL-based database management system (e.g., MySQL) is used.

[0217] 4. Automatic generation and updating of information pages

[0218] The server automatically generates and updates information pages based on business terms stored in the database, including detailed explanations and related links for each term.

[0219] 5. Spelling Variation Detection and Notification

[0220] The server periodically monitors the file sharing system and analyzes newly uploaded documents. If it detects a spelling variation that does not match the existing term list, it sends a notification to the relevant user urging them to correct the variation. The notification system includes email notifications and push notifications.

[0221] 6. Real-time voice annotation

[0222] The device receives voice input during the meeting and converts the voice data into text. The server detects business terms from the text and generates annotations in real time and sends them to the device, allowing meeting participants to instantly understand the meaning of the terms.

[0223] 7. Visual Device Camera Feed Analysis

[0224] A visual device (e.g., smart glasses) extracts text from a document through a camera feed. After extracting the text using Tesseract OCR, the server analyzes the text and displays real-time annotations for business terms.

[0225] 8. Analysis and presentation of voice input

[0226] The voice input from the visual device is analyzed and speech recognition is performed using the Google Cloud Speech-to-Text API. For recognized terms, the server retrieves detailed descriptions and presents them as audio through the visual device's speaker.

[0227] Specific examples

[0228] For example, suppose a newly assigned employee views a document containing the term "ROI (Return on Investment)." The camera in the vision device captures this, and the server performs text analysis, displaying the annotation "Return on Investment" for the term "ROI." Furthermore, if the employee voice-inputs, "Tell me more about Return on Investment," the vision device recognizes the voice and provides a detailed explanation.

[0229] This will generate a prompt like this:

[0230] "A new document has been uploaded. This document contains the term 'return on investment'. Could you please explain this term in more detail?"

[0231] In this way, the present invention promotes understanding of business terms in a factory work environment, enabling efficient work execution.

[0232] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0233] Step 1:

[0234] The server monitors the file sharing system within the factory and detects newly uploaded materials and documents. As input, it receives directory information and file lists from the file sharing system and identifies new files based on timestamps. After processing this data, it outputs a list of newly uploaded files.

[0235] Step 2:

[0236] The server uses a natural language processing model (e.g., SpaCy) to extract business terms from newly detected materials and documents. As input, it receives the contents of the files detected in step 1 in text format and analyzes the text data. This data processing results in the output of a list of extracted business terms.

[0237] Step 3:

[0238] The server stores the extracted business terms in a database, checks for matches with existing terms, and standardizes notation. As input, it receives the list of business terms extracted in step 2 and compares them with existing terms in the database. This data processing results in a consistent list of terms being output.

[0239] Step 4:

[0240] The server automatically generates and updates information pages based on business terms stored in the database. It receives an updated term list as input and adds detailed explanations and related links for each term. This data calculation results in the output of the latest information page.

[0241] Step 5:

[0242] The server periodically monitors the file sharing system, detects spelling variations that do not match the existing terminology list, and sends a notification prompting correction. It receives the content of newly uploaded documents as input and compares it with the existing terminology list. After processing this data, it outputs a notification containing mismatched terms and suggested corrections.

[0243] Step 6:

[0244] The device converts voice input received during the meeting into text and sends it to the server. It receives the voice data of the meeting participants as input and performs speech recognition using the Google Cloud Speech-to-Text API. This data is then processed and the text data is output.

[0245] Step 7:

[0246] The server generates annotations in real time for business terms detected from the voice input and sends them to the terminal. As input, it receives the text data output in step 6, analyzes and extracts business terms, and generates annotations for them. This data processing results in the output of annotated text.

[0247] Step 8:

[0248] The vision device analyzes the camera feed and extracts text from the document. It receives the video data from the camera feed as input and performs character recognition using Tesseract OCR. It then processes this data and outputs the extracted text data.

[0249] Step 9:

[0250] The server generates annotations of business terms from the extracted text and displays them on a visual device. It receives the text data output in step 8 as input, analyzes and extracts business terms, and generates annotations for them. This data processing results in the output of annotated text.

[0251] Step 10:

[0252] The visual device analyzes voice input and performs speech recognition. It receives the user's voice data as input and performs speech recognition using the Google Cloud Speech-to-Text API. It then processes this data and outputs the recognized text data.

[0253] Step 11:

[0254] The server obtains detailed explanations for the recognized terms and presents them as audio through the speaker of the visual device. As input, it receives the text data output in step 10, generates detailed explanations, and sends them as audio data to the visual device. This data processing results in the audio data of the detailed explanation being output.

[0255] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0256] MODE FOR CARRYING OUT THE INVENTION

[0257] The present invention is a system that uses AI to automatically create a wiki that summarizes business terminology, resolves issues such as lack of maintenance and inconsistent spelling, and combines it with an emotion engine that recognizes user emotions. An embodiment of this system is described in detail below.

[0258] System Overview

[0259] The system has the following main functions:

[0260] 1. Monitoring and detection of materials and documents

[0261] 2. Extracting business terms using NLP models

[0262] 3. Storing terms in the database and standardizing notation

[0263] 4. Automatic generation and updating of Wiki pages

[0264] 5. Spelling Variation Detection and Notification

[0265] 6. Real-time annotation generation during meetings

[0266] 7. Emotion engine that recognizes user emotions

[0267] Monitoring and detection of materials and documents

[0268] The server monitors the file sharing system and automatically detects newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[0269] Extracting business terms using NLP models

[0270] The server uses natural language processing (NLP) models to analyze materials and documents and extract business terms, abbreviations, and technical terms, which then lists specific business terms and allows for efficient analysis.

[0271] Terminology storage in database and standardization of notation

[0272] The server stores the extracted business terms in a database, checks whether they match existing terms, and corrects any inconsistencies in spelling to create a consistent term list.

[0273] Automatic generation and updating of Wiki pages

[0274] The server automatically generates and updates Wiki pages based on business terms stored in the database, adding detailed explanations and related links for each term to make them easily accessible to employees.

[0275] Spelling variation detection and notification

[0276] The server periodically monitors the file sharing system and analyzes newly uploaded documents. If it detects a spelling variation that does not match the existing terminology list, it notifies the user and encourages them to correct the spelling.

[0277] Real-time annotation during meetings

[0278] The device receives voice input during the meeting and converts the voice data into text. The server detects business terms from the converted data and generates annotations in real time, which are sent to the device. The device then displays the annotations on the screen of the meeting participants to support understanding.

[0279] Emotion engine that recognizes user emotions

[0280] The server analyzes the user's voice and text data, recognizes the user's emotions using an emotion engine, and provides appropriate feedback and advice based on the recognized emotions.

[0281] Specific examples

[0282] Consider the process flow when a newly assigned employee downloads "Project Meeting Materials." This material contains specific technical terms and abbreviations.

[0283] 1. Material detection and analysis:

[0284] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[0285] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[0286] 2. Terminology database storage and standardization:

[0287] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[0288] 3. Generate Wiki pages:

[0289] The server automatically generates Wiki pages based on the extracted terms, adding detailed descriptions and related links.

[0290] 4. Notification of spelling variations:

[0291] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[0292] 5. Real-time annotation during meetings:

[0293] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this as "ROI (Return on Investment)" and displays it in real time.

[0294] 6. Leveraging the Emotion Engine:

[0295] The server analyzes the voice data and text data entered by the user during the meeting and uses an emotion engine to recognize the user's emotions. For example, if the user is feeling stressed during the meeting, the server will provide appropriate feedback to support the progress of the meeting.

[0296] The processing flow will be explained below.

[0297] MODE FOR CARRYING OUT THE INVENTION

[0298] Monitoring and detection of materials and documents

[0299] Step 1:

[0300] The server accesses the file sharing system and monitors newly uploaded materials and documents.

[0301] Step 2:

[0302] The server scans the directory of the file sharing system at regular time intervals to check whether new files exist.

[0303] Step 3:

[0304] The server detects the newly uploaded file and retrieves the file's metadata (such as name, upload date and time, file type, etc.).

[0305] Extracting business terms using NLP models

[0306] Step 4:

[0307] The server inputs the detected new files into a natural language processing (NLP) model.

[0308] Step 5:

[0309] The server uses NLP models to analyze the content of the files and extract business terms, abbreviations, and technical terms.

[0310] Step 6:

[0311] The server organizes the extracted business terms in list format and temporarily stores the analysis results.

[0312] Terminology storage in database and standardization of notation

[0313] Step 7:

[0314] The server stores the extracted business terms in a database.

[0315] Step 8:

[0316] The server checks the newly extracted terms against existing terms in the database to see if there is a match.

[0317] Step 9:

[0318] If the term is spelled differently, the server will correct it to a unified spelling for consistency.

[0319] Automatic generation and updating of Wiki pages

[0320] Step 10:

[0321] The server runs a script that automatically generates Wiki pages based on business terms in the database.

[0322] Step 11:

[0323] The server adds detailed descriptions and related links to each term to the Wiki page.

[0324] Step 12:

[0325] The server uploads the generated Wiki page to the company's internal Wiki system so that users can view it.

[0326] Spelling variation detection and notification

[0327] Step 13:

[0328] The server periodically monitors the file sharing system and analyzes newly uploaded documents.

[0329] Step 14:

[0330] The server detects spelling variations when a term in the parsed document does not match an existing list of terms.

[0331] Step 15:

[0332] When the server detects a spelling variation, it sends a notification email to the relevant user, prompting them to correct the spelling to the recommended spelling.

[0333] Real-time annotation during meetings

[0334] Step 16:

[0335] The terminal receives voice input during the conference and acquires the voice data.

[0336] Step 17:

[0337] The server passes the received voice data to a voice recognition service and converts it into text.

[0338] Step 18:

[0339] The server detects business terms from the text data and generates annotations in real time.

[0340] Step 19:

[0341] The terminal displays the generated annotations on the screens of the conference participants to support their understanding of the terminology.

[0342] Emotion engine that recognizes user emotions

[0343] Step 20:

[0344] The terminal collects voices during the conference and text data entered by the user.

[0345] Step 21:

[0346] The server passes the collected voice and text data to an emotion engine to analyze the user's emotions.

[0347] Step 22:

[0348] The server generates appropriate feedback and advice based on the recognized emotions.

[0349] Step 23:

[0350] The terminal displays the generated feedback and advice on the screens of the conference participants to support the progress of the conference.

[0351] Example 2

[0352] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0353] When managing materials and documents, it is necessary to efficiently monitor and detect newly uploaded information. There is also the issue of eliminating inconsistencies in notation due to the existence of unstandardized business terms and abbreviations, and creating and maintaining a consistent terminology list. Furthermore, there is a need to streamline the provision of information during meetings and to implement a feedback system based on user sentiment.

[0354] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0355] In this invention, the server includes: means for monitoring materials and documents and detecting newly uploaded materials and documents; means for extracting business terms from materials and documents using a natural language processing model; means for storing the extracted business terms in a database, checking for matches with existing terms, and standardizing notation; means for automatically generating and updating information pages based on the business terms and adding explanations of the terms; means for periodically monitoring the information sharing system, detecting notation variations in new documents, and sending notifications prompting corrections; means for receiving voice input during a meeting, detecting business terms from the voice data, and generating and displaying annotations in real time; and means for analyzing user voice and text data, recognizing user emotions, and providing feedback and advice. This enables efficient document management, standardization of notation, maintaining consistency, efficient information provision during meetings, and appropriate feedback based on user emotions.

[0356] "Materials and documents" refers to electronic files containing text data and graphical data, such as reports, plans, technical manuals, procedure manuals, e-mails, memos, and meeting materials used inside and outside a company.

[0357] "Monitoring" means constantly checking operations such as adding, changing, or deleting files that are performed on a specific system or network, and detecting them when certain conditions are met.

[0358] A "natural language processing model" refers to an algorithm or machine learning model designed to understand, analyze, and generate human language, specifically extracting business terms and analyzing text.

[0359] "Business jargon" refers to the technical terms, abbreviations, and phrases used within a particular company or industry that are an important part of business processes and communication.

[0360] "Database" refers to a data storage system that stores data in an organized manner and is designed to make it easy to access and manage.

[0361] "Unification of notation" refers to the process of changing identical terms with different notations or expressions into a consistent format to maintain data integrity and consistency.

[0362] An "information page" is a web page or digital document that aggregates business terms and related information and displays them in an easy-to-reference format.

[0363] An "information sharing system" is a platform that allows multiple users to upload, download, and share files and information, and includes cloud storage and corporate network drives.

[0364] "Spelling variation" refers to the use of different expressions or spellings that have the same meaning, which creates problems that make understanding and management complicated.

[0365] "Sending a notification" refers to the act of communicating alerts or information to a user based on a specific event or condition detected by the system.

[0366] "Receiving audio input" refers to the process of collecting audio data through a microphone or other audio capture device and passing that data to the system.

[0367] "Audio data" is data that is a digital representation of human speech recorded via a voice input device.

[0368] "Generating annotations in real time" refers to the process of instantly adding relevant explanations and information to input speech or text data and displaying them.

[0369] "Recognizing user emotions" refers to the technical process of determining the emotions a user is feeling at that time through analysis of voice and text data.

[0370] "Providing feedback and advice" refers to the process by which the system suggests appropriate responses or recommended actions based on the user's perceived emotions.

[0371] MODE FOR CARRYING OUT THE INVENTION

[0372] The present invention is a system that uses AI to automatically create information pages summarizing business terms, eliminating problems such as lack of maintenance and inconsistent spelling, and combining this with an emotion engine that recognizes user emotions. An embodiment of this system is described in detail below.

[0373] Hardware and software used

[0374] The system utilizes the following hardware and software:

[0375] Hardware: Server, terminal, audio input device (microphone)

[0376] File sharing systems: Dropbox, Google Drive, etc.

[0377] Natural language processing models: Google BERT, OpenAI GPT-3, etc.

[0378] Database: MySQL, PostgreSQL

[0379] Emotion engine: IBM Watson, Azure Cognitive Services

[0380] Speech recognition software: Google Speech-to-Text, Microsoft Azure Speech Service

[0381] Overall flow

[0382] The server monitors the file sharing system to detect and analyze newly uploaded materials and documents. A natural language processing model is used for the analysis, which extracts business terms, abbreviations, and technical terms from the materials. Before storing the extracted terms in the database, they are compared with existing terms, and any variations in spelling are unified.

[0383] Next, the server automatically generates and updates information pages based on the business terms stored in the database. The generated information pages include detailed explanations of the terms and related links. The server also periodically monitors the information sharing system to detect spelling variations in new documents. If any are detected, a notification is sent to the relevant user.

[0384] During a meeting, the device receives voice input and converts the voice data into text. The text data is then sent to a server, which detects business terms and generates annotations in real time. The annotations are then sent to the device and displayed on the screens of meeting participants.

[0385] Finally, the server analyzes the user's voice and text data, and uses an emotion engine to recognize the user's emotions. Based on the recognized emotions, the server provides appropriate feedback and advice.

[0386] Specific examples

[0387] The process flow when a newly assigned employee downloads "Project Meeting Materials" is as follows: This material contains specific technical terms and abbreviations.

[0388] 1. Material detection and analysis:

[0389] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[0390] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[0391] 2. Terminology database storage and standardization:

[0392] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[0393] 3. Generate information page:

[0394] The server automatically generates information pages based on the extracted terms, adding detailed descriptions and related links.

[0395] 4. Notification of spelling variations:

[0396] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[0397] 5. Real-time annotation during meetings:

[0398] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this in real time as "ROI (Return on Investment)" and displays it.

[0399] 6. Leveraging the Emotion Engine:

[0400] The server analyzes the voice data and text data entered by the user during the meeting and uses an emotion engine to recognize the user's emotions. For example, if the user is feeling stressed during the meeting, the server will provide appropriate feedback to support the progress of the meeting.

[0401] Prompt Sentence Examples

[0402] Here are some example prompts to input to a generative AI model:

[0403] Please analyze the business terms contained in newly uploaded documents and standardize the notation of the extracted business terms to a consistent format.

[0404] "Analyze audio during meetings in real time and annotate and display important business terms."

[0405] "Recognize emotions based on the user's voice and text data and provide appropriate feedback if they are feeling stressed."

[0406] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0407] Step 1:

[0408] The server monitors the file sharing system to detect newly uploaded materials or documents. In this step, the server checks whether a new file has been uploaded from the file sharing system (e.g., Dropbox, Google Drive) and obtains the path and name of the file. Specifically, the server scans the file list and detects the new file "Project Meeting Materials.pdf."

[0409] Input: A file uploaded to a file sharing system.

[0410] Output: The path and name of the file found.

[0411] Step 2:

[0412] The server analyzes the detected file using a natural language processing (NLP) model to extract business terms. In this step, the server uses an NLP model (e.g., Google BERT, OpenAI GPT-3) to list business terms such as "ROI" and "critical path" from the text data in the file. Specifically, the server analyzes "Project Meeting Materials.pdf" and lists business terms.

[0413] Input: The path and name of the detected file.

[0414] Output: A list of extracted business terms.

[0415] Step 3:

[0416] The server stores the extracted business terms in a database, checks whether they match existing terms, and unifies the spelling. In this step, before storing the extracted terms in a database (e.g., MySQL, PostgreSQL), the server compares them with existing terms and unifies any spelling variations. Specifically, it checks whether "ROI" already exists in the database and whether it matches "Return on Investment," and unifies them.

[0417] Input: A list of extracted business terms.

[0418] Output: A consistent list of terms stored in a database.

[0419] Step 4:

[0420] The server automatically generates and updates information pages based on the business terms stored in the database. In this step, the server generates and updates information pages containing detailed explanations and related links based on the newly extracted business terms. Specifically, a detailed explanation of "ROI (Return on Investment)" is added to the information page.

[0421] Input: A consistent list of terms stored in a database.

[0422] Output: Generated and updated information page.

[0423] Step 5:

[0424] The server periodically monitors the information sharing system, detects spelling variations in new documents, and sends a notification to prompt correction. In this step, the server compares the newly detected document with the existing term list, and if a spelling variation is found, it sends a notification to the relevant user. Specifically, if "return on investment" is newly detected, it sends a notification to unify it to "ROI."

[0425] Input: The newly discovered document.

[0426] Output: Notification of spelling variations.

[0427] Step 6:

[0428] The device receives voice input during the meeting and converts the voice data into text. In this step, the device uses speech recognition software (e.g., Google Speech-to-Text, Microsoft Azure Speech Service) to convert the voice during the meeting into text in real time. Specifically, the device converts the part where "ROI" is spoken during the meeting into text.

[0429] Input: Audio input during a meeting.

[0430] Output: The audio data converted to text.

[0431] Step 7:

[0432] The server detects business terms from the converted speech data, generates annotations in real time, and sends them to the terminal. In this step, the server detects "ROI" from the converted data and generates the annotation "ROI (Return on Investment)." After generating the annotation, it sends it to the terminal and displays it. Specifically, the annotation "ROI (Return on Investment)" is displayed on the screen of the conference participants.

[0433] Input: Transcribed audio data.

[0434] Output: The generated annotations.

[0435] Step 8:

[0436] The server analyzes the user's voice and text data and recognizes the user's emotions using an emotion engine. In this step, the server analyzes the user's emotions using an emotion engine (e.g., IBM Watson, Azure Cognitive Services). Based on the recognized emotions, the server provides appropriate feedback and advice. Specifically, if the server determines that the user is feeling stressed, it sends feedback encouraging the user to relax.

[0437] Input: User voice and text data.

[0438] Output: Emotion-based feedback and advice.

[0439] (Application example 2)

[0440] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0441] Modern factories generate a large number of documents and materials every day, and these documents contain a large amount of technical terminology. However, if these terms are not used consistently, it becomes difficult to understand and share information. In addition, there is a lack of a way to recognize workers' emotions and provide appropriate feedback, which can lead to reduced work efficiency and the risk of mistakes. There is a need for a system that can resolve these issues and provide support based on workers' emotions while maintaining consistency in work terminology.

[0442] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for monitoring materials and documents and detecting newly uploaded materials and documents; means for extracting business terms from materials and documents using a natural language processing model; means for storing the extracted business terms in a database, checking for matches with existing terms, and unifying notations; means for automatically generating and updating knowledge pages based on the business terms and adding term explanations; means for periodically monitoring the data storage system, detecting notation inconsistencies in new documents, and sending notifications prompting corrections; means for receiving voice input during work, detecting business terms from the voice data, and generating and displaying annotations in real time; and means for recognizing the user's emotions and providing appropriate instructions and feedback. This makes it possible to improve work efficiency while ensuring consistency in business terminology.

[0443] "Materials and documents" refers to all documents generated within the factory, such as reports, plans, procedures, and manuals.

[0444] "Means for detecting newly uploaded files" refers to the ability to monitor the data storage system and automatically detect newly uploaded files.

[0445] A "natural language processing model" is a software algorithm for analyzing the content of a document and extracting specific business terms.

[0446] "Business terms" refers to technical terms, abbreviations, and keywords used in a particular business or field of expertise.

[0447] "Means of storing in a database, checking for matches with existing terms, and standardizing notation" refers to a function that saves extracted business terms and compares them with existing data to ensure consistency in notation.

[0448] "Means for generating and updating knowledge pages and adding explanations of terms" is a function that automatically creates and updates information pages (Wiki pages) that include detailed explanations of business terms.

[0449] "Means for periodically monitoring the data storage system, detecting inconsistencies in notation in new documents, and sending notifications prompting corrections" refers to a notification function that monitors new document uploads, detects inconsistencies in notation, and requests the user to make corrections.

[0450] "Means for receiving voice input during work, detecting business terms from the voice data, and generating and displaying annotations in real time" refers to a function that analyzes voice information generated during work, recognizes technical terms, and immediately displays their definitions and explanations.

[0451] "Means for recognizing the user's emotions and providing appropriate instructions and feedback" refers to a function that detects the user's emotions from their words and actions, and provides advice and support according to the situation.

[0452] MODE FOR CARRYING OUT THE INVENTION

[0453] This invention is a system that maintains consistency in business terminology within a factory, recognizes the emotions of workers, and provides appropriate feedback. The system's main functions are monitoring materials and documents, extracting business terminology using natural language processing, storing it in a database and standardizing its notation, automatically generating Wiki pages, detecting and notifying inconsistencies in notation, generating annotations in real time during work, and recognizing emotions using an emotion engine.

[0454] System Overview

[0455] 1. Monitoring and detection of materials and documents

[0456] The server monitors the factory's data storage system and automatically detects newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[0457] 2. Extracting business terms using natural language processing models

[0458] The server uses a natural language processing model (NLP model) to analyze materials and documents and extract business terms, a process carried out using automated algorithms.

[0459] 3. Storing terms in the database and standardizing notation

[0460] The extracted business terms are stored in a database by the server. A consistent term list is maintained by checking whether they match existing terms and correcting any inconsistencies in their spellings to make them consistent.

[0461] 4. Automatic generation and updating of knowledge pages

[0462] The server automatically generates and updates knowledge pages based on the business terms stored in the database, adding detailed explanations and related links for each term to make them easily accessible to workers.

[0463] 5. Detecting and notifying inconsistencies

[0464] The server periodically monitors the data storage system and analyzes newly uploaded documents. If it detects any inconsistencies in spelling that do not match the existing terminology list, it notifies the user and prompts them to make corrections.

[0465] 6. Real-time annotation generation while working

[0466] The device receives voice input during work and converts the voice data into text. The server detects business terms from the converted data and generates annotations in real time, which are sent to the device. The device then displays these on the worker's screen to help them understand the work.

[0467] 7. Emotion Recognition with Emotion Engine

[0468] The server analyzes voice and text data during work and uses an emotion engine to recognize the worker's emotions. Based on the recognized emotions, the server provides appropriate feedback and advice.

[0469] Specific examples

[0470] For example, if a factory uploads a document called "New Product Line Plan.pdf", the following process will be performed based on this file:

[0471] 1. The server detects "New Product Line Plan.pdf" and analyzes its contents.

[0472] 2. Use a natural language processing model to extract technical terms such as "TPM (Total Productive Maintenance)."

[0473] 3. Store the extracted business terms in a database, and check for and correct any inconsistencies in notation.

[0474] 4. Automatically generate knowledge pages based on terms and provide related information.

[0475] 5. If there are any inconsistencies in the notation of newly uploaded materials, notify the worker and request corrections.

[0476] 6. When you encounter technical terms while working, their definitions are displayed in real time.

[0477] 7. Provide feedback if workers are stressed.

[0478] Prompt Sentence Examples

[0479] For example, you can input the following prompts into a generative AI model:

[0480] Extract technical terms from the factory's "New Product Line Plan.pdf" and generate knowledge pages based on their definitions. Also, analyze the emotions of workers and display the message "Calm down, everything is under control" when they are feeling stressed.

[0481] The above is a detailed description of the mode for carrying out the invention.

[0482] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0483] Step 1:

[0484] The server monitors the data storage system in the factory to detect newly uploaded materials or documents. When a new document is detected, it obtains its path. This step requires the directory path of the data storage system as input and the path of the detected new document as output.

[0485] Step 2:

[0486] The server uses a natural language processing model (NLP model) to analyze newly detected materials and documents and extract business terms. A list of extracted business terms is generated. This step requires the text data of the document as input and produces a list of business terms as output.

[0487] Step 3:

[0488] The server accesses the database and stores the extracted business terms. It also checks for matches with existing terms and for inconsistencies in spelling. If necessary, it makes corrections to unify spelling. This step requires a list of business terms as input and an updated database as output.

[0489] Step 4:

[0490] The server automatically generates and updates knowledge pages based on the business terms in the database. Specifically, it creates pages that include detailed explanations and related information for each term. This step requires business terms in the database as input, and the generated knowledge pages as output.

[0491] Step 5:

[0492] The server periodically monitors the data storage system and analyzes newly uploaded documents. If it detects a spelling inconsistency that does not match the existing terminology list, it sends a notification to the relevant user to prompt them to correct it. This step requires the new document and the existing terminology list as input, and a spelling inconsistency notification as output.

[0493] Step 6:

[0494] The terminal receives voice input during work and converts the voice data into text. The server then detects business terms from the converted data and generates annotations in real time, which are sent to the terminal. The terminal then displays the annotations on the worker's screen. This step requires voice data as input and produces annotations to be displayed as output.

[0495] Step 7:

[0496] The server analyzes the voice and text data during work and recognizes the worker's emotions using an emotion engine. Based on the recognized emotions, the server provides appropriate feedback and advice. This step requires voice and text data as input, and feedback and advice are obtained as output.

[0497] This will enable the entire system to function consistently, ensuring consistency in business terminology, improving work efficiency, and creating a system that provides support based on the emotions of workers.

[0498] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0499] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0500] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0501] [Second embodiment]

[0502] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0503] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0504] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0505] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0506] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0507] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0508] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0509] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0510] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0511] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0512] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0513] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0514] MODE FOR CARRYING OUT THE INVENTION

[0515] The present invention is a system that uses AI to automatically create a Wiki that summarizes business terminology, eliminating problems such as lack of maintenance and inconsistent spelling. An embodiment of this system is described in detail below.

[0516] System Overview

[0517] The system has the following main functions:

[0518] 1. Monitoring and detection of materials and documents

[0519] 2. Extracting business terms using NLP models

[0520] 3. Storing terms in the database and standardizing notation

[0521] 4. Automatic generation and updating of Wiki pages

[0522] 5. Spelling Variation Detection and Notification

[0523] 6. Real-time annotation generation during meetings

[0524] Monitoring and detection of materials and documents

[0525] The server monitors the file sharing system and automatically detects newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[0526] Extracting business terms using NLP models

[0527] The server uses natural language processing (NLP) models to analyze materials and documents and extract business terms, abbreviations, and technical terms, which then lists specific business terms and allows for efficient analysis.

[0528] Terminology storage in database and standardization of notation

[0529] The server stores the extracted business terms in a database, checks whether they match existing terms, and corrects any inconsistencies in spelling to create a consistent term list.

[0530] Automatic generation and updating of Wiki pages

[0531] The server automatically generates and updates Wiki pages based on business terms stored in the database, adding detailed explanations and related links for each term to make them easily accessible to employees.

[0532] Spelling variation detection and notification

[0533] The server periodically monitors the file sharing system and analyzes newly uploaded documents. If it detects a spelling variation that does not match the existing terminology list, it notifies the user and encourages them to correct the spelling.

[0534] Real-time annotation during meetings

[0535] The device receives voice input during the meeting and converts the voice data into text. The server detects business terms from the converted data and generates annotations in real time, which are sent to the device. The device then displays the annotations on the screen of the meeting participants to support understanding.

[0536] Specific examples

[0537] Consider the process flow when a newly assigned employee downloads "Project Meeting Materials." This material contains specific technical terms and abbreviations.

[0538] 1. Material detection and analysis:

[0539] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[0540] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[0541] 2. Terminology database storage and standardization:

[0542] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[0543] 3. Generate Wiki pages:

[0544] The server automatically generates Wiki pages based on the extracted terms, adding detailed descriptions and related links.

[0545] 4. Notification of spelling variations:

[0546] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[0547] 5. Real-time annotation during meetings:

[0548] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this as "ROI (Return on Investment)" and displays it in real time.

[0549] In this way, the system helps new employees quickly adapt to their work and contributes to standardizing and promoting understanding of work terminology.

[0550] The processing flow will be explained below.

[0551] Program processing

[0552] 1. Monitoring and detection of materials and documents

[0553] Step 1:

[0554] The server accesses the file sharing system and monitors newly uploaded materials and documents.

[0555] Step 2:

[0556] The server scans the directory of the file sharing system at regular time intervals to check whether new files exist.

[0557] Step 3:

[0558] The server detects the newly uploaded file and retrieves the file's metadata (such as name, upload date and time, file type, etc.).

[0559] 2. Extracting business terms using NLP models

[0560] Step 4:

[0561] The server inputs the detected new files into a natural language processing (NLP) model.

[0562] Step 5:

[0563] The server uses NLP models to analyze the content of the files and extract business terms, abbreviations, and technical terms.

[0564] Step 6:

[0565] The server organizes the extracted business terms in list format and temporarily stores the analysis results.

[0566] 3. Storing terms in the database and standardizing notation

[0567] Step 7:

[0568] The server stores the extracted business terms in a database.

[0569] Step 8:

[0570] The server checks the newly extracted terms against existing terms in the database to see if there is a match.

[0571] Step 9:

[0572] If the term is spelled differently, the server will correct it to a unified spelling for consistency.

[0573] 4. Automatic generation and updating of Wiki pages

[0574] Step 10:

[0575] The server runs a script that automatically generates Wiki pages based on business terms in the database.

[0576] Step 11:

[0577] The server adds detailed descriptions and related links to each term to the Wiki page.

[0578] Step 12:

[0579] The server uploads the generated Wiki page to the company's internal Wiki system so that users can view it.

[0580] 5. Spelling Variation Detection and Notification

[0581] Step 13:

[0582] The server periodically monitors the file sharing system and analyzes newly uploaded documents.

[0583] Step 14:

[0584] The server detects spelling variations when a term in the parsed document does not match an existing list of terms.

[0585] Step 15:

[0586] When the server detects a spelling variation, it sends a notification email to the relevant user, prompting them to correct the spelling to the recommended spelling.

[0587] 6. Real-time annotation generation during meetings

[0588] Step 16:

[0589] The terminal receives voice input during the conference and acquires the voice data.

[0590] Step 17:

[0591] The server passes the received voice data to a voice recognition service and converts it into text.

[0592] Step 18:

[0593] The server detects business terms from the text data and generates annotations in real time.

[0594] Step 19:

[0595] The terminal displays the generated annotations on the screens of the conference participants to support their understanding of the terminology.

[0596] Example 1

[0597] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0598] Inconsistent notation and lack of maintenance in materials and documents are often problems in business activities. New employees and those transferred from other departments often have difficulty understanding business terms and abbreviations, which can lead to reduced work efficiency. Also, the use of technical terms during meetings often reduces participants' understanding. It is necessary to solve these problems and provide information quickly and clearly while maintaining consistency in business terminology.

[0599] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0600] In this invention, the server includes a means for monitoring documents and automatically detecting newly uploaded documents, a means for extracting business terms from documents using a natural language processing model, and a means for storing the extracted business terms in a database, checking for matches with existing terms, and standardizing their spellings to a consistent format. This allows for the consolidation and standardization of business terms. The server also includes a means for automatically generating and updating Wiki pages based on business terms and adding term explanations, a means for periodically monitoring the file system, detecting spelling variations in new documents, and sending notifications prompting corrections. The server also includes a means for receiving voice input during meetings, detecting business terms from the voice data, and generating and displaying annotations in real time. This allows for the rapid sharing and understanding of new information, improving work efficiency and accuracy.

[0601] "Materials" refers to information media such as documents, files, and reports used in business activities.

[0602] A "natural language processing model" refers to a model that makes full use of artificial intelligence technology to analyze and understand human language.

[0603] "Business jargon" refers to specialized words and abbreviations used in specific industries or jobs.

[0604] A "database" refers to a system for organizing and storing extracted information and data so that it can be used efficiently.

[0605] "Orthographic variation" refers to differences between words or terms that have the same meaning but are expressed in different forms.

[0606] A "Wiki page" refers to a web page that can be collaboratively edited on the Internet or an intranet.

[0607] A "file system" refers to the method and implementation for managing and storing data and files.

[0608] "Annotation" refers to explanations or explanatory text added to aid understanding.

[0609] "Voice input" refers to a method of inputting voice information into a computer via a microphone or the like.

[0610] MODE FOR CARRYING OUT THE INVENTION

[0611] The present invention is a system that uses AI to automatically create summaries of business terms and resolves issues such as lack of maintenance and inconsistent spelling. Specific embodiments of this system are described in detail below.

[0612] System configuration

[0613] This system is mainly composed of three elements: a server, a terminal, and a user. The server processes and monitors data, the terminal functions as an input and display device, and the user provides and uses information.

[0614] Key Features

[0615] Material monitoring and detection

[0616] The server monitors the file sharing system and automatically detects newly uploaded files. Specifically, the server uses a Python script to check the file system for changes at regular intervals and obtain a list of new or updated files. This operation utilizes the OS's file monitoring function and the cloud service's API.

[0617] Extracting business terms using NLP models

[0618] The server analyzes the documents using a natural language processing (NLP) model to extract business terms. This operation uses a pre-trained NLP model (e.g., BERT or spaCy). Depending on the file format, the server converts the file (e.g., PDF or Word) into text format, and then uses the NLP model to analyze and extract key terms.

[0619] Terminology storage in database and standardization of notation

[0620] The server stores the extracted business terms in a dedicated database. Here, it checks whether they match existing terms and standardizes them into consistent notations. MySQL, PostgreSQL, or other database systems are used. The server executes queries to check and update the contents of the database.

[0621] Automatic generation and updating of Wiki pages

[0622] The server automatically generates and updates Wiki pages based on business terms stored in the database. Wiki pages are generated using Markdown format or HTML templates. Specifically, the server retrieves terminology information from the database and applies it to templates to generate Wiki pages.

[0623] Spelling variation detection and notification

[0624] The server periodically monitors the file sharing system and compares new material with the existing database. If a spelling variation is detected, the server notifies the user via email or instant messaging services (e.g., Slack).

[0625] Real-time annotation during meetings

[0626] The device receives voice input during the meeting and converts the voice data into text. The Google Speech-to-Text API is used for speech recognition. For example, if someone says "ROI" during a meeting, the server analyzes it in real time and generates the annotation "ROI (Return on Investment)." The generated annotation is then displayed on the device of each meeting participant.

[0627] Specific examples

[0628] The following shows the process flow when a newly assigned employee downloads "Project Meeting Materials."

[0629] 1. Material detection and analysis:

[0630] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[0631] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[0632] 2. Terminology database storage and standardization:

[0633] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[0634] 3. Generate Wiki pages:

[0635] The server automatically generates Wiki pages based on the extracted terms, adding detailed descriptions and related links.

[0636] 4. Notification of spelling variations:

[0637] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[0638] 5. Real-time annotation during meetings:

[0639] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this as "ROI (Return on Investment)" and displays it in real time.

[0640] Prompt Sentence Examples

[0641] 1. Material detection and analysis:

[0642] "Detect newly uploaded conference materials and extract the technical terms they contain."

[0643] 2. Terminology database storage and standardization:

[0644] "Store the extracted terms in a database and check if they match existing terms."

[0645] 3. Generate Wiki pages:

[0646] "Generate wiki pages based on terms stored in your database and add detailed descriptions and related links."

[0647] 4. Notification of spelling variations:

[0648] "If you detect spelling variations in newly uploaded materials, please notify us of the recommended spelling."

[0649] 5. Real-time annotation during meetings:

[0650] "Extract technical terms from audio during meetings in real time and generate and display annotations."

[0651] The above describes a specific embodiment of the present invention. This system makes it possible to standardize business terminology and promote understanding, thereby realizing efficient business operations.

[0652] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0653] Step 1: Monitoring and detecting material

[0654] The server monitors the file sharing system. Specifically, the server scans the contents of the file system at regular intervals to detect newly added or updated files. The input is the file system path and the monitoring interval. The output is a list of new or updated files. The server also records file metadata (file name, creation date and time, and update date and time).

[0655] Step 2: Prepare the file for analysis

[0656] The server obtains the paths of the detected files and prepares them for analysis. First, it checks the file format (PDF, Word, etc.) and converts it to text data using an appropriate text conversion tool (e.g., PyPDF2 library for PDF, python-docx library for Word). The input is a list of detected files, and the output is text data.

[0657] Step 3: Extracting business terms using an NLP model

[0658] The server inputs the converted text data into an NLP model to extract business terms, abbreviations, and technical terms. This process uses a pre-trained NLP model (e.g., BERT or spaCy). The input is the text data, and the output is a list of extracted business terms. Specifically, terms such as "ROI" and "critical path" are extracted.

[0659] Step 4: Store business terms in a database

[0660] The server stores the extracted business terms in a database. During this process, it checks for matches with existing terms and standardizes the notation. The database used is MySQL or PostgreSQL. The input is the extracted business term list, and the output is the updated database. For example, if "ROI" and "return on investment" match, they are unified to "ROI."

[0661] Step 5: Automatically generate and update Wiki pages

[0662] The server automatically generates and updates Wiki pages based on business terms stored in the database. Markdown format and HTML templates are used for this process. The input is the business term information in the database, and the output is the generated or updated Wiki page. For example, the "ROI" page contains the description "Return on Investment."

[0663] Step 6: Detecting and reporting spelling variations

[0664] The server compares newly uploaded materials with the existing database to detect variations in notation. If any discrepancies are found during this process, a notification is sent to the user. Notification methods include email and instant messaging services (e.g., Slack). The input is the new material and database information, and the output is a notification to the user. For example, a notification recommending that "return on investment" be changed to "ROI" is sent.

[0665] Step 7: Real-time annotation during the meeting

[0666] The device receives voice input during the meeting and converts the voice data into text. This process uses the Google Speech-to-Text API. The input is the voice data from the meeting, and the output is text data. The converted text data is sent to the server, and if a business term is detected, an annotation is generated in real time and displayed on the device. For example, if "ROI" is spoken during a meeting, the annotation "ROI (Return on Investment)" is generated and displayed.

[0667] (Application example 1)

[0668] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0669] In today's business environment, multiple documents and materials are frequently uploaded, often containing many technical terms and abbreviations. This makes it difficult for newly assigned employees and external participants to quickly adapt to the work. Furthermore, inconsistencies in the spelling of business terms occur, reducing the consistency and efficiency of information. There is a need to solve these problems and improve business efficiency.

[0670] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0671] In this invention, the server includes: means for monitoring materials and documents and detecting newly uploaded materials and documents; means for extracting business terms from materials and documents using a natural language processing model; means for storing the extracted business terms in a database, checking for matches with existing terms, and standardizing spellings; means for automatically generating and updating information pages based on the business terms and adding term explanations; means for periodically monitoring the file sharing system, detecting spelling variations in new documents, and sending notifications prompting corrections; means for receiving voice input during meetings, detecting business terms from the voice data, and generating and displaying annotations in real time; means for analyzing the camera feed of a visual device and extracting text from documents; means for displaying business terms as annotations from the extracted text; and means for analyzing the voice input of the visual device, performing speech recognition, and providing detailed explanations of the relevant business terms by voice. This enables newly assigned employees to quickly adapt to their work, promotes standardization of business terminology, and promotes understanding.

[0672] "Materials" refers to documents and other information related to business.

[0673] "Document" refers to a medium containing information stored as printed material or digital data.

[0674] "Monitoring" means that a system periodically or continuously checks a particular file sharing system or database to detect newly added information.

[0675] A "natural language processing model" refers to a computer program or algorithm that analyzes human language and extracts useful information from text data.

[0676] "Business jargon" refers to technical terms and abbreviations frequently used in a particular business or industry.

[0677] A "database" refers to a system for efficiently storing, retrieving, and managing information.

[0678] "Unification of spelling" refers to the standardization of words with the same meaning that are expressed in different spellings or writing styles into a consistent format.

[0679] An "information page" is a web page or document that provides detailed explanations and related information about a particular term or concept.

[0680] A "file sharing system" refers to a system that allows multiple users within a specific network to share and access files.

[0681] "Spelling variation" refers to the situation where the same term or concept is written in different ways.

[0682] "Notification" refers to a system-generated message that informs the user of a particular event or situation.

[0683] "Voice input during a meeting" refers to the process by which voice data spoken during a meeting is received by the system.

[0684] "Audio Data" refers to a digital recording of sound collected by a microphone or the like.

[0685] "Generating and displaying annotations in real time" refers to the process of instantly analyzing voice or text input and displaying additional information based on the results.

[0686] "Visual devices" refer to devices worn by users that visually present information, such as smart glasses and head-mounted displays.

[0687] "Camera feed" refers to video data captured by a camera in real time.

[0688] "Extracting text from a document" refers to the process of analyzing character information from camera images or scanned data and extracting that text data.

[0689] "Speech recognition" refers to the technology and algorithms that analyze audio signals and identify meaningful words and sentences from them.

[0690] This invention is a system that uses AI to consolidate business terminology, particularly to improve efficiency in factory work environments. This system is composed of a server, terminals, and users. The specific configuration and functions of this system are described below.

[0691] System configuration and functions

[0692] 1. Monitoring and detection of materials and documents

[0693] The server monitors the factory's file sharing system and database, automatically detecting newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[0694] 2. Extracting business terms using natural language processing models

[0695] The server uses natural language processing (NLP) models to analyze documents and extract business terms, abbreviations, and technical terms, using NLP libraries such as SpaCy.

[0696] 3. Storing terms in the database and standardizing notation

[0697] The server stores the extracted business terms in a database and checks whether they match existing terms. If there are any variations in spelling, they are unified to create a consistent term list. For this purpose, an SQL-based database management system (e.g., MySQL) is used.

[0698] 4. Automatic generation and updating of information pages

[0699] The server automatically generates and updates information pages based on business terms stored in the database, including detailed explanations and related links for each term.

[0700] 5. Spelling Variation Detection and Notification

[0701] The server periodically monitors the file sharing system and analyzes newly uploaded documents. If it detects a spelling variation that does not match the existing term list, it sends a notification to the relevant user urging them to correct the variation. The notification system includes email notifications and push notifications.

[0702] 6. Real-time voice annotation

[0703] The device receives voice input during the meeting and converts the voice data into text. The server detects business terms from the text and generates annotations in real time and sends them to the device, allowing meeting participants to instantly understand the meaning of the terms.

[0704] 7. Visual Device Camera Feed Analysis

[0705] A visual device (e.g., smart glasses) extracts text from a document through a camera feed. After extracting the text using Tesseract OCR, the server analyzes the text and displays real-time annotations for business terms.

[0706] 8. Analysis and presentation of voice input

[0707] The voice input from the visual device is analyzed and speech recognition is performed using the Google Cloud Speech-to-Text API. For recognized terms, the server retrieves detailed descriptions and presents them as audio through the visual device's speaker.

[0708] Specific examples

[0709] For example, suppose a newly assigned employee views a document containing the term "ROI (Return on Investment)." The camera in the vision device captures this, and the server performs text analysis, displaying the annotation "Return on Investment" for the term "ROI." Furthermore, if the employee voice-inputs, "Tell me more about Return on Investment," the vision device recognizes the voice and provides a detailed explanation.

[0710] This will generate a prompt like this:

[0711] "A new document has been uploaded. This document contains the term 'return on investment'. Could you please explain this term in more detail?"

[0712] In this way, the present invention promotes understanding of business terms in a factory work environment, enabling efficient work execution.

[0713] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0714] Step 1:

[0715] The server monitors the file sharing system within the factory and detects newly uploaded materials and documents. As input, it receives directory information and file lists from the file sharing system and identifies new files based on timestamps. After processing this data, it outputs a list of newly uploaded files.

[0716] Step 2:

[0717] The server uses a natural language processing model (e.g., SpaCy) to extract business terms from newly detected materials and documents. As input, it receives the contents of the files detected in step 1 in text format and analyzes the text data. This data processing results in the output of a list of extracted business terms.

[0718] Step 3:

[0719] The server stores the extracted business terms in a database, checks for matches with existing terms, and standardizes notation. As input, it receives the list of business terms extracted in step 2 and compares them with existing terms in the database. This data processing results in a consistent list of terms being output.

[0720] Step 4:

[0721] The server automatically generates and updates information pages based on business terms stored in the database. It receives an updated term list as input and adds detailed explanations and related links for each term. This data calculation results in the output of the latest information page.

[0722] Step 5:

[0723] The server periodically monitors the file sharing system, detects spelling variations that do not match the existing terminology list, and sends a notification prompting correction. It receives the content of newly uploaded documents as input and compares it with the existing terminology list. After processing this data, it outputs a notification containing mismatched terms and suggested corrections.

[0724] Step 6:

[0725] The device converts voice input received during the meeting into text and sends it to the server. It receives the voice data of the meeting participants as input and performs speech recognition using the Google Cloud Speech-to-Text API. This data is then processed and the text data is output.

[0726] Step 7:

[0727] The server generates annotations in real time for business terms detected from the voice input and sends them to the terminal. As input, it receives the text data output in step 6, analyzes and extracts business terms, and generates annotations for them. This data processing results in the output of annotated text.

[0728] Step 8:

[0729] The vision device analyzes the camera feed and extracts text from the document. It receives the video data from the camera feed as input and performs character recognition using Tesseract OCR. It then processes this data and outputs the extracted text data.

[0730] Step 9:

[0731] The server generates annotations of business terms from the extracted text and displays them on a visual device. It receives the text data output in step 8 as input, analyzes and extracts business terms, and generates annotations for them. This data processing results in the output of annotated text.

[0732] Step 10:

[0733] The visual device analyzes voice input and performs speech recognition. It receives the user's voice data as input and performs speech recognition using the Google Cloud Speech-to-Text API. It then processes this data and outputs the recognized text data.

[0734] Step 11:

[0735] The server obtains detailed explanations for the recognized terms and presents them as audio through the speaker of the visual device. As input, it receives the text data output in step 10, generates detailed explanations, and sends them as audio data to the visual device. This data processing results in the audio data of the detailed explanation being output.

[0736] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0737] MODE FOR CARRYING OUT THE INVENTION

[0738] The present invention is a system that uses AI to automatically create a wiki that summarizes business terminology, resolves issues such as lack of maintenance and inconsistent spelling, and combines it with an emotion engine that recognizes user emotions. An embodiment of this system is described in detail below.

[0739] System Overview

[0740] The system has the following main functions:

[0741] 1. Monitoring and detection of materials and documents

[0742] 2. Extracting business terms using NLP models

[0743] 3. Storing terms in the database and standardizing notation

[0744] 4. Automatic generation and updating of Wiki pages

[0745] 5. Spelling Variation Detection and Notification

[0746] 6. Real-time annotation generation during meetings

[0747] 7. Emotion engine that recognizes user emotions

[0748] Monitoring and detection of materials and documents

[0749] The server monitors the file sharing system and automatically detects newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[0750] Extracting business terms using NLP models

[0751] The server uses natural language processing (NLP) models to analyze materials and documents and extract business terms, abbreviations, and technical terms, which then lists specific business terms and allows for efficient analysis.

[0752] Terminology storage in database and standardization of notation

[0753] The server stores the extracted business terms in a database, checks whether they match existing terms, and corrects any inconsistencies in spelling to create a consistent term list.

[0754] Automatic generation and updating of Wiki pages

[0755] The server automatically generates and updates Wiki pages based on business terms stored in the database, adding detailed explanations and related links for each term to make them easily accessible to employees.

[0756] Spelling variation detection and notification

[0757] The server periodically monitors the file sharing system and analyzes newly uploaded documents. If it detects a spelling variation that does not match the existing terminology list, it notifies the user and encourages them to correct the spelling.

[0758] Real-time annotation during meetings

[0759] The device receives voice input during the meeting and converts the voice data into text. The server detects business terms from the converted data and generates annotations in real time, which are sent to the device. The device then displays the annotations on the screen of the meeting participants to support understanding.

[0760] Emotion engine that recognizes user emotions

[0761] The server analyzes the user's voice and text data, recognizes the user's emotions using an emotion engine, and provides appropriate feedback and advice based on the recognized emotions.

[0762] Specific examples

[0763] Consider the process flow when a newly assigned employee downloads "Project Meeting Materials." This material contains specific technical terms and abbreviations.

[0764] 1. Material detection and analysis:

[0765] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[0766] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[0767] 2. Terminology database storage and standardization:

[0768] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[0769] 3. Generate Wiki pages:

[0770] The server automatically generates Wiki pages based on the extracted terms, adding detailed descriptions and related links.

[0771] 4. Notification of spelling variations:

[0772] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[0773] 5. Real-time annotation during meetings:

[0774] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this as "ROI (Return on Investment)" and displays it in real time.

[0775] 6. Leveraging the Emotion Engine:

[0776] The server analyzes the voice data and text data entered by the user during the meeting and uses an emotion engine to recognize the user's emotions. For example, if the user is feeling stressed during the meeting, the server will provide appropriate feedback to support the progress of the meeting.

[0777] The processing flow will be explained below.

[0778] MODE FOR CARRYING OUT THE INVENTION

[0779] Monitoring and detection of materials and documents

[0780] Step 1:

[0781] The server accesses the file sharing system and monitors newly uploaded materials and documents.

[0782] Step 2:

[0783] The server scans the directory of the file sharing system at regular time intervals to check whether new files exist.

[0784] Step 3:

[0785] The server detects the newly uploaded file and retrieves the file's metadata (such as name, upload date and time, file type, etc.).

[0786] Extracting business terms using NLP models

[0787] Step 4:

[0788] The server inputs the detected new files into a natural language processing (NLP) model.

[0789] Step 5:

[0790] The server uses NLP models to analyze the content of the files and extract business terms, abbreviations, and technical terms.

[0791] Step 6:

[0792] The server organizes the extracted business terms in list format and temporarily stores the analysis results.

[0793] Terminology storage in database and standardization of notation

[0794] Step 7:

[0795] The server stores the extracted business terms in a database.

[0796] Step 8:

[0797] The server checks the newly extracted terms against existing terms in the database to see if there is a match.

[0798] Step 9:

[0799] If the term is spelled differently, the server will correct it to a unified spelling for consistency.

[0800] Automatic generation and updating of Wiki pages

[0801] Step 10:

[0802] The server runs a script that automatically generates Wiki pages based on business terms in the database.

[0803] Step 11:

[0804] The server adds detailed descriptions and related links to each term to the Wiki page.

[0805] Step 12:

[0806] The server uploads the generated Wiki page to the company's internal Wiki system so that users can view it.

[0807] Spelling variation detection and notification

[0808] Step 13:

[0809] The server periodically monitors the file sharing system and analyzes newly uploaded documents.

[0810] Step 14:

[0811] The server detects spelling variations when a term in the parsed document does not match an existing list of terms.

[0812] Step 15:

[0813] When the server detects a spelling variation, it sends a notification email to the relevant user, prompting them to correct the spelling to the recommended spelling.

[0814] Real-time annotation during meetings

[0815] Step 16:

[0816] The terminal receives voice input during the conference and acquires the voice data.

[0817] Step 17:

[0818] The server passes the received voice data to a voice recognition service and converts it into text.

[0819] Step 18:

[0820] The server detects business terms from the text data and generates annotations in real time.

[0821] Step 19:

[0822] The terminal displays the generated annotations on the screens of the conference participants to support their understanding of the terminology.

[0823] Emotion engine that recognizes user emotions

[0824] Step 20:

[0825] The terminal collects voices during the conference and text data entered by the user.

[0826] Step 21:

[0827] The server passes the collected voice and text data to an emotion engine to analyze the user's emotions.

[0828] Step 22:

[0829] The server generates appropriate feedback and advice based on the recognized emotions.

[0830] Step 23:

[0831] The terminal displays the generated feedback and advice on the screens of the conference participants to support the progress of the conference.

[0832] Example 2

[0833] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0834] When managing materials and documents, it is necessary to efficiently monitor and detect newly uploaded information. There is also the issue of eliminating inconsistencies in notation due to the existence of unstandardized business terms and abbreviations, and creating and maintaining a consistent terminology list. Furthermore, there is a need to streamline the provision of information during meetings and to implement a feedback system based on user sentiment.

[0835] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0836] In this invention, the server includes: means for monitoring materials and documents and detecting newly uploaded materials and documents; means for extracting business terms from materials and documents using a natural language processing model; means for storing the extracted business terms in a database, checking for matches with existing terms, and standardizing notation; means for automatically generating and updating information pages based on the business terms and adding explanations of the terms; means for periodically monitoring the information sharing system, detecting notation variations in new documents, and sending notifications prompting corrections; means for receiving voice input during a meeting, detecting business terms from the voice data, and generating and displaying annotations in real time; and means for analyzing user voice and text data, recognizing user emotions, and providing feedback and advice. This enables efficient document management, standardization of notation, maintaining consistency, efficient information provision during meetings, and appropriate feedback based on user emotions.

[0837] "Materials and documents" refers to electronic files containing text data and graphical data, such as reports, plans, technical manuals, procedure manuals, e-mails, memos, and meeting materials used inside and outside a company.

[0838] "Monitoring" means constantly checking operations such as adding, changing, or deleting files that are performed on a specific system or network, and detecting them when certain conditions are met.

[0839] A "natural language processing model" refers to an algorithm or machine learning model designed to understand, analyze, and generate human language, specifically extracting business terms and analyzing text.

[0840] "Business jargon" refers to the technical terms, abbreviations, and phrases used within a particular company or industry that are an important part of business processes and communication.

[0841] "Database" refers to a data storage system that stores data in an organized manner and is designed to make it easy to access and manage.

[0842] "Unification of notation" refers to the process of changing identical terms with different notations or expressions into a consistent format to maintain data integrity and consistency.

[0843] An "information page" is a web page or digital document that aggregates business terms and related information and displays them in an easy-to-reference format.

[0844] An "information sharing system" is a platform that allows multiple users to upload, download, and share files and information, and includes cloud storage and corporate network drives.

[0845] "Spelling variation" refers to the use of different expressions or spellings that have the same meaning, which creates problems that make understanding and management complicated.

[0846] "Sending a notification" refers to the act of communicating alerts or information to a user based on a specific event or condition detected by the system.

[0847] "Receiving audio input" refers to the process of collecting audio data through a microphone or other audio capture device and passing that data to the system.

[0848] "Audio data" is data that is a digital representation of human speech recorded via a voice input device.

[0849] "Generating annotations in real time" refers to the process of instantly adding relevant explanations and information to input speech or text data and displaying them.

[0850] "Recognizing user emotions" refers to the technical process of determining the emotions a user is feeling at that time through analysis of voice and text data.

[0851] "Providing feedback and advice" refers to the process by which the system suggests appropriate responses or recommended actions based on the user's perceived emotions.

[0852] MODE FOR CARRYING OUT THE INVENTION

[0853] The present invention is a system that uses AI to automatically create information pages summarizing business terms, eliminating problems such as lack of maintenance and inconsistent spelling, and combining this with an emotion engine that recognizes user emotions. An embodiment of this system is described in detail below.

[0854] Hardware and software used

[0855] The system utilizes the following hardware and software:

[0856] Hardware: Server, terminal, audio input device (microphone)

[0857] File sharing systems: Dropbox, Google Drive, etc.

[0858] Natural language processing models: Google BERT, OpenAI GPT-3, etc.

[0859] Database: MySQL, PostgreSQL

[0860] Emotion engine: IBM Watson, Azure Cognitive Services

[0861] Speech recognition software: Google Speech-to-Text, Microsoft Azure Speech Service

[0862] Overall flow

[0863] The server monitors the file sharing system to detect and analyze newly uploaded materials and documents. A natural language processing model is used for the analysis, which extracts business terms, abbreviations, and technical terms from the materials. Before storing the extracted terms in the database, they are compared with existing terms, and any variations in spelling are unified.

[0864] Next, the server automatically generates and updates information pages based on the business terms stored in the database. The generated information pages include detailed explanations of the terms and related links. The server also periodically monitors the information sharing system to detect spelling variations in new documents. If any are detected, a notification is sent to the relevant user.

[0865] During a meeting, the device receives voice input and converts the voice data into text. The text data is then sent to a server, which detects business terms and generates annotations in real time. The annotations are then sent to the device and displayed on the screens of meeting participants.

[0866] Finally, the server analyzes the user's voice and text data, and uses an emotion engine to recognize the user's emotions. Based on the recognized emotions, the server provides appropriate feedback and advice.

[0867] Specific examples

[0868] The process flow when a newly assigned employee downloads "Project Meeting Materials" is as follows: This material contains specific technical terms and abbreviations.

[0869] 1. Material detection and analysis:

[0870] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[0871] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[0872] 2. Terminology database storage and standardization:

[0873] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[0874] 3. Generate information page:

[0875] The server automatically generates information pages based on the extracted terms, adding detailed descriptions and related links.

[0876] 4. Notification of spelling variations:

[0877] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[0878] 5. Real-time annotation during meetings:

[0879] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this in real time as "ROI (Return on Investment)" and displays it.

[0880] 6. Leveraging the Emotion Engine:

[0881] The server analyzes the voice data and text data entered by the user during the meeting and uses an emotion engine to recognize the user's emotions. For example, if the user is feeling stressed during the meeting, the server will provide appropriate feedback to support the progress of the meeting.

[0882] Prompt Sentence Examples

[0883] Here are some example prompts to input to a generative AI model:

[0884] Please analyze the business terms contained in newly uploaded documents and standardize the notation of the extracted business terms to a consistent format.

[0885] "Analyze audio during meetings in real time and annotate and display important business terms."

[0886] "Recognize emotions based on the user's voice and text data and provide appropriate feedback if they are feeling stressed."

[0887] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0888] Step 1:

[0889] The server monitors the file sharing system to detect newly uploaded materials or documents. In this step, the server checks whether a new file has been uploaded from the file sharing system (e.g., Dropbox, Google Drive) and obtains the path and name of the file. Specifically, the server scans the file list and detects the new file "Project Meeting Materials.pdf."

[0890] Input: A file uploaded to a file sharing system.

[0891] Output: The path and name of the file found.

[0892] Step 2:

[0893] The server analyzes the detected file using a natural language processing (NLP) model to extract business terms. In this step, the server uses an NLP model (e.g., Google BERT, OpenAI GPT-3) to list business terms such as "ROI" and "critical path" from the text data in the file. Specifically, the server analyzes "Project Meeting Materials.pdf" and lists business terms.

[0894] Input: The path and name of the detected file.

[0895] Output: A list of extracted business terms.

[0896] Step 3:

[0897] The server stores the extracted business terms in a database, checks whether they match existing terms, and unifies the spelling. In this step, before storing the extracted terms in a database (e.g., MySQL, PostgreSQL), the server compares them with existing terms and unifies any spelling variations. Specifically, it checks whether "ROI" already exists in the database and whether it matches "Return on Investment," and unifies them.

[0898] Input: A list of extracted business terms.

[0899] Output: A consistent list of terms stored in a database.

[0900] Step 4:

[0901] The server automatically generates and updates information pages based on the business terms stored in the database. In this step, the server generates and updates information pages containing detailed explanations and related links based on the newly extracted business terms. Specifically, a detailed explanation of "ROI (Return on Investment)" is added to the information page.

[0902] Input: A consistent list of terms stored in a database.

[0903] Output: Generated and updated information page.

[0904] Step 5:

[0905] The server periodically monitors the information sharing system, detects spelling variations in new documents, and sends a notification to prompt correction. In this step, the server compares the newly detected document with the existing term list, and if a spelling variation is found, it sends a notification to the relevant user. Specifically, if "return on investment" is newly detected, it sends a notification to unify it to "ROI."

[0906] Input: The newly discovered document.

[0907] Output: Notification of spelling variations.

[0908] Step 6:

[0909] The device receives voice input during the meeting and converts the voice data into text. In this step, the device uses speech recognition software (e.g., Google Speech-to-Text, Microsoft Azure Speech Service) to convert the voice during the meeting into text in real time. Specifically, the device converts the part where "ROI" is spoken during the meeting into text.

[0910] Input: Audio input during a meeting.

[0911] Output: The audio data converted to text.

[0912] Step 7:

[0913] The server detects business terms from the converted speech data, generates annotations in real time, and sends them to the terminal. In this step, the server detects "ROI" from the converted data and generates the annotation "ROI (Return on Investment)." After generating the annotation, it sends it to the terminal and displays it. Specifically, the annotation "ROI (Return on Investment)" is displayed on the screen of the conference participants.

[0914] Input: Transcribed audio data.

[0915] Output: The generated annotations.

[0916] Step 8:

[0917] The server analyzes the user's voice and text data and recognizes the user's emotions using an emotion engine. In this step, the server analyzes the user's emotions using an emotion engine (e.g., IBM Watson, Azure Cognitive Services). Based on the recognized emotions, the server provides appropriate feedback and advice. Specifically, if the server determines that the user is feeling stressed, it sends feedback encouraging the user to relax.

[0918] Input: User voice and text data.

[0919] Output: Emotion-based feedback and advice.

[0920] (Application example 2)

[0921] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0922] Modern factories generate a large number of documents and materials every day, and these documents contain a large amount of technical terminology. However, if these terms are not used consistently, it becomes difficult to understand and share information. In addition, there is a lack of a way to recognize workers' emotions and provide appropriate feedback, which can lead to reduced work efficiency and the risk of mistakes. There is a need for a system that can resolve these issues and provide support based on workers' emotions while maintaining consistency in work terminology.

[0923] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for monitoring materials and documents and detecting newly uploaded materials and documents; means for extracting business terms from materials and documents using a natural language processing model; means for storing the extracted business terms in a database, checking for matches with existing terms, and unifying notations; means for automatically generating and updating knowledge pages based on the business terms and adding term explanations; means for periodically monitoring the data storage system, detecting notation inconsistencies in new documents, and sending notifications prompting corrections; means for receiving voice input during work, detecting business terms from the voice data, and generating and displaying annotations in real time; and means for recognizing the user's emotions and providing appropriate instructions and feedback. This makes it possible to improve work efficiency while ensuring consistency in business terminology.

[0924] "Materials and documents" refers to all documents generated within the factory, such as reports, plans, procedures, and manuals.

[0925] "Means for detecting newly uploaded files" refers to the ability to monitor the data storage system and automatically detect newly uploaded files.

[0926] A "natural language processing model" is a software algorithm for analyzing the content of a document and extracting specific business terms.

[0927] "Business terms" refers to technical terms, abbreviations, and keywords used in a particular business or field of expertise.

[0928] "Means of storing in a database, checking for matches with existing terms, and standardizing notation" refers to a function that saves extracted business terms and compares them with existing data to ensure consistency in notation.

[0929] "Means for generating and updating knowledge pages and adding explanations of terms" is a function that automatically creates and updates information pages (Wiki pages) that include detailed explanations of business terms.

[0930] "Means for periodically monitoring the data storage system, detecting inconsistencies in notation in new documents, and sending notifications prompting corrections" refers to a notification function that monitors new document uploads, detects inconsistencies in notation, and requests the user to make corrections.

[0931] "Means for receiving voice input during work, detecting business terms from the voice data, and generating and displaying annotations in real time" refers to a function that analyzes voice information generated during work, recognizes technical terms, and immediately displays their definitions and explanations.

[0932] "Means for recognizing the user's emotions and providing appropriate instructions and feedback" refers to a function that detects the user's emotions from their words and actions, and provides advice and support according to the situation.

[0933] MODE FOR CARRYING OUT THE INVENTION

[0934] This invention is a system that maintains consistency in business terminology within a factory, recognizes the emotions of workers, and provides appropriate feedback. The system's main functions are monitoring materials and documents, extracting business terminology using natural language processing, storing it in a database and standardizing its notation, automatically generating Wiki pages, detecting and notifying inconsistencies in notation, generating annotations in real time during work, and recognizing emotions using an emotion engine.

[0935] System Overview

[0936] 1. Monitoring and detection of materials and documents

[0937] The server monitors the factory's data storage system and automatically detects newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[0938] 2. Extracting business terms using natural language processing models

[0939] The server uses a natural language processing model (NLP model) to analyze materials and documents and extract business terms, a process carried out using automated algorithms.

[0940] 3. Storing terms in the database and standardizing notation

[0941] The extracted business terms are stored in a database by the server. A consistent term list is maintained by checking whether they match existing terms and correcting any inconsistencies in their spellings to make them consistent.

[0942] 4. Automatic generation and updating of knowledge pages

[0943] The server automatically generates and updates knowledge pages based on the business terms stored in the database, adding detailed explanations and related links for each term to make them easily accessible to workers.

[0944] 5. Detecting and notifying inconsistencies

[0945] The server periodically monitors the data storage system and analyzes newly uploaded documents. If it detects any inconsistencies in spelling that do not match the existing terminology list, it notifies the user and prompts them to make corrections.

[0946] 6. Real-time annotation generation while working

[0947] The device receives voice input during work and converts the voice data into text. The server detects business terms from the converted data and generates annotations in real time, which are sent to the device. The device then displays these on the worker's screen to help them understand the work.

[0948] 7. Emotion Recognition with Emotion Engine

[0949] The server analyzes voice and text data during work and uses an emotion engine to recognize the worker's emotions. Based on the recognized emotions, the server provides appropriate feedback and advice.

[0950] Specific examples

[0951] For example, if a factory uploads a document called "New Product Line Plan.pdf", the following process will be performed based on this file:

[0952] 1. The server detects "New Product Line Plan.pdf" and analyzes its contents.

[0953] 2. Use a natural language processing model to extract technical terms such as "TPM (Total Productive Maintenance)."

[0954] 3. Store the extracted business terms in a database, and check for and correct any inconsistencies in notation.

[0955] 4. Automatically generate knowledge pages based on terms and provide related information.

[0956] 5. If there are any inconsistencies in the notation of newly uploaded materials, notify the worker and request corrections.

[0957] 6. When you encounter technical terms while working, their definitions are displayed in real time.

[0958] 7. Provide feedback if workers are stressed.

[0959] Prompt Sentence Examples

[0960] For example, you can input the following prompts into a generative AI model:

[0961] Extract technical terms from the factory's "New Product Line Plan.pdf" and generate knowledge pages based on their definitions. Also, analyze the emotions of workers and display the message "Calm down, everything is under control" when they are feeling stressed.

[0962] The above is a detailed description of the mode for carrying out the invention.

[0963] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0964] Step 1:

[0965] The server monitors the data storage system in the factory to detect newly uploaded materials or documents. When a new document is detected, it obtains its path. This step requires the directory path of the data storage system as input and the path of the detected new document as output.

[0966] Step 2:

[0967] The server uses a natural language processing model (NLP model) to analyze newly detected materials and documents and extract business terms. A list of extracted business terms is generated. This step requires the text data of the document as input and produces a list of business terms as output.

[0968] Step 3:

[0969] The server accesses the database and stores the extracted business terms. It also checks for matches with existing terms and for inconsistencies in spelling. If necessary, it makes corrections to unify spelling. This step requires a list of business terms as input and an updated database as output.

[0970] Step 4:

[0971] The server automatically generates and updates knowledge pages based on the business terms in the database. Specifically, it creates pages that include detailed explanations and related information for each term. This step requires business terms in the database as input, and the generated knowledge pages as output.

[0972] Step 5:

[0973] The server periodically monitors the data storage system and analyzes newly uploaded documents. If it detects a spelling inconsistency that does not match the existing terminology list, it sends a notification to the relevant user to prompt them to correct it. This step requires the new document and the existing terminology list as input, and a spelling inconsistency notification as output.

[0974] Step 6:

[0975] The terminal receives voice input during work and converts the voice data into text. The server then detects business terms from the converted data and generates annotations in real time, which are sent to the terminal. The terminal then displays the annotations on the worker's screen. This step requires voice data as input and produces annotations to be displayed as output.

[0976] Step 7:

[0977] The server analyzes the voice and text data during work and recognizes the worker's emotions using an emotion engine. Based on the recognized emotions, the server provides appropriate feedback and advice. This step requires voice and text data as input, and feedback and advice are obtained as output.

[0978] This will enable the entire system to function consistently, ensuring consistency in business terminology, improving work efficiency, and creating a system that provides support based on the emotions of workers.

[0979] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0980] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0981] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0982] [Third embodiment]

[0983] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0984] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0985] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0986] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0987] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0988] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0989] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0990] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0991] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0992] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0993] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0994] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0995] MODE FOR CARRYING OUT THE INVENTION

[0996] The present invention is a system that uses AI to automatically create a Wiki that summarizes business terminology, eliminating problems such as lack of maintenance and inconsistent spelling. An embodiment of this system is described in detail below.

[0997] System Overview

[0998] The system has the following main functions:

[0999] 1. Monitoring and detection of materials and documents

[1000] 2. Extracting business terms using NLP models

[1001] 3. Storing terms in the database and standardizing notation

[1002] 4. Automatic generation and updating of Wiki pages

[1003] 5. Spelling Variation Detection and Notification

[1004] 6. Real-time annotation generation during meetings

[1005] Monitoring and detection of materials and documents

[1006] The server monitors the file sharing system and automatically detects newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[1007] Extracting business terms using NLP models

[1008] The server uses natural language processing (NLP) models to analyze materials and documents and extract business terms, abbreviations, and technical terms, which then lists specific business terms and allows for efficient analysis.

[1009] Terminology storage in database and standardization of notation

[1010] The server stores the extracted business terms in a database, checks whether they match existing terms, and corrects any inconsistencies in spelling to create a consistent term list.

[1011] Automatic generation and updating of Wiki pages

[1012] The server automatically generates and updates Wiki pages based on business terms stored in the database, adding detailed explanations and related links for each term to make them easily accessible to employees.

[1013] Spelling variation detection and notification

[1014] The server periodically monitors the file sharing system and analyzes newly uploaded documents. If it detects a spelling variation that does not match the existing terminology list, it notifies the user and encourages them to correct the spelling.

[1015] Real-time annotation during meetings

[1016] The device receives voice input during the meeting and converts the voice data into text. The server detects business terms from the converted data and generates annotations in real time, which are sent to the device. The device then displays the annotations on the screen of the meeting participants to support understanding.

[1017] Specific examples

[1018] Consider the process flow when a newly assigned employee downloads "Project Meeting Materials." This material contains specific technical terms and abbreviations.

[1019] 1. Material detection and analysis:

[1020] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[1021] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[1022] 2. Terminology database storage and standardization:

[1023] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[1024] 3. Generate Wiki pages:

[1025] The server automatically generates Wiki pages based on the extracted terms, adding detailed descriptions and related links.

[1026] 4. Notification of spelling variations:

[1027] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[1028] 5. Real-time annotation during meetings:

[1029] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this as "ROI (Return on Investment)" and displays it in real time.

[1030] In this way, the system helps new employees quickly adapt to their work and contributes to standardizing and promoting understanding of work terminology.

[1031] The processing flow will be explained below.

[1032] Program processing

[1033] 1. Monitoring and detection of materials and documents

[1034] Step 1:

[1035] The server accesses the file sharing system and monitors newly uploaded materials and documents.

[1036] Step 2:

[1037] The server scans the directory of the file sharing system at regular time intervals to check whether new files exist.

[1038] Step 3:

[1039] The server detects the newly uploaded file and retrieves the file's metadata (such as name, upload date and time, file type, etc.).

[1040] 2. Extracting business terms using NLP models

[1041] Step 4:

[1042] The server inputs the detected new files into a natural language processing (NLP) model.

[1043] Step 5:

[1044] The server uses NLP models to analyze the content of the files and extract business terms, abbreviations, and technical terms.

[1045] Step 6:

[1046] The server organizes the extracted business terms in list format and temporarily stores the analysis results.

[1047] 3. Storing terms in the database and standardizing notation

[1048] Step 7:

[1049] The server stores the extracted business terms in a database.

[1050] Step 8:

[1051] The server checks the newly extracted terms against existing terms in the database to see if there is a match.

[1052] Step 9:

[1053] If the term is spelled differently, the server will correct it to a unified spelling for consistency.

[1054] 4. Automatic generation and updating of Wiki pages

[1055] Step 10:

[1056] The server runs a script that automatically generates Wiki pages based on business terms in the database.

[1057] Step 11:

[1058] The server adds detailed descriptions and related links to each term to the Wiki page.

[1059] Step 12:

[1060] The server uploads the generated Wiki page to the company's internal Wiki system so that users can view it.

[1061] 5. Spelling Variation Detection and Notification

[1062] Step 13:

[1063] The server periodically monitors the file sharing system and analyzes newly uploaded documents.

[1064] Step 14:

[1065] The server detects spelling variations when a term in the parsed document does not match an existing list of terms.

[1066] Step 15:

[1067] When the server detects a spelling variation, it sends a notification email to the relevant user, prompting them to correct the spelling to the recommended spelling.

[1068] 6. Real-time annotation generation during meetings

[1069] Step 16:

[1070] The terminal receives voice input during the conference and acquires the voice data.

[1071] Step 17:

[1072] The server passes the received voice data to a voice recognition service and converts it into text.

[1073] Step 18:

[1074] The server detects business terms from the text data and generates annotations in real time.

[1075] Step 19:

[1076] The terminal displays the generated annotations on the screens of the conference participants to support their understanding of the terminology.

[1077] Example 1

[1078] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1079] Inconsistent notation and lack of maintenance in materials and documents are often problems in business activities. New employees and those transferred from other departments often have difficulty understanding business terms and abbreviations, which can lead to reduced work efficiency. Also, the use of technical terms during meetings often reduces participants' understanding. It is necessary to solve these problems and provide information quickly and clearly while maintaining consistency in business terminology.

[1080] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1081] In this invention, the server includes a means for monitoring documents and automatically detecting newly uploaded documents, a means for extracting business terms from documents using a natural language processing model, and a means for storing the extracted business terms in a database, checking for matches with existing terms, and standardizing their spellings to a consistent format. This allows for the consolidation and standardization of business terms. The server also includes a means for automatically generating and updating Wiki pages based on business terms and adding term explanations, a means for periodically monitoring the file system, detecting spelling variations in new documents, and sending notifications prompting corrections. The server also includes a means for receiving voice input during meetings, detecting business terms from the voice data, and generating and displaying annotations in real time. This allows for the rapid sharing and understanding of new information, improving work efficiency and accuracy.

[1082] "Materials" refers to information media such as documents, files, and reports used in business activities.

[1083] A "natural language processing model" refers to a model that makes full use of artificial intelligence technology to analyze and understand human language.

[1084] "Business jargon" refers to specialized words and abbreviations used in specific industries or jobs.

[1085] A "database" refers to a system for organizing and storing extracted information and data so that it can be used efficiently.

[1086] "Orthographic variation" refers to differences between words or terms that have the same meaning but are expressed in different forms.

[1087] A "Wiki page" refers to a web page that can be collaboratively edited on the Internet or an intranet.

[1088] A "file system" refers to the method and implementation for managing and storing data and files.

[1089] "Annotation" refers to explanations or explanatory text added to aid understanding.

[1090] "Voice input" refers to a method of inputting voice information into a computer via a microphone or the like.

[1091] MODE FOR CARRYING OUT THE INVENTION

[1092] The present invention is a system that uses AI to automatically create summaries of business terms and resolves issues such as lack of maintenance and inconsistent spelling. Specific embodiments of this system are described in detail below.

[1093] System configuration

[1094] This system is mainly composed of three elements: a server, a terminal, and a user. The server processes and monitors data, the terminal functions as an input and display device, and the user provides and uses information.

[1095] Key Features

[1096] Material monitoring and detection

[1097] The server monitors the file sharing system and automatically detects newly uploaded files. Specifically, the server uses a Python script to check the file system for changes at regular intervals and obtain a list of new or updated files. This operation utilizes the OS's file monitoring function and the cloud service's API.

[1098] Extracting business terms using NLP models

[1099] The server analyzes the documents using a natural language processing (NLP) model to extract business terms. This operation uses a pre-trained NLP model (e.g., BERT or spaCy). Depending on the file format, the server converts the file (e.g., PDF or Word) into text format, and then uses the NLP model to analyze and extract key terms.

[1100] Terminology storage in database and standardization of notation

[1101] The server stores the extracted business terms in a dedicated database. Here, it checks whether they match existing terms and standardizes them into consistent notations. MySQL, PostgreSQL, or other database systems are used. The server executes queries to check and update the contents of the database.

[1102] Automatic generation and updating of Wiki pages

[1103] The server automatically generates and updates Wiki pages based on business terms stored in the database. Wiki pages are generated using Markdown format or HTML templates. Specifically, the server retrieves terminology information from the database and applies it to templates to generate Wiki pages.

[1104] Spelling variation detection and notification

[1105] The server periodically monitors the file sharing system and compares new material with the existing database. If a spelling variation is detected, the server notifies the user via email or instant messaging services (e.g., Slack).

[1106] Real-time annotation during meetings

[1107] The device receives voice input during the meeting and converts the voice data into text. The Google Speech-to-Text API is used for speech recognition. For example, if someone says "ROI" during a meeting, the server analyzes it in real time and generates the annotation "ROI (Return on Investment)." The generated annotation is then displayed on the device of each meeting participant.

[1108] Specific examples

[1109] The following shows the process flow when a newly assigned employee downloads "Project Meeting Materials."

[1110] 1. Material detection and analysis:

[1111] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[1112] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[1113] 2. Terminology database storage and standardization:

[1114] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[1115] 3. Generate Wiki pages:

[1116] The server automatically generates Wiki pages based on the extracted terms, adding detailed descriptions and related links.

[1117] 4. Notification of spelling variations:

[1118] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[1119] 5. Real-time annotation during meetings:

[1120] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this as "ROI (Return on Investment)" and displays it in real time.

[1121] Prompt Sentence Examples

[1122] 1. Material detection and analysis:

[1123] "Detect newly uploaded conference materials and extract the technical terms they contain."

[1124] 2. Terminology database storage and standardization:

[1125] "Store the extracted terms in a database and check if they match existing terms."

[1126] 3. Generate Wiki pages:

[1127] "Generate wiki pages based on terms stored in your database and add detailed descriptions and related links."

[1128] 4. Notification of spelling variations:

[1129] "If you detect spelling variations in newly uploaded materials, please notify us of the recommended spelling."

[1130] 5. Real-time annotation during meetings:

[1131] "Extract technical terms from audio during meetings in real time and generate and display annotations."

[1132] The above describes a specific embodiment of the present invention. This system makes it possible to standardize business terminology and promote understanding, thereby realizing efficient business operations.

[1133] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1134] Step 1: Monitoring and detecting material

[1135] The server monitors the file sharing system. Specifically, the server scans the contents of the file system at regular intervals to detect newly added or updated files. The input is the file system path and the monitoring interval. The output is a list of new or updated files. The server also records file metadata (file name, creation date and time, and update date and time).

[1136] Step 2: Prepare the file for analysis

[1137] The server obtains the paths of the detected files and prepares them for analysis. First, it checks the file format (PDF, Word, etc.) and converts it to text data using an appropriate text conversion tool (e.g., PyPDF2 library for PDF, python-docx library for Word). The input is a list of detected files, and the output is text data.

[1138] Step 3: Extracting business terms using an NLP model

[1139] The server inputs the converted text data into an NLP model to extract business terms, abbreviations, and technical terms. This process uses a pre-trained NLP model (e.g., BERT or spaCy). The input is the text data, and the output is a list of extracted business terms. Specifically, terms such as "ROI" and "critical path" are extracted.

[1140] Step 4: Store business terms in a database

[1141] The server stores the extracted business terms in a database. During this process, it checks for matches with existing terms and standardizes the notation. The database used is MySQL or PostgreSQL. The input is the extracted business term list, and the output is the updated database. For example, if "ROI" and "return on investment" match, they are unified to "ROI."

[1142] Step 5: Automatically generate and update Wiki pages

[1143] The server automatically generates and updates Wiki pages based on business terms stored in the database. Markdown format and HTML templates are used for this process. The input is the business term information in the database, and the output is the generated or updated Wiki page. For example, the "ROI" page contains the description "Return on Investment."

[1144] Step 6: Detecting and reporting spelling variations

[1145] The server compares newly uploaded materials with the existing database to detect variations in notation. If any discrepancies are found during this process, a notification is sent to the user. Notification methods include email and instant messaging services (e.g., Slack). The input is the new material and database information, and the output is a notification to the user. For example, a notification recommending that "return on investment" be changed to "ROI" is sent.

[1146] Step 7: Real-time annotation during the meeting

[1147] The device receives voice input during the meeting and converts the voice data into text. This process uses the Google Speech-to-Text API. The input is the voice data from the meeting, and the output is text data. The converted text data is sent to the server, and if a business term is detected, an annotation is generated in real time and displayed on the device. For example, if "ROI" is spoken during a meeting, the annotation "ROI (Return on Investment)" is generated and displayed.

[1148] (Application example 1)

[1149] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1150] In today's business environment, multiple documents and materials are frequently uploaded, often containing many technical terms and abbreviations. This makes it difficult for newly assigned employees and external participants to quickly adapt to the work. Furthermore, inconsistencies in the spelling of business terms occur, reducing the consistency and efficiency of information. There is a need to solve these problems and improve business efficiency.

[1151] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1152] In this invention, the server includes: means for monitoring materials and documents and detecting newly uploaded materials and documents; means for extracting business terms from materials and documents using a natural language processing model; means for storing the extracted business terms in a database, checking for matches with existing terms, and standardizing spellings; means for automatically generating and updating information pages based on the business terms and adding term explanations; means for periodically monitoring the file sharing system, detecting spelling variations in new documents, and sending notifications prompting corrections; means for receiving voice input during meetings, detecting business terms from the voice data, and generating and displaying annotations in real time; means for analyzing the camera feed of a visual device and extracting text from documents; means for displaying business terms as annotations from the extracted text; and means for analyzing the voice input of the visual device, performing speech recognition, and providing detailed explanations of the relevant business terms by voice. This enables newly assigned employees to quickly adapt to their work, promotes standardization of business terminology, and promotes understanding.

[1153] "Materials" refers to documents and other information related to business.

[1154] "Document" refers to a medium containing information stored as printed material or digital data.

[1155] "Monitoring" means that a system periodically or continuously checks a particular file sharing system or database to detect newly added information.

[1156] A "natural language processing model" refers to a computer program or algorithm that analyzes human language and extracts useful information from text data.

[1157] "Business jargon" refers to technical terms and abbreviations frequently used in a particular business or industry.

[1158] A "database" refers to a system for efficiently storing, retrieving, and managing information.

[1159] "Unification of spelling" refers to the standardization of words with the same meaning that are expressed in different spellings or writing styles into a consistent format.

[1160] An "information page" is a web page or document that provides detailed explanations and related information about a particular term or concept.

[1161] A "file sharing system" refers to a system that allows multiple users within a specific network to share and access files.

[1162] "Spelling variation" refers to the situation where the same term or concept is written in different ways.

[1163] "Notification" refers to a system-generated message that informs the user of a particular event or situation.

[1164] "Voice input during a meeting" refers to the process by which voice data spoken during a meeting is received by the system.

[1165] "Audio Data" refers to a digital recording of sound collected by a microphone or the like.

[1166] "Generating and displaying annotations in real time" refers to the process of instantly analyzing voice or text input and displaying additional information based on the results.

[1167] "Visual devices" refer to devices worn by users that visually present information, such as smart glasses and head-mounted displays.

[1168] "Camera feed" refers to video data captured by a camera in real time.

[1169] "Extracting text from a document" refers to the process of analyzing character information from camera images or scanned data and extracting that text data.

[1170] "Speech recognition" refers to the technology and algorithms that analyze audio signals and identify meaningful words and sentences from them.

[1171] This invention is a system that uses AI to consolidate business terminology, particularly to improve efficiency in factory work environments. This system is composed of a server, terminals, and users. The specific configuration and functions of this system are described below.

[1172] System configuration and functions

[1173] 1. Monitoring and detection of materials and documents

[1174] The server monitors the factory's file sharing system and database, automatically detecting newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[1175] 2. Extracting business terms using natural language processing models

[1176] The server uses natural language processing (NLP) models to analyze documents and extract business terms, abbreviations, and technical terms, using NLP libraries such as SpaCy.

[1177] 3. Storing terms in the database and standardizing notation

[1178] The server stores the extracted business terms in a database and checks whether they match existing terms. If there are any variations in spelling, they are unified to create a consistent term list. For this purpose, an SQL-based database management system (e.g., MySQL) is used.

[1179] 4. Automatic generation and updating of information pages

[1180] The server automatically generates and updates information pages based on business terms stored in the database, including detailed explanations and related links for each term.

[1181] 5. Spelling Variation Detection and Notification

[1182] The server periodically monitors the file sharing system and analyzes newly uploaded documents. If it detects a spelling variation that does not match the existing term list, it sends a notification to the relevant user urging them to correct the variation. The notification system includes email notifications and push notifications.

[1183] 6. Real-time voice annotation

[1184] The device receives voice input during the meeting and converts the voice data into text. The server detects business terms from the text and generates annotations in real time and sends them to the device, allowing meeting participants to instantly understand the meaning of the terms.

[1185] 7. Visual Device Camera Feed Analysis

[1186] A visual device (e.g., smart glasses) extracts text from a document through a camera feed. After extracting the text using Tesseract OCR, the server analyzes the text and displays real-time annotations for business terms.

[1187] 8. Analysis and presentation of voice input

[1188] The voice input from the visual device is analyzed and speech recognition is performed using the Google Cloud Speech-to-Text API. For recognized terms, the server retrieves detailed descriptions and presents them as audio through the visual device's speaker.

[1189] Specific examples

[1190] For example, suppose a newly assigned employee views a document containing the term "ROI (Return on Investment)." The camera in the vision device captures this, and the server performs text analysis, displaying the annotation "Return on Investment" for the term "ROI." Furthermore, if the employee voice-inputs, "Tell me more about Return on Investment," the vision device recognizes the voice and provides a detailed explanation.

[1191] This will generate a prompt like this:

[1192] "A new document has been uploaded. This document contains the term 'return on investment'. Could you please explain this term in more detail?"

[1193] In this way, the present invention promotes understanding of business terms in a factory work environment, enabling efficient work execution.

[1194] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1195] Step 1:

[1196] The server monitors the file sharing system within the factory and detects newly uploaded materials and documents. As input, it receives directory information and file lists from the file sharing system and identifies new files based on timestamps. After processing this data, it outputs a list of newly uploaded files.

[1197] Step 2:

[1198] The server uses a natural language processing model (e.g., SpaCy) to extract business terms from newly detected materials and documents. As input, it receives the contents of the files detected in step 1 in text format and analyzes the text data. This data processing results in the output of a list of extracted business terms.

[1199] Step 3:

[1200] The server stores the extracted business terms in a database, checks for matches with existing terms, and standardizes notation. As input, it receives the list of business terms extracted in step 2 and compares them with existing terms in the database. This data processing results in a consistent list of terms being output.

[1201] Step 4:

[1202] The server automatically generates and updates information pages based on business terms stored in the database. It receives an updated term list as input and adds detailed explanations and related links for each term. This data calculation results in the output of the latest information page.

[1203] Step 5:

[1204] The server periodically monitors the file sharing system, detects spelling variations that do not match the existing terminology list, and sends a notification prompting correction. It receives the content of newly uploaded documents as input and compares it with the existing terminology list. After processing this data, it outputs a notification containing mismatched terms and suggested corrections.

[1205] Step 6:

[1206] The device converts voice input received during the meeting into text and sends it to the server. It receives the voice data of the meeting participants as input and performs speech recognition using the Google Cloud Speech-to-Text API. This data is then processed and the text data is output.

[1207] Step 7:

[1208] The server generates annotations in real time for business terms detected from the voice input and sends them to the terminal. As input, it receives the text data output in step 6, analyzes and extracts business terms, and generates annotations for them. This data processing results in the output of annotated text.

[1209] Step 8:

[1210] The vision device analyzes the camera feed and extracts text from the document. It receives the video data from the camera feed as input and performs character recognition using Tesseract OCR. It then processes this data and outputs the extracted text data.

[1211] Step 9:

[1212] The server generates annotations of business terms from the extracted text and displays them on a visual device. It receives the text data output in step 8 as input, analyzes and extracts business terms, and generates annotations for them. This data processing results in the output of annotated text.

[1213] Step 10:

[1214] The visual device analyzes voice input and performs speech recognition. It receives the user's voice data as input and performs speech recognition using the Google Cloud Speech-to-Text API. It then processes this data and outputs the recognized text data.

[1215] Step 11:

[1216] The server obtains detailed explanations for the recognized terms and presents them as audio through the speaker of the visual device. As input, it receives the text data output in step 10, generates detailed explanations, and sends them as audio data to the visual device. This data processing results in the audio data of the detailed explanation being output.

[1217] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1218] MODE FOR CARRYING OUT THE INVENTION

[1219] The present invention is a system that uses AI to automatically create a wiki that summarizes business terminology, resolves issues such as lack of maintenance and inconsistent spelling, and combines it with an emotion engine that recognizes user emotions. An embodiment of this system is described in detail below.

[1220] System Overview

[1221] The system has the following main functions:

[1222] 1. Monitoring and detection of materials and documents

[1223] 2. Extracting business terms using NLP models

[1224] 3. Storing terms in the database and standardizing notation

[1225] 4. Automatic generation and updating of Wiki pages

[1226] 5. Spelling Variation Detection and Notification

[1227] 6. Real-time annotation generation during meetings

[1228] 7. Emotion engine that recognizes user emotions

[1229] Monitoring and detection of materials and documents

[1230] The server monitors the file sharing system and automatically detects newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[1231] Extracting business terms using NLP models

[1232] The server uses natural language processing (NLP) models to analyze materials and documents and extract business terms, abbreviations, and technical terms, which then lists specific business terms and allows for efficient analysis.

[1233] Terminology storage in database and standardization of notation

[1234] The server stores the extracted business terms in a database, checks whether they match existing terms, and corrects any inconsistencies in spelling to create a consistent term list.

[1235] Automatic generation and updating of Wiki pages

[1236] The server automatically generates and updates Wiki pages based on business terms stored in the database, adding detailed explanations and related links for each term to make them easily accessible to employees.

[1237] Spelling variation detection and notification

[1238] The server periodically monitors the file sharing system and analyzes newly uploaded documents. If it detects a spelling variation that does not match the existing terminology list, it notifies the user and encourages them to correct the spelling.

[1239] Real-time annotation during meetings

[1240] The device receives voice input during the meeting and converts the voice data into text. The server detects business terms from the converted data and generates annotations in real time, which are sent to the device. The device then displays the annotations on the screen of the meeting participants to support understanding.

[1241] Emotion engine that recognizes user emotions

[1242] The server analyzes the user's voice and text data, recognizes the user's emotions using an emotion engine, and provides appropriate feedback and advice based on the recognized emotions.

[1243] Specific examples

[1244] Consider the process flow when a newly assigned employee downloads "Project Meeting Materials." This material contains specific technical terms and abbreviations.

[1245] 1. Material detection and analysis:

[1246] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[1247] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[1248] 2. Terminology database storage and standardization:

[1249] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[1250] 3. Generate Wiki pages:

[1251] The server automatically generates Wiki pages based on the extracted terms, adding detailed descriptions and related links.

[1252] 4. Notification of spelling variations:

[1253] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[1254] 5. Real-time annotation during meetings:

[1255] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this as "ROI (Return on Investment)" and displays it in real time.

[1256] 6. Leveraging the Emotion Engine:

[1257] The server analyzes the voice data and text data entered by the user during the meeting and uses an emotion engine to recognize the user's emotions. For example, if the user is feeling stressed during the meeting, the server will provide appropriate feedback to support the progress of the meeting.

[1258] The processing flow will be explained below.

[1259] MODE FOR CARRYING OUT THE INVENTION

[1260] Monitoring and detection of materials and documents

[1261] Step 1:

[1262] The server accesses the file sharing system and monitors newly uploaded materials and documents.

[1263] Step 2:

[1264] The server scans the directory of the file sharing system at regular time intervals to check whether new files exist.

[1265] Step 3:

[1266] The server detects the newly uploaded file and retrieves the file's metadata (such as name, upload date and time, file type, etc.).

[1267] Extracting business terms using NLP models

[1268] Step 4:

[1269] The server inputs the detected new files into a natural language processing (NLP) model.

[1270] Step 5:

[1271] The server uses NLP models to analyze the content of the files and extract business terms, abbreviations, and technical terms.

[1272] Step 6:

[1273] The server organizes the extracted business terms in list format and temporarily stores the analysis results.

[1274] Terminology storage in database and standardization of notation

[1275] Step 7:

[1276] The server stores the extracted business terms in a database.

[1277] Step 8:

[1278] The server checks the newly extracted terms against existing terms in the database to see if there is a match.

[1279] Step 9:

[1280] If the term is spelled differently, the server will correct it to a unified spelling for consistency.

[1281] Automatic generation and updating of Wiki pages

[1282] Step 10:

[1283] The server runs a script that automatically generates Wiki pages based on business terms in the database.

[1284] Step 11:

[1285] The server adds detailed descriptions and related links to each term to the Wiki page.

[1286] Step 12:

[1287] The server uploads the generated Wiki page to the company's internal Wiki system so that users can view it.

[1288] Spelling variation detection and notification

[1289] Step 13:

[1290] The server periodically monitors the file sharing system and analyzes newly uploaded documents.

[1291] Step 14:

[1292] The server detects spelling variations when a term in the parsed document does not match an existing list of terms.

[1293] Step 15:

[1294] When the server detects a spelling variation, it sends a notification email to the relevant user, prompting them to correct the spelling to the recommended spelling.

[1295] Real-time annotation during meetings

[1296] Step 16:

[1297] The terminal receives voice input during the conference and acquires the voice data.

[1298] Step 17:

[1299] The server passes the received voice data to a voice recognition service and converts it into text.

[1300] Step 18:

[1301] The server detects business terms from the text data and generates annotations in real time.

[1302] Step 19:

[1303] The terminal displays the generated annotations on the screens of the conference participants to support their understanding of the terminology.

[1304] Emotion engine that recognizes user emotions

[1305] Step 20:

[1306] The terminal collects voices during the conference and text data entered by the user.

[1307] Step 21:

[1308] The server passes the collected voice and text data to an emotion engine to analyze the user's emotions.

[1309] Step 22:

[1310] The server generates appropriate feedback and advice based on the recognized emotions.

[1311] Step 23:

[1312] The terminal displays the generated feedback and advice on the screens of the conference participants to support the progress of the conference.

[1313] Example 2

[1314] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1315] When managing materials and documents, it is necessary to efficiently monitor and detect newly uploaded information. There is also the issue of eliminating inconsistencies in notation due to the existence of unstandardized business terms and abbreviations, and creating and maintaining a consistent terminology list. Furthermore, there is a need to streamline the provision of information during meetings and to implement a feedback system based on user sentiment.

[1316] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1317] In this invention, the server includes: means for monitoring materials and documents and detecting newly uploaded materials and documents; means for extracting business terms from materials and documents using a natural language processing model; means for storing the extracted business terms in a database, checking for matches with existing terms, and standardizing notation; means for automatically generating and updating information pages based on the business terms and adding explanations of the terms; means for periodically monitoring the information sharing system, detecting notation variations in new documents, and sending notifications prompting corrections; means for receiving voice input during a meeting, detecting business terms from the voice data, and generating and displaying annotations in real time; and means for analyzing user voice and text data, recognizing user emotions, and providing feedback and advice. This enables efficient document management, standardization of notation, maintaining consistency, efficient information provision during meetings, and appropriate feedback based on user emotions.

[1318] "Materials and documents" refers to electronic files containing text data and graphical data, such as reports, plans, technical manuals, procedure manuals, e-mails, memos, and meeting materials used inside and outside a company.

[1319] "Monitoring" means constantly checking operations such as adding, changing, or deleting files that are performed on a specific system or network, and detecting them when certain conditions are met.

[1320] A "natural language processing model" refers to an algorithm or machine learning model designed to understand, analyze, and generate human language, specifically extracting business terms and analyzing text.

[1321] "Business jargon" refers to the technical terms, abbreviations, and phrases used within a particular company or industry that are an important part of business processes and communication.

[1322] "Database" refers to a data storage system that stores data in an organized manner and is designed to make it easy to access and manage.

[1323] "Unification of notation" refers to the process of changing identical terms with different notations or expressions into a consistent format to maintain data integrity and consistency.

[1324] An "information page" is a web page or digital document that aggregates business terms and related information and displays them in an easy-to-reference format.

[1325] An "information sharing system" is a platform that allows multiple users to upload, download, and share files and information, and includes cloud storage and corporate network drives.

[1326] "Spelling variation" refers to the use of different expressions or spellings that have the same meaning, which creates problems that make understanding and management complicated.

[1327] "Sending a notification" refers to the act of communicating alerts or information to a user based on a specific event or condition detected by the system.

[1328] "Receiving audio input" refers to the process of collecting audio data through a microphone or other audio capture device and passing that data to the system.

[1329] "Audio data" is data that is a digital representation of human speech recorded via a voice input device.

[1330] "Generating annotations in real time" refers to the process of instantly adding relevant explanations and information to input speech or text data and displaying them.

[1331] "Recognizing user emotions" refers to the technical process of determining the emotions a user is feeling at that time through analysis of voice and text data.

[1332] "Providing feedback and advice" refers to the process by which the system suggests appropriate responses or recommended actions based on the user's perceived emotions.

[1333] MODE FOR CARRYING OUT THE INVENTION

[1334] The present invention is a system that uses AI to automatically create information pages summarizing business terms, eliminating problems such as lack of maintenance and inconsistent spelling, and combining this with an emotion engine that recognizes user emotions. An embodiment of this system is described in detail below.

[1335] Hardware and software used

[1336] The system utilizes the following hardware and software:

[1337] Hardware: Server, terminal, audio input device (microphone)

[1338] File sharing systems: Dropbox, Google Drive, etc.

[1339] Natural language processing models: Google BERT, OpenAI GPT-3, etc.

[1340] Database: MySQL, PostgreSQL

[1341] Emotion engine: IBM Watson, Azure Cognitive Services

[1342] Speech recognition software: Google Speech-to-Text, Microsoft Azure Speech Service

[1343] Overall flow

[1344] The server monitors the file sharing system to detect and analyze newly uploaded materials and documents. A natural language processing model is used for the analysis, which extracts business terms, abbreviations, and technical terms from the materials. Before storing the extracted terms in the database, they are compared with existing terms, and any variations in spelling are unified.

[1345] Next, the server automatically generates and updates information pages based on the business terms stored in the database. The generated information pages include detailed explanations of the terms and related links. The server also periodically monitors the information sharing system to detect spelling variations in new documents. If any are detected, a notification is sent to the relevant user.

[1346] During a meeting, the device receives voice input and converts the voice data into text. The text data is then sent to a server, which detects business terms and generates annotations in real time. The annotations are then sent to the device and displayed on the screens of meeting participants.

[1347] Finally, the server analyzes the user's voice and text data, and uses an emotion engine to recognize the user's emotions. Based on the recognized emotions, the server provides appropriate feedback and advice.

[1348] Specific examples

[1349] The process flow when a newly assigned employee downloads "Project Meeting Materials" is as follows: This material contains specific technical terms and abbreviations.

[1350] 1. Material detection and analysis:

[1351] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[1352] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[1353] 2. Terminology database storage and standardization:

[1354] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[1355] 3. Generate information page:

[1356] The server automatically generates information pages based on the extracted terms, adding detailed descriptions and related links.

[1357] 4. Notification of spelling variations:

[1358] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[1359] 5. Real-time annotation during meetings:

[1360] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this in real time as "ROI (Return on Investment)" and displays it.

[1361] 6. Leveraging the Emotion Engine:

[1362] The server analyzes the voice data and text data entered by the user during the meeting and uses an emotion engine to recognize the user's emotions. For example, if the user is feeling stressed during the meeting, the server will provide appropriate feedback to support the progress of the meeting.

[1363] Prompt Sentence Examples

[1364] Here are some example prompts to input to a generative AI model:

[1365] Please analyze the business terms contained in newly uploaded documents and standardize the notation of the extracted business terms to a consistent format.

[1366] "Analyze audio during meetings in real time and annotate and display important business terms."

[1367] "Recognize emotions based on the user's voice and text data and provide appropriate feedback if they are feeling stressed."

[1368] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1369] Step 1:

[1370] The server monitors the file sharing system to detect newly uploaded materials or documents. In this step, the server checks whether a new file has been uploaded from the file sharing system (e.g., Dropbox, Google Drive) and obtains the path and name of the file. Specifically, the server scans the file list and detects the new file "Project Meeting Materials.pdf."

[1371] Input: A file uploaded to a file sharing system.

[1372] Output: The path and name of the file found.

[1373] Step 2:

[1374] The server analyzes the detected file using a natural language processing (NLP) model to extract business terms. In this step, the server uses an NLP model (e.g., Google BERT, OpenAI GPT-3) to list business terms such as "ROI" and "critical path" from the text data in the file. Specifically, the server analyzes "Project Meeting Materials.pdf" and lists business terms.

[1375] Input: The path and name of the detected file.

[1376] Output: A list of extracted business terms.

[1377] Step 3:

[1378] The server stores the extracted business terms in a database, checks whether they match existing terms, and unifies the spelling. In this step, before storing the extracted terms in a database (e.g., MySQL, PostgreSQL), the server compares them with existing terms and unifies any spelling variations. Specifically, it checks whether "ROI" already exists in the database and whether it matches "Return on Investment," and unifies them.

[1379] Input: A list of extracted business terms.

[1380] Output: A consistent list of terms stored in a database.

[1381] Step 4:

[1382] The server automatically generates and updates information pages based on the business terms stored in the database. In this step, the server generates and updates information pages containing detailed explanations and related links based on the newly extracted business terms. Specifically, a detailed explanation of "ROI (Return on Investment)" is added to the information page.

[1383] Input: A consistent list of terms stored in a database.

[1384] Output: Generated and updated information page.

[1385] Step 5:

[1386] The server periodically monitors the information sharing system, detects spelling variations in new documents, and sends a notification to prompt correction. In this step, the server compares the newly detected document with the existing term list, and if a spelling variation is found, it sends a notification to the relevant user. Specifically, if "return on investment" is newly detected, it sends a notification to unify it to "ROI."

[1387] Input: The newly discovered document.

[1388] Output: Notification of spelling variations.

[1389] Step 6:

[1390] The device receives voice input during the meeting and converts the voice data into text. In this step, the device uses speech recognition software (e.g., Google Speech-to-Text, Microsoft Azure Speech Service) to convert the voice during the meeting into text in real time. Specifically, the device converts the part where "ROI" is spoken during the meeting into text.

[1391] Input: Audio input during a meeting.

[1392] Output: The audio data converted to text.

[1393] Step 7:

[1394] The server detects business terms from the converted speech data, generates annotations in real time, and sends them to the terminal. In this step, the server detects "ROI" from the converted data and generates the annotation "ROI (Return on Investment)." After generating the annotation, it sends it to the terminal and displays it. Specifically, the annotation "ROI (Return on Investment)" is displayed on the screen of the conference participants.

[1395] Input: Transcribed audio data.

[1396] Output: The generated annotations.

[1397] Step 8:

[1398] The server analyzes the user's voice and text data and recognizes the user's emotions using an emotion engine. In this step, the server analyzes the user's emotions using an emotion engine (e.g., IBM Watson, Azure Cognitive Services). Based on the recognized emotions, the server provides appropriate feedback and advice. Specifically, if the server determines that the user is feeling stressed, it sends feedback encouraging the user to relax.

[1399] Input: User voice and text data.

[1400] Output: Emotion-based feedback and advice.

[1401] (Application example 2)

[1402] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1403] Modern factories generate a large number of documents and materials every day, and these documents contain a large amount of technical terminology. However, if these terms are not used consistently, it becomes difficult to understand and share information. In addition, there is a lack of a way to recognize workers' emotions and provide appropriate feedback, which can lead to reduced work efficiency and the risk of mistakes. There is a need for a system that can resolve these issues and provide support based on workers' emotions while maintaining consistency in work terminology.

[1404] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for monitoring materials and documents and detecting newly uploaded materials and documents; means for extracting business terms from materials and documents using a natural language processing model; means for storing the extracted business terms in a database, checking for matches with existing terms, and unifying notations; means for automatically generating and updating knowledge pages based on the business terms and adding term explanations; means for periodically monitoring the data storage system, detecting notation inconsistencies in new documents, and sending notifications prompting corrections; means for receiving voice input during work, detecting business terms from the voice data, and generating and displaying annotations in real time; and means for recognizing the user's emotions and providing appropriate instructions and feedback. This makes it possible to improve work efficiency while ensuring consistency in business terminology.

[1405] "Materials and documents" refers to all documents generated within the factory, such as reports, plans, procedures, and manuals.

[1406] "Means for detecting newly uploaded files" refers to the ability to monitor the data storage system and automatically detect newly uploaded files.

[1407] A "natural language processing model" is a software algorithm for analyzing the content of a document and extracting specific business terms.

[1408] "Business terms" refers to technical terms, abbreviations, and keywords used in a particular business or field of expertise.

[1409] "Means of storing in a database, checking for matches with existing terms, and standardizing notation" refers to a function that saves extracted business terms and compares them with existing data to ensure consistency in notation.

[1410] "Means for generating and updating knowledge pages and adding explanations of terms" is a function that automatically creates and updates information pages (Wiki pages) that include detailed explanations of business terms.

[1411] "Means for periodically monitoring the data storage system, detecting inconsistencies in notation in new documents, and sending notifications prompting corrections" refers to a notification function that monitors new document uploads, detects inconsistencies in notation, and requests the user to make corrections.

[1412] "Means for receiving voice input during work, detecting business terms from the voice data, and generating and displaying annotations in real time" refers to a function that analyzes voice information generated during work, recognizes technical terms, and immediately displays their definitions and explanations.

[1413] "Means for recognizing the user's emotions and providing appropriate instructions and feedback" refers to a function that detects the user's emotions from their words and actions, and provides advice and support according to the situation.

[1414] MODE FOR CARRYING OUT THE INVENTION

[1415] This invention is a system that maintains consistency in business terminology within a factory, recognizes the emotions of workers, and provides appropriate feedback. The system's main functions are monitoring materials and documents, extracting business terminology using natural language processing, storing it in a database and standardizing its notation, automatically generating Wiki pages, detecting and notifying inconsistencies in notation, generating annotations in real time during work, and recognizing emotions using an emotion engine.

[1416] System Overview

[1417] 1. Monitoring and detection of materials and documents

[1418] The server monitors the factory's data storage system and automatically detects newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[1419] 2. Extracting business terms using natural language processing models

[1420] The server uses a natural language processing model (NLP model) to analyze materials and documents and extract business terms, a process carried out using automated algorithms.

[1421] 3. Storing terms in the database and standardizing notation

[1422] The extracted business terms are stored in a database by the server. A consistent term list is maintained by checking whether they match existing terms and correcting any inconsistencies in their spellings to make them consistent.

[1423] 4. Automatic generation and updating of knowledge pages

[1424] The server automatically generates and updates knowledge pages based on the business terms stored in the database, adding detailed explanations and related links for each term to make them easily accessible to workers.

[1425] 5. Detecting and notifying inconsistencies

[1426] The server periodically monitors the data storage system and analyzes newly uploaded documents. If it detects any inconsistencies in spelling that do not match the existing terminology list, it notifies the user and prompts them to make corrections.

[1427] 6. Real-time annotation generation while working

[1428] The device receives voice input during work and converts the voice data into text. The server detects business terms from the converted data and generates annotations in real time, which are sent to the device. The device then displays these on the worker's screen to help them understand the work.

[1429] 7. Emotion Recognition with Emotion Engine

[1430] The server analyzes voice and text data during work and uses an emotion engine to recognize the worker's emotions. Based on the recognized emotions, the server provides appropriate feedback and advice.

[1431] Specific examples

[1432] For example, if a factory uploads a document called "New Product Line Plan.pdf", the following process will be performed based on this file:

[1433] 1. The server detects "New Product Line Plan.pdf" and analyzes its contents.

[1434] 2. Use a natural language processing model to extract technical terms such as "TPM (Total Productive Maintenance)."

[1435] 3. Store the extracted business terms in a database, and check for and correct any inconsistencies in notation.

[1436] 4. Automatically generate knowledge pages based on terms and provide related information.

[1437] 5. If there are any inconsistencies in the notation of newly uploaded materials, notify the worker and request corrections.

[1438] 6. When you encounter technical terms while working, their definitions are displayed in real time.

[1439] 7. Provide feedback if workers are stressed.

[1440] Prompt Sentence Examples

[1441] For example, you can input the following prompts into a generative AI model:

[1442] Extract technical terms from the factory's "New Product Line Plan.pdf" and generate knowledge pages based on their definitions. Also, analyze the emotions of workers and display the message "Calm down, everything is under control" when they are feeling stressed.

[1443] The above is a detailed description of the mode for carrying out the invention.

[1444] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1445] Step 1:

[1446] The server monitors the data storage system in the factory to detect newly uploaded materials or documents. When a new document is detected, it obtains its path. This step requires the directory path of the data storage system as input and the path of the detected new document as output.

[1447] Step 2:

[1448] The server uses a natural language processing model (NLP model) to analyze newly detected materials and documents and extract business terms. A list of extracted business terms is generated. This step requires the text data of the document as input and produces a list of business terms as output.

[1449] Step 3:

[1450] The server accesses the database and stores the extracted business terms. It also checks for matches with existing terms and for inconsistencies in spelling. If necessary, it makes corrections to unify spelling. This step requires a list of business terms as input and an updated database as output.

[1451] Step 4:

[1452] The server automatically generates and updates knowledge pages based on the business terms in the database. Specifically, it creates pages that include detailed explanations and related information for each term. This step requires business terms in the database as input, and the generated knowledge pages as output.

[1453] Step 5:

[1454] The server periodically monitors the data storage system and analyzes newly uploaded documents. If it detects a spelling inconsistency that does not match the existing terminology list, it sends a notification to the relevant user to prompt them to correct it. This step requires the new document and the existing terminology list as input, and a spelling inconsistency notification as output.

[1455] Step 6:

[1456] The terminal receives voice input during work and converts the voice data into text. The server then detects business terms from the converted data and generates annotations in real time, which are sent to the terminal. The terminal then displays the annotations on the worker's screen. This step requires voice data as input and produces annotations to be displayed as output.

[1457] Step 7:

[1458] The server analyzes the voice and text data during work and recognizes the worker's emotions using an emotion engine. Based on the recognized emotions, the server provides appropriate feedback and advice. This step requires voice and text data as input, and feedback and advice are obtained as output.

[1459] This will enable the entire system to function consistently, ensuring consistency in business terminology, improving work efficiency, and creating a system that provides support based on the emotions of workers.

[1460] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1461] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1462] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1463] [Fourth embodiment]

[1464] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1465] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1466] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1467] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1468] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1469] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1470] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1471] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1472] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1473] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1474] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1475] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1476] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1477] MODE FOR CARRYING OUT THE INVENTION

[1478] The present invention is a system that uses AI to automatically create a Wiki that summarizes business terminology, eliminating problems such as lack of maintenance and inconsistent spelling. An embodiment of this system is described in detail below.

[1479] System Overview

[1480] The system has the following main functions:

[1481] 1. Monitoring and detection of materials and documents

[1482] 2. Extracting business terms using NLP models

[1483] 3. Storing terms in the database and standardizing notation

[1484] 4. Automatic generation and updating of Wiki pages

[1485] 5. Spelling Variation Detection and Notification

[1486] 6. Real-time annotation generation during meetings

[1487] Monitoring and detection of materials and documents

[1488] The server monitors the file sharing system and automatically detects newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[1489] Extracting business terms using NLP models

[1490] The server uses natural language processing (NLP) models to analyze materials and documents and extract business terms, abbreviations, and technical terms, which then lists specific business terms and allows for efficient analysis.

[1491] Terminology storage in database and standardization of notation

[1492] The server stores the extracted business terms in a database, checks whether they match existing terms, and corrects any inconsistencies in spelling to create a consistent term list.

[1493] Automatic generation and updating of Wiki pages

[1494] The server automatically generates and updates Wiki pages based on business terms stored in the database, adding detailed explanations and related links for each term to make them easily accessible to employees.

[1495] Spelling variation detection and notification

[1496] The server periodically monitors the file sharing system and analyzes newly uploaded documents. If it detects a spelling variation that does not match the existing terminology list, it notifies the user and encourages them to correct the spelling.

[1497] Real-time annotation during meetings

[1498] The device receives voice input during the meeting and converts the voice data into text. The server detects business terms from the converted data and generates annotations in real time, which are sent to the device. The device then displays the annotations on the screen of the meeting participants to support understanding.

[1499] Specific examples

[1500] Consider the process flow when a newly assigned employee downloads "Project Meeting Materials." This material contains specific technical terms and abbreviations.

[1501] 1. Material detection and analysis:

[1502] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[1503] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[1504] 2. Terminology database storage and standardization:

[1505] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[1506] 3. Generate Wiki pages:

[1507] The server automatically generates Wiki pages based on the extracted terms, adding detailed descriptions and related links.

[1508] 4. Notification of spelling variations:

[1509] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[1510] 5. Real-time annotation during meetings:

[1511] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this as "ROI (Return on Investment)" and displays it in real time.

[1512] In this way, the system helps new employees quickly adapt to their work and contributes to standardizing and promoting understanding of work terminology.

[1513] The processing flow will be explained below.

[1514] Program processing

[1515] 1. Monitoring and detection of materials and documents

[1516] Step 1:

[1517] The server accesses the file sharing system and monitors newly uploaded materials and documents.

[1518] Step 2:

[1519] The server scans the directory of the file sharing system at regular time intervals to check whether new files exist.

[1520] Step 3:

[1521] The server detects the newly uploaded file and retrieves the file's metadata (such as name, upload date and time, file type, etc.).

[1522] 2. Extracting business terms using NLP models

[1523] Step 4:

[1524] The server inputs the detected new files into a natural language processing (NLP) model.

[1525] Step 5:

[1526] The server uses NLP models to analyze the content of the files and extract business terms, abbreviations, and technical terms.

[1527] Step 6:

[1528] The server organizes the extracted business terms in list format and temporarily stores the analysis results.

[1529] 3. Storing terms in the database and standardizing notation

[1530] Step 7:

[1531] The server stores the extracted business terms in a database.

[1532] Step 8:

[1533] The server checks the newly extracted terms against existing terms in the database to see if there is a match.

[1534] Step 9:

[1535] If the term is spelled differently, the server will correct it to a unified spelling for consistency.

[1536] 4. Automatic generation and updating of Wiki pages

[1537] Step 10:

[1538] The server runs a script that automatically generates Wiki pages based on business terms in the database.

[1539] Step 11:

[1540] The server adds detailed descriptions and related links to each term to the Wiki page.

[1541] Step 12:

[1542] The server uploads the generated Wiki page to the company's internal Wiki system so that users can view it.

[1543] 5. Spelling Variation Detection and Notification

[1544] Step 13:

[1545] The server periodically monitors the file sharing system and analyzes newly uploaded documents.

[1546] Step 14:

[1547] The server detects spelling variations when a term in the parsed document does not match an existing list of terms.

[1548] Step 15:

[1549] When the server detects a spelling variation, it sends a notification email to the relevant user, prompting them to correct the spelling to the recommended spelling.

[1550] 6. Real-time annotation generation during meetings

[1551] Step 16:

[1552] The terminal receives voice input during the conference and acquires the voice data.

[1553] Step 17:

[1554] The server passes the received voice data to a voice recognition service and converts it into text.

[1555] Step 18:

[1556] The server detects business terms from the text data and generates annotations in real time.

[1557] Step 19:

[1558] The terminal displays the generated annotations on the screens of the conference participants to support their understanding of the terminology.

[1559] Example 1

[1560] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1561] Inconsistent notation and lack of maintenance in materials and documents are often problems in business activities. New employees and those transferred from other departments often have difficulty understanding business terms and abbreviations, which can lead to reduced work efficiency. Also, the use of technical terms during meetings often reduces participants' understanding. It is necessary to solve these problems and provide information quickly and clearly while maintaining consistency in business terminology.

[1562] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1563] In this invention, the server includes a means for monitoring documents and automatically detecting newly uploaded documents, a means for extracting business terms from documents using a natural language processing model, and a means for storing the extracted business terms in a database, checking for matches with existing terms, and standardizing their spellings to a consistent format. This allows for the consolidation and standardization of business terms. The server also includes a means for automatically generating and updating Wiki pages based on business terms and adding term explanations, a means for periodically monitoring the file system, detecting spelling variations in new documents, and sending notifications prompting corrections. The server also includes a means for receiving voice input during meetings, detecting business terms from the voice data, and generating and displaying annotations in real time. This allows for the rapid sharing and understanding of new information, improving work efficiency and accuracy.

[1564] "Materials" refers to information media such as documents, files, and reports used in business activities.

[1565] A "natural language processing model" refers to a model that makes full use of artificial intelligence technology to analyze and understand human language.

[1566] "Business jargon" refers to specialized words and abbreviations used in specific industries or jobs.

[1567] A "database" refers to a system for organizing and storing extracted information and data so that it can be used efficiently.

[1568] "Orthographic variation" refers to differences between words or terms that have the same meaning but are expressed in different forms.

[1569] A "Wiki page" refers to a web page that can be collaboratively edited on the Internet or an intranet.

[1570] A "file system" refers to the method and implementation for managing and storing data and files.

[1571] "Annotation" refers to explanations or explanatory text added to aid understanding.

[1572] "Voice input" refers to a method of inputting voice information into a computer via a microphone or the like.

[1573] MODE FOR CARRYING OUT THE INVENTION

[1574] The present invention is a system that uses AI to automatically create summaries of business terms and resolves issues such as lack of maintenance and inconsistent spelling. Specific embodiments of this system are described in detail below.

[1575] System configuration

[1576] This system is mainly composed of three elements: a server, a terminal, and a user. The server processes and monitors data, the terminal functions as an input and display device, and the user provides and uses information.

[1577] Key Features

[1578] Material monitoring and detection

[1579] The server monitors the file sharing system and automatically detects newly uploaded files. Specifically, the server uses a Python script to check the file system for changes at regular intervals and obtain a list of new or updated files. This operation utilizes the OS's file monitoring function and the cloud service's API.

[1580] Extracting business terms using NLP models

[1581] The server analyzes the documents using a natural language processing (NLP) model to extract business terms. This operation uses a pre-trained NLP model (e.g., BERT or spaCy). Depending on the file format, the server converts the file (e.g., PDF or Word) into text format, and then uses the NLP model to analyze and extract key terms.

[1582] Terminology storage in database and standardization of notation

[1583] The server stores the extracted business terms in a dedicated database. Here, it checks whether they match existing terms and standardizes them into consistent notations. MySQL, PostgreSQL, or other database systems are used. The server executes queries to check and update the contents of the database.

[1584] Automatic generation and updating of Wiki pages

[1585] The server automatically generates and updates Wiki pages based on business terms stored in the database. Wiki pages are generated using Markdown format or HTML templates. Specifically, the server retrieves terminology information from the database and applies it to templates to generate Wiki pages.

[1586] Spelling variation detection and notification

[1587] The server periodically monitors the file sharing system and compares new material with the existing database. If a spelling variation is detected, the server notifies the user via email or instant messaging services (e.g., Slack).

[1588] Real-time annotation during meetings

[1589] The device receives voice input during the meeting and converts the voice data into text. The Google Speech-to-Text API is used for speech recognition. For example, if someone says "ROI" during a meeting, the server analyzes it in real time and generates the annotation "ROI (Return on Investment)." The generated annotation is then displayed on the device of each meeting participant.

[1590] Specific examples

[1591] The following shows the process flow when a newly assigned employee downloads "Project Meeting Materials."

[1592] 1. Material detection and analysis:

[1593] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[1594] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[1595] 2. Terminology database storage and standardization:

[1596] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[1597] 3. Generate Wiki pages:

[1598] The server automatically generates Wiki pages based on the extracted terms, adding detailed descriptions and related links.

[1599] 4. Notification of spelling variations:

[1600] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[1601] 5. Real-time annotation during meetings:

[1602] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this as "ROI (Return on Investment)" and displays it in real time.

[1603] Prompt Sentence Examples

[1604] 1. Material detection and analysis:

[1605] "Detect newly uploaded conference materials and extract the technical terms they contain."

[1606] 2. Terminology database storage and standardization:

[1607] "Store the extracted terms in a database and check if they match existing terms."

[1608] 3. Generate Wiki pages:

[1609] "Generate wiki pages based on terms stored in your database and add detailed descriptions and related links."

[1610] 4. Notification of spelling variations:

[1611] "If you detect spelling variations in newly uploaded materials, please notify us of the recommended spelling."

[1612] 5. Real-time annotation during meetings:

[1613] "Extract technical terms from audio during meetings in real time and generate and display annotations."

[1614] The above describes a specific embodiment of the present invention. This system makes it possible to standardize business terminology and promote understanding, thereby realizing efficient business operations.

[1615] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1616] Step 1: Monitoring and detecting material

[1617] The server monitors the file sharing system. Specifically, the server scans the contents of the file system at regular intervals to detect newly added or updated files. The input is the file system path and the monitoring interval. The output is a list of new or updated files. The server also records file metadata (file name, creation date and time, and update date and time).

[1618] Step 2: Prepare the file for analysis

[1619] The server obtains the paths of the detected files and prepares them for analysis. First, it checks the file format (PDF, Word, etc.) and converts it to text data using an appropriate text conversion tool (e.g., PyPDF2 library for PDF, python-docx library for Word). The input is a list of detected files, and the output is text data.

[1620] Step 3: Extracting business terms using an NLP model

[1621] The server inputs the converted text data into an NLP model to extract business terms, abbreviations, and technical terms. This process uses a pre-trained NLP model (e.g., BERT or spaCy). The input is the text data, and the output is a list of extracted business terms. Specifically, terms such as "ROI" and "critical path" are extracted.

[1622] Step 4: Store business terms in a database

[1623] The server stores the extracted business terms in a database. During this process, it checks for matches with existing terms and standardizes the notation. The database used is MySQL or PostgreSQL. The input is the extracted business term list, and the output is the updated database. For example, if "ROI" and "return on investment" match, they are unified to "ROI."

[1624] Step 5: Automatically generate and update Wiki pages

[1625] The server automatically generates and updates Wiki pages based on business terms stored in the database. Markdown format and HTML templates are used for this process. The input is the business term information in the database, and the output is the generated or updated Wiki page. For example, the "ROI" page contains the description "Return on Investment."

[1626] Step 6: Detecting and reporting spelling variations

[1627] The server compares newly uploaded materials with the existing database to detect variations in notation. If any discrepancies are found during this process, a notification is sent to the user. Notification methods include email and instant messaging services (e.g., Slack). The input is the new material and database information, and the output is a notification to the user. For example, a notification recommending that "return on investment" be changed to "ROI" is sent.

[1628] Step 7: Real-time annotation during the meeting

[1629] The device receives voice input during the meeting and converts the voice data into text. This process uses the Google Speech-to-Text API. The input is the voice data from the meeting, and the output is text data. The converted text data is sent to the server, and if a business term is detected, an annotation is generated in real time and displayed on the device. For example, if "ROI" is spoken during a meeting, the annotation "ROI (Return on Investment)" is generated and displayed.

[1630] (Application example 1)

[1631] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1632] In today's business environment, multiple documents and materials are frequently uploaded, often containing many technical terms and abbreviations. This makes it difficult for newly assigned employees and external participants to quickly adapt to the work. Furthermore, inconsistencies in the spelling of business terms occur, reducing the consistency and efficiency of information. There is a need to solve these problems and improve business efficiency.

[1633] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1634] In this invention, the server includes: means for monitoring materials and documents and detecting newly uploaded materials and documents; means for extracting business terms from materials and documents using a natural language processing model; means for storing the extracted business terms in a database, checking for matches with existing terms, and standardizing spellings; means for automatically generating and updating information pages based on the business terms and adding term explanations; means for periodically monitoring the file sharing system, detecting spelling variations in new documents, and sending notifications prompting corrections; means for receiving voice input during meetings, detecting business terms from the voice data, and generating and displaying annotations in real time; means for analyzing the camera feed of a visual device and extracting text from documents; means for displaying business terms as annotations from the extracted text; and means for analyzing the voice input of the visual device, performing speech recognition, and providing detailed explanations of the relevant business terms by voice. This enables newly assigned employees to quickly adapt to their work, promotes standardization of business terminology, and promotes understanding.

[1635] "Materials" refers to documents and other information related to business.

[1636] "Document" refers to a medium containing information stored as printed material or digital data.

[1637] "Monitoring" means that a system periodically or continuously checks a particular file sharing system or database to detect newly added information.

[1638] A "natural language processing model" refers to a computer program or algorithm that analyzes human language and extracts useful information from text data.

[1639] "Business jargon" refers to technical terms and abbreviations frequently used in a particular business or industry.

[1640] A "database" refers to a system for efficiently storing, retrieving, and managing information.

[1641] "Unification of spelling" refers to the standardization of words with the same meaning that are expressed in different spellings or writing styles into a consistent format.

[1642] An "information page" is a web page or document that provides detailed explanations and related information about a particular term or concept.

[1643] A "file sharing system" refers to a system that allows multiple users within a specific network to share and access files.

[1644] "Spelling variation" refers to the situation where the same term or concept is written in different ways.

[1645] "Notification" refers to a system-generated message that informs the user of a particular event or situation.

[1646] "Voice input during a meeting" refers to the process by which voice data spoken during a meeting is received by the system.

[1647] "Audio Data" refers to a digital recording of sound collected by a microphone or the like.

[1648] "Generating and displaying annotations in real time" refers to the process of instantly analyzing voice or text input and displaying additional information based on the results.

[1649] "Visual devices" refer to devices worn by users that visually present information, such as smart glasses and head-mounted displays.

[1650] "Camera feed" refers to video data captured by a camera in real time.

[1651] "Extracting text from a document" refers to the process of analyzing character information from camera images or scanned data and extracting that text data.

[1652] "Speech recognition" refers to the technology and algorithms that analyze audio signals and identify meaningful words and sentences from them.

[1653] This invention is a system that uses AI to consolidate business terminology, particularly to improve efficiency in factory work environments. This system is composed of a server, terminals, and users. The specific configuration and functions of this system are described below.

[1654] System configuration and functions

[1655] 1. Monitoring and detection of materials and documents

[1656] The server monitors the factory's file sharing system and database, automatically detecting newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[1657] 2. Extracting business terms using natural language processing models

[1658] The server uses natural language processing (NLP) models to analyze documents and extract business terms, abbreviations, and technical terms, using NLP libraries such as SpaCy.

[1659] 3. Storing terms in the database and standardizing notation

[1660] The server stores the extracted business terms in a database and checks whether they match existing terms. If there are any variations in spelling, they are unified to create a consistent term list. For this purpose, an SQL-based database management system (e.g., MySQL) is used.

[1661] 4. Automatic generation and updating of information pages

[1662] The server automatically generates and updates information pages based on business terms stored in the database, including detailed explanations and related links for each term.

[1663] 5. Spelling Variation Detection and Notification

[1664] The server periodically monitors the file sharing system and analyzes newly uploaded documents. If it detects a spelling variation that does not match the existing term list, it sends a notification to the relevant user urging them to correct the variation. The notification system includes email notifications and push notifications.

[1665] 6. Real-time voice annotation

[1666] The device receives voice input during the meeting and converts the voice data into text. The server detects business terms from the text and generates annotations in real time and sends them to the device, allowing meeting participants to instantly understand the meaning of the terms.

[1667] 7. Visual Device Camera Feed Analysis

[1668] A visual device (e.g., smart glasses) extracts text from a document through a camera feed. After extracting the text using Tesseract OCR, the server analyzes the text and displays real-time annotations for business terms.

[1669] 8. Analysis and presentation of voice input

[1670] The voice input from the visual device is analyzed and speech recognition is performed using the Google Cloud Speech-to-Text API. For recognized terms, the server retrieves detailed descriptions and presents them as audio through the visual device's speaker.

[1671] Specific examples

[1672] For example, suppose a newly assigned employee views a document containing the term "ROI (Return on Investment)." The camera in the vision device captures this, and the server performs text analysis, displaying the annotation "Return on Investment" for the term "ROI." Furthermore, if the employee voice-inputs, "Tell me more about Return on Investment," the vision device recognizes the voice and provides a detailed explanation.

[1673] This will generate a prompt like this:

[1674] "A new document has been uploaded. This document contains the term 'return on investment'. Could you please explain this term in more detail?"

[1675] In this way, the present invention promotes understanding of business terms in a factory work environment, enabling efficient work execution.

[1676] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1677] Step 1:

[1678] The server monitors the file sharing system within the factory and detects newly uploaded materials and documents. As input, it receives directory information and file lists from the file sharing system and identifies new files based on timestamps. After processing this data, it outputs a list of newly uploaded files.

[1679] Step 2:

[1680] The server uses a natural language processing model (e.g., SpaCy) to extract business terms from newly detected materials and documents. As input, it receives the contents of the files detected in step 1 in text format and analyzes the text data. This data processing results in the output of a list of extracted business terms.

[1681] Step 3:

[1682] The server stores the extracted business terms in a database, checks for matches with existing terms, and standardizes notation. As input, it receives the list of business terms extracted in step 2 and compares them with existing terms in the database. This data processing results in a consistent list of terms being output.

[1683] Step 4:

[1684] The server automatically generates and updates information pages based on business terms stored in the database. It receives an updated term list as input and adds detailed explanations and related links for each term. This data calculation results in the output of the latest information page.

[1685] Step 5:

[1686] The server periodically monitors the file sharing system, detects spelling variations that do not match the existing terminology list, and sends a notification prompting correction. It receives the content of newly uploaded documents as input and compares it with the existing terminology list. After processing this data, it outputs a notification containing mismatched terms and suggested corrections.

[1687] Step 6:

[1688] The device converts voice input received during the meeting into text and sends it to the server. It receives the voice data of the meeting participants as input and performs speech recognition using the Google Cloud Speech-to-Text API. This data is then processed and the text data is output.

[1689] Step 7:

[1690] The server generates annotations in real time for business terms detected from the voice input and sends them to the terminal. As input, it receives the text data output in step 6, analyzes and extracts business terms, and generates annotations for them. This data processing results in the output of annotated text.

[1691] Step 8:

[1692] The vision device analyzes the camera feed and extracts text from the document. It receives the video data from the camera feed as input and performs character recognition using Tesseract OCR. It then processes this data and outputs the extracted text data.

[1693] Step 9:

[1694] The server generates annotations of business terms from the extracted text and displays them on a visual device. It receives the text data output in step 8 as input, analyzes and extracts business terms, and generates annotations for them. This data processing results in the output of annotated text.

[1695] Step 10:

[1696] The visual device analyzes voice input and performs speech recognition. It receives the user's voice data as input and performs speech recognition using the Google Cloud Speech-to-Text API. It then processes this data and outputs the recognized text data.

[1697] Step 11:

[1698] The server obtains detailed explanations for the recognized terms and presents them as audio through the speaker of the visual device. As input, it receives the text data output in step 10, generates detailed explanations, and sends them as audio data to the visual device. This data processing results in the audio data of the detailed explanation being output.

[1699] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1700] MODE FOR CARRYING OUT THE INVENTION

[1701] The present invention is a system that uses AI to automatically create a wiki that summarizes business terminology, resolves issues such as lack of maintenance and inconsistent spelling, and combines it with an emotion engine that recognizes user emotions. An embodiment of this system is described in detail below.

[1702] System Overview

[1703] The system has the following main functions:

[1704] 1. Monitoring and detection of materials and documents

[1705] 2. Extracting business terms using NLP models

[1706] 3. Storing terms in the database and standardizing notation

[1707] 4. Automatic generation and updating of Wiki pages

[1708] 5. Spelling Variation Detection and Notification

[1709] 6. Real-time annotation generation during meetings

[1710] 7. Emotion engine that recognizes user emotions

[1711] Monitoring and detection of materials and documents

[1712] The server monitors the file sharing system and automatically detects newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[1713] Extracting business terms using NLP models

[1714] The server uses natural language processing (NLP) models to analyze materials and documents and extract business terms, abbreviations, and technical terms, which then lists specific business terms and allows for efficient analysis.

[1715] Terminology storage in database and standardization of notation

[1716] The server stores the extracted business terms in a database, checks whether they match existing terms, and corrects any inconsistencies in spelling to create a consistent term list.

[1717] Automatic generation and updating of Wiki pages

[1718] The server automatically generates and updates Wiki pages based on business terms stored in the database, adding detailed explanations and related links for each term to make them easily accessible to employees.

[1719] Spelling variation detection and notification

[1720] The server periodically monitors the file sharing system and analyzes newly uploaded documents. If it detects a spelling variation that does not match the existing terminology list, it notifies the user and encourages them to correct the spelling.

[1721] Real-time annotation during meetings

[1722] The device receives voice input during the meeting and converts the voice data into text. The server detects business terms from the converted data and generates annotations in real time, which are sent to the device. The device then displays the annotations on the screen of the meeting participants to support understanding.

[1723] Emotion engine that recognizes user emotions

[1724] The server analyzes the user's voice and text data, recognizes the user's emotions using an emotion engine, and provides appropriate feedback and advice based on the recognized emotions.

[1725] Specific examples

[1726] Consider the process flow when a newly assigned employee downloads "Project Meeting Materials." This material contains specific technical terms and abbreviations.

[1727] 1. Material detection and analysis:

[1728] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[1729] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[1730] 2. Terminology database storage and standardization:

[1731] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[1732] 3. Generate Wiki pages:

[1733] The server automatically generates Wiki pages based on the extracted terms, adding detailed descriptions and related links.

[1734] 4. Notification of spelling variations:

[1735] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[1736] 5. Real-time annotation during meetings:

[1737] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this as "ROI (Return on Investment)" and displays it in real time.

[1738] 6. Leveraging the Emotion Engine:

[1739] The server analyzes the voice data and text data entered by the user during the meeting and uses an emotion engine to recognize the user's emotions. For example, if the user is feeling stressed during the meeting, the server will provide appropriate feedback to support the progress of the meeting.

[1740] The processing flow will be explained below.

[1741] MODE FOR CARRYING OUT THE INVENTION

[1742] Monitoring and detection of materials and documents

[1743] Step 1:

[1744] The server accesses the file sharing system and monitors newly uploaded materials and documents.

[1745] Step 2:

[1746] The server scans the directory of the file sharing system at regular time intervals to check whether new files exist.

[1747] Step 3:

[1748] The server detects the newly uploaded file and retrieves the file's metadata (such as name, upload date and time, file type, etc.).

[1749] Extracting business terms using NLP models

[1750] Step 4:

[1751] The server inputs the detected new files into a natural language processing (NLP) model.

[1752] Step 5:

[1753] The server uses NLP models to analyze the content of the files and extract business terms, abbreviations, and technical terms.

[1754] Step 6:

[1755] The server organizes the extracted business terms in list format and temporarily stores the analysis results.

[1756] Terminology storage in database and standardization of notation

[1757] Step 7:

[1758] The server stores the extracted business terms in a database.

[1759] Step 8:

[1760] The server checks the newly extracted terms against existing terms in the database to see if there is a match.

[1761] Step 9:

[1762] If the term is spelled differently, the server will correct it to a unified spelling for consistency.

[1763] Automatic generation and updating of Wiki pages

[1764] Step 10:

[1765] The server runs a script that automatically generates Wiki pages based on business terms in the database.

[1766] Step 11:

[1767] The server adds detailed descriptions and related links to each term to the Wiki page.

[1768] Step 12:

[1769] The server uploads the generated Wiki page to the company's internal Wiki system so that users can view it.

[1770] Spelling variation detection and notification

[1771] Step 13:

[1772] The server periodically monitors the file sharing system and analyzes newly uploaded documents.

[1773] Step 14:

[1774] The server detects spelling variations when a term in the parsed document does not match an existing list of terms.

[1775] Step 15:

[1776] When the server detects a spelling variation, it sends a notification email to the relevant user, prompting them to correct the spelling to the recommended spelling.

[1777] Real-time annotation during meetings

[1778] Step 16:

[1779] The terminal receives voice input during the conference and acquires the voice data.

[1780] Step 17:

[1781] The server passes the received voice data to a voice recognition service and converts it into text.

[1782] Step 18:

[1783] The server detects business terms from the text data and generates annotations in real time.

[1784] Step 19:

[1785] The terminal displays the generated annotations on the screens of the conference participants to support their understanding of the terminology.

[1786] Emotion engine that recognizes user emotions

[1787] Step 20:

[1788] The terminal collects voices during the conference and text data entered by the user.

[1789] Step 21:

[1790] The server passes the collected voice and text data to an emotion engine to analyze the user's emotions.

[1791] Step 22:

[1792] The server generates appropriate feedback and advice based on the recognized emotions.

[1793] Step 23:

[1794] The terminal displays the generated feedback and advice on the screens of the conference participants to support the progress of the conference.

[1795] Example 2

[1796] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1797] When managing materials and documents, it is necessary to efficiently monitor and detect newly uploaded information. There is also the issue of eliminating inconsistencies in notation due to the existence of unstandardized business terms and abbreviations, and creating and maintaining a consistent terminology list. Furthermore, there is a need to streamline the provision of information during meetings and to implement a feedback system based on user sentiment.

[1798] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1799] In this invention, the server includes: means for monitoring materials and documents and detecting newly uploaded materials and documents; means for extracting business terms from materials and documents using a natural language processing model; means for storing the extracted business terms in a database, checking for matches with existing terms, and standardizing notation; means for automatically generating and updating information pages based on the business terms and adding explanations of the terms; means for periodically monitoring the information sharing system, detecting notation variations in new documents, and sending notifications prompting corrections; means for receiving voice input during a meeting, detecting business terms from the voice data, and generating and displaying annotations in real time; and means for analyzing user voice and text data, recognizing user emotions, and providing feedback and advice. This enables efficient document management, standardization of notation, maintaining consistency, efficient information provision during meetings, and appropriate feedback based on user emotions.

[1800] "Materials and documents" refers to electronic files containing text data and graphical data, such as reports, plans, technical manuals, procedure manuals, e-mails, memos, and meeting materials used inside and outside a company.

[1801] "Monitoring" means constantly checking operations such as adding, changing, or deleting files that are performed on a specific system or network, and detecting them when certain conditions are met.

[1802] A "natural language processing model" refers to an algorithm or machine learning model designed to understand, analyze, and generate human language, specifically extracting business terms and analyzing text.

[1803] "Business jargon" refers to the technical terms, abbreviations, and phrases used within a particular company or industry that are an important part of business processes and communication.

[1804] "Database" refers to a data storage system that stores data in an organized manner and is designed to make it easy to access and manage.

[1805] "Unification of notation" refers to the process of changing identical terms with different notations or expressions into a consistent format to maintain data integrity and consistency.

[1806] An "information page" is a web page or digital document that aggregates business terms and related information and displays them in an easy-to-reference format.

[1807] An "information sharing system" is a platform that allows multiple users to upload, download, and share files and information, and includes cloud storage and corporate network drives.

[1808] "Spelling variation" refers to the use of different expressions or spellings that have the same meaning, which creates problems that make understanding and management complicated.

[1809] "Sending a notification" refers to the act of communicating alerts or information to a user based on a specific event or condition detected by the system.

[1810] "Receiving audio input" refers to the process of collecting audio data through a microphone or other audio capture device and passing that data to the system.

[1811] "Audio data" is data that is a digital representation of human speech recorded via a voice input device.

[1812] "Generating annotations in real time" refers to the process of instantly adding relevant explanations and information to input speech or text data and displaying them.

[1813] "Recognizing user emotions" refers to the technical process of determining the emotions a user is feeling at that time through analysis of voice and text data.

[1814] "Providing feedback and advice" refers to the process by which the system suggests appropriate responses or recommended actions based on the user's perceived emotions.

[1815] MODE FOR CARRYING OUT THE INVENTION

[1816] The present invention is a system that uses AI to automatically create information pages summarizing business terms, eliminating problems such as lack of maintenance and inconsistent spelling, and combining this with an emotion engine that recognizes user emotions. An embodiment of this system is described in detail below.

[1817] Hardware and software used

[1818] The system utilizes the following hardware and software:

[1819] Hardware: Server, terminal, audio input device (microphone)

[1820] File sharing systems: Dropbox, Google Drive, etc.

[1821] Natural language processing models: Google BERT, OpenAI GPT-3, etc.

[1822] Database: MySQL, PostgreSQL

[1823] Emotion engine: IBM Watson, Azure Cognitive Services

[1824] Speech recognition software: Google Speech-to-Text, Microsoft Azure Speech Service

[1825] Overall flow

[1826] The server monitors the file sharing system to detect and analyze newly uploaded materials and documents. A natural language processing model is used for the analysis, which extracts business terms, abbreviations, and technical terms from the materials. Before storing the extracted terms in the database, they are compared with existing terms, and any variations in spelling are unified.

[1827] Next, the server automatically generates and updates information pages based on the business terms stored in the database. The generated information pages include detailed explanations of the terms and related links. The server also periodically monitors the information sharing system to detect spelling variations in new documents. If any are detected, a notification is sent to the relevant user.

[1828] During a meeting, the device receives voice input and converts the voice data into text. The text data is then sent to a server, which detects business terms and generates annotations in real time. The annotations are then sent to the device and displayed on the screens of meeting participants.

[1829] Finally, the server analyzes the user's voice and text data, and uses an emotion engine to recognize the user's emotions. Based on the recognized emotions, the server provides appropriate feedback and advice.

[1830] Specific examples

[1831] The process flow when a newly assigned employee downloads "Project Meeting Materials" is as follows: This material contains specific technical terms and abbreviations.

[1832] 1. Material detection and analysis:

[1833] The server detects "Project Meeting Materials.pdf" from the file sharing system.

[1834] The server uses an NLP model to extract business terms such as "ROI" and "critical path" from the documents.

[1835] 2. Terminology database storage and standardization:

[1836] The server stores terms such as "ROI (Return on Investment)" and "Critical Path" in a database and checks whether they match existing terms. If they do not match, it registers a new term, and if they do match, it standardizes the notation.

[1837] 3. Generate information page:

[1838] The server automatically generates information pages based on the extracted terms, adding detailed descriptions and related links.

[1839] 4. Notification of spelling variations:

[1840] If a newly uploaded document contains a "Return on Investment", the server will detect this and notify the user to modify it to the recommended "ROI".

[1841] 5. Real-time annotation during meetings:

[1842] The terminal receives the part of the audio during the conference where "ROI" is spoken, and the server annotates this in real time as "ROI (Return on Investment)" and displays it.

[1843] 6. Leveraging the Emotion Engine:

[1844] The server analyzes the voice data and text data entered by the user during the meeting and uses an emotion engine to recognize the user's emotions. For example, if the user is feeling stressed during the meeting, the server will provide appropriate feedback to support the progress of the meeting.

[1845] Prompt Sentence Examples

[1846] Here are some example prompts to input to a generative AI model:

[1847] Please analyze the business terms contained in newly uploaded documents and standardize the notation of the extracted business terms to a consistent format.

[1848] "Analyze audio during meetings in real time and annotate and display important business terms."

[1849] "Recognize emotions based on the user's voice and text data and provide appropriate feedback if they are feeling stressed."

[1850] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1851] Step 1:

[1852] The server monitors the file sharing system to detect newly uploaded materials or documents. In this step, the server checks whether a new file has been uploaded from the file sharing system (e.g., Dropbox, Google Drive) and obtains the path and name of the file. Specifically, the server scans the file list and detects the new file "Project Meeting Materials.pdf."

[1853] Input: A file uploaded to a file sharing system.

[1854] Output: The path and name of the file found.

[1855] Step 2:

[1856] The server analyzes the detected file using a natural language processing (NLP) model to extract business terms. In this step, the server uses an NLP model (e.g., Google BERT, OpenAI GPT-3) to list business terms such as "ROI" and "critical path" from the text data in the file. Specifically, the server analyzes "Project Meeting Materials.pdf" and lists business terms.

[1857] Input: The path and name of the detected file.

[1858] Output: A list of extracted business terms.

[1859] Step 3:

[1860] The server stores the extracted business terms in a database, checks whether they match existing terms, and unifies the spelling. In this step, before storing the extracted terms in a database (e.g., MySQL, PostgreSQL), the server compares them with existing terms and unifies any spelling variations. Specifically, it checks whether "ROI" already exists in the database and whether it matches "Return on Investment," and unifies them.

[1861] Input: A list of extracted business terms.

[1862] Output: A consistent list of terms stored in a database.

[1863] Step 4:

[1864] The server automatically generates and updates information pages based on the business terms stored in the database. In this step, the server generates and updates information pages containing detailed explanations and related links based on the newly extracted business terms. Specifically, a detailed explanation of "ROI (Return on Investment)" is added to the information page.

[1865] Input: A consistent list of terms stored in a database.

[1866] Output: Generated and updated information page.

[1867] Step 5:

[1868] The server periodically monitors the information sharing system, detects spelling variations in new documents, and sends a notification to prompt correction. In this step, the server compares the newly detected document with the existing term list, and if a spelling variation is found, it sends a notification to the relevant user. Specifically, if "return on investment" is newly detected, it sends a notification to unify it to "ROI."

[1869] Input: The newly discovered document.

[1870] Output: Notification of spelling variations.

[1871] Step 6:

[1872] The device receives voice input during the meeting and converts the voice data into text. In this step, the device uses speech recognition software (e.g., Google Speech-to-Text, Microsoft Azure Speech Service) to convert the voice during the meeting into text in real time. Specifically, the device converts the part where "ROI" is spoken during the meeting into text.

[1873] Input: Audio input during a meeting.

[1874] Output: The audio data converted to text.

[1875] Step 7:

[1876] The server detects business terms from the converted speech data, generates annotations in real time, and sends them to the terminal. In this step, the server detects "ROI" from the converted data and generates the annotation "ROI (Return on Investment)." After generating the annotation, it sends it to the terminal and displays it. Specifically, the annotation "ROI (Return on Investment)" is displayed on the screen of the conference participants.

[1877] Input: Transcribed audio data.

[1878] Output: The generated annotations.

[1879] Step 8:

[1880] The server analyzes the user's voice and text data and recognizes the user's emotions using an emotion engine. In this step, the server analyzes the user's emotions using an emotion engine (e.g., IBM Watson, Azure Cognitive Services). Based on the recognized emotions, the server provides appropriate feedback and advice. Specifically, if the server determines that the user is feeling stressed, it sends feedback encouraging the user to relax.

[1881] Input: User voice and text data.

[1882] Output: Emotion-based feedback and advice.

[1883] (Application example 2)

[1884] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1885] Modern factories generate a large number of documents and materials every day, and these documents contain a large amount of technical terminology. However, if these terms are not used consistently, it becomes difficult to understand and share information. In addition, there is a lack of a way to recognize workers' emotions and provide appropriate feedback, which can lead to reduced work efficiency and the risk of mistakes. There is a need for a system that can resolve these issues and provide support based on workers' emotions while maintaining consistency in work terminology.

[1886] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for monitoring materials and documents and detecting newly uploaded materials and documents; means for extracting business terms from materials and documents using a natural language processing model; means for storing the extracted business terms in a database, checking for matches with existing terms, and unifying notations; means for automatically generating and updating knowledge pages based on the business terms and adding term explanations; means for periodically monitoring the data storage system, detecting notation inconsistencies in new documents, and sending notifications prompting corrections; means for receiving voice input during work, detecting business terms from the voice data, and generating and displaying annotations in real time; and means for recognizing the user's emotions and providing appropriate instructions and feedback. This makes it possible to improve work efficiency while ensuring consistency in business terminology.

[1887] "Materials and documents" refers to all documents generated within the factory, such as reports, plans, procedures, and manuals.

[1888] "Means for detecting newly uploaded files" refers to the ability to monitor the data storage system and automatically detect newly uploaded files.

[1889] A "natural language processing model" is a software algorithm for analyzing the content of a document and extracting specific business terms.

[1890] "Business terms" refers to technical terms, abbreviations, and keywords used in a particular business or field of expertise.

[1891] "Means of storing in a database, checking for matches with existing terms, and standardizing notation" refers to a function that saves extracted business terms and compares them with existing data to ensure consistency in notation.

[1892] "Means for generating and updating knowledge pages and adding explanations of terms" is a function that automatically creates and updates information pages (Wiki pages) that include detailed explanations of business terms.

[1893] "Means for periodically monitoring the data storage system, detecting inconsistencies in notation in new documents, and sending notifications prompting corrections" refers to a notification function that monitors new document uploads, detects inconsistencies in notation, and requests the user to make corrections.

[1894] "Means for receiving voice input during work, detecting business terms from the voice data, and generating and displaying annotations in real time" refers to a function that analyzes voice information generated during work, recognizes technical terms, and immediately displays their definitions and explanations.

[1895] "Means for recognizing the user's emotions and providing appropriate instructions and feedback" refers to a function that detects the user's emotions from their words and actions, and provides advice and support according to the situation.

[1896] MODE FOR CARRYING OUT THE INVENTION

[1897] This invention is a system that maintains consistency in business terminology within a factory, recognizes the emotions of workers, and provides appropriate feedback. The system's main functions are monitoring materials and documents, extracting business terminology using natural language processing, storing it in a database and standardizing its notation, automatically generating Wiki pages, detecting and notifying inconsistencies in notation, generating annotations in real time during work, and recognizing emotions using an emotion engine.

[1898] System Overview

[1899] 1. Monitoring and detection of materials and documents

[1900] The server monitors the factory's data storage system and automatically detects newly uploaded materials and documents, ensuring that the latest information is always available for analysis.

[1901] 2. Extracting business terms using natural language processing models

[1902] The server uses a natural language processing model (NLP model) to analyze materials and documents and extract business terms, a process carried out using automated algorithms.

[1903] 3. Storing terms in the database and standardizing notation

[1904] The extracted business terms are stored in a database by the server. A consistent term list is maintained by checking whether they match existing terms and correcting any inconsistencies in their spellings to make them consistent.

[1905] 4. Automatic generation and updating of knowledge pages

[1906] The server automatically generates and updates knowledge pages based on the business terms stored in the database, adding detailed explanations and related links for each term to make them easily accessible to workers.

[1907] 5. Detecting and notifying inconsistencies

[1908] The server periodically monitors the data storage system and analyzes newly uploaded documents. If it detects any inconsistencies in spelling that do not match the existing terminology list, it notifies the user and prompts them to make corrections.

[1909] 6. Real-time annotation generation while working

[1910] The device receives voice input during work and converts the voice data into text. The server detects business terms from the converted data and generates annotations in real time, which are sent to the device. The device then displays these on the worker's screen to help them understand the work.

[1911] 7. Emotion Recognition with Emotion Engine

[1912] The server analyzes voice and text data during work and uses an emotion engine to recognize the worker's emotions. Based on the recognized emotions, the server provides appropriate feedback and advice.

[1913] Specific examples

[1914] For example, if a factory uploads a document called "New Product Line Plan.pdf", the following process will be performed based on this file:

[1915] 1. The server detects "New Product Line Plan.pdf" and analyzes its contents.

[1916] 2. Use a natural language processing model to extract technical terms such as "TPM (Total Productive Maintenance)."

[1917] 3. Store the extracted business terms in a database, and check for and correct any inconsistencies in notation.

[1918] 4. Automatically generate knowledge pages based on terms and provide related information.

[1919] 5. If there are any inconsistencies in the notation of newly uploaded materials, notify the worker and request corrections.

[1920] 6. When you encounter technical terms while working, their definitions are displayed in real time.

[1921] 7. Provide feedback if workers are stressed.

[1922] Prompt Sentence Examples

[1923] For example, you can input the following prompts into a generative AI model:

[1924] Extract technical terms from the factory's "New Product Line Plan.pdf" and generate knowledge pages based on their definitions. Also, analyze the emotions of workers and display the message "Calm down, everything is under control" when they are feeling stressed.

[1925] The above is a detailed description of the mode for carrying out the invention.

[1926] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1927] Step 1:

[1928] The server monitors the data storage system in the factory to detect newly uploaded materials or documents. When a new document is detected, it obtains its path. This step requires the directory path of the data storage system as input and the path of the detected new document as output.

[1929] Step 2:

[1930] The server uses a natural language processing model (NLP model) to analyze newly detected materials and documents and extract business terms. A list of extracted business terms is generated. This step requires the text data of the document as input and produces a list of business terms as output.

[1931] Step 3:

[1932] The server accesses the database and stores the extracted business terms. It also checks for matches with existing terms and for inconsistencies in spelling. If necessary, it makes corrections to unify spelling. This step requires a list of business terms as input and an updated database as output.

[1933] Step 4:

[1934] The server automatically generates and updates knowledge pages based on the business terms in the database. Specifically, it creates pages that include detailed explanations and related information for each term. This step requires business terms in the database as input, and the generated knowledge pages as output.

[1935] Step 5:

[1936] The server periodically monitors the data storage system and analyzes newly uploaded documents. If it detects a spelling inconsistency that does not match the existing terminology list, it sends a notification to the relevant user to prompt them to correct it. This step requires the new document and the existing terminology list as input, and a spelling inconsistency notification as output.

[1937] Step 6:

[1938] The terminal receives voice input during work and converts the voice data into text. The server then detects business terms from the converted data and generates annotations in real time, which are sent to the terminal. The terminal then displays the annotations on the worker's screen. This step requires voice data as input and produces annotations to be displayed as output.

[1939] Step 7:

[1940] The server analyzes the voice and text data during work and recognizes the worker's emotions using an emotion engine. Based on the recognized emotions, the server provides appropriate feedback and advice. This step requires voice and text data as input, and feedback and advice are obtained as output.

[1941] This will enable the entire system to function consistently, ensuring consistency in business terminology, improving work efficiency, and creating a system that provides support based on the emotions of workers.

[1942] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1943] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1944] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1945] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1946] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other. 【1...

Claims

1. A means of monitoring materials and documents to detect newly uploaded materials and documents; A means of extracting business terms from materials and documents using natural language processing models; A method for storing the extracted business terms in a database, checking for matches with existing terms, and standardizing notation. A method to automatically generate and update Wiki pages based on business terms and add explanations of terms, A method for periodically monitoring the file sharing system, detecting spelling variations in new documents, and sending notifications prompting corrections; A means for receiving voice input during a meeting, detecting business terms from the voice data, and generating and displaying annotations in real time; A system including:

2. The system of claim 1 further comprising means for a user who receives the notification to modify the document to the recommended notation.

3. 2. The system of claim 1, further comprising a speech recognition means for converting voice data during the conference into text.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A