system

The system automates video data analysis by detecting and indexing events in natural language, generating reports, and incorporating emotional feedback, addressing inefficiencies in conventional methods.

JP2026070272APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Conventional methods for analyzing video data are time-consuming and labor-intensive, requiring significant effort to find specific events and create reports, leading to reduced work efficiency and increased burden on supervisors.

Method used

A system that receives, stores, analyzes, and indexes video data to detect specific events, converts them into natural language, and automatically generates reports, incorporating user feedback and emotional analysis for enhanced user experience.

Benefits of technology

The system efficiently automates the analysis and reporting process, reducing supervisor burden and enabling rapid, accurate detection of events with integrated emotional insights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070272000001_ABST
    Figure 2026070272000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of receiving video data and saving it to storage, A means for analyzing stored video data and detecting specific events within the video, A means of converting detected events into natural language, A method for indexing events within a video based on verbalized information, A means of providing an interface for users to view and edit indexed events, A means of automatically generating reports using edited information, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional method, it takes a huge amount of time and labor to visually check video data, and particularly when there is a large amount of data, the work efficiency is significantly reduced. Also, it is difficult to find a specific event, and it takes a lot of effort to create a work report. The present invention aims to solve these problems and reduce the burden on work supervisors.

Means for Solving the Problems

[0005] The present invention solves these problems by providing a system that includes means for receiving video data and storing it in storage, means for analyzing the stored video data and detecting specific events in the video, means for converting the detected events into natural language, means for indexing events in the video based on the verbalized information, means for providing an interface for users to review and edit the indexed events, and means for automatically generating a report using the edited information.

[0006] "Video data" refers to information that records visual information in digital format.

[0007] "Storage" refers to hardware or software used to store digital information.

[0008] "Analysis" is the process of examining digital data and identifying its constituent elements and characteristics.

[0009] An "event" refers to a specific situation or series of actions, specifically the concrete actions or occurrences within the video being analyzed.

[0010] "Natural language" refers to the language that humans use on a daily basis, and it is the language that is used without being converted into a form that is easily understood by machines through programs.

[0011] Indexing is the process of adding reference markers to data to enable quick access to information.

[0012] An "interface" is a user interface that provides means of input and output, enabling interaction between a user and a system.

[0013] A "report" is a document created to record and communicate the details, results, and analysis of an activity. [Brief explanation of the drawing]

[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

MODE FOR CARRYING OUT THE INVENTION

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention is a system that efficiently analyzes video data, verbalizes and indexes specific events, thereby reducing the burden on supervisors and automating report creation. The system operates as follows:

[0036] First, the user uploads a video of their work from their device to the server. The uploaded video is automatically saved to storage by the server. The storage system maintains metadata about the data and is managed to ensure that subsequent processing proceeds smoothly.

[0037] Next, the server sequentially analyzes the stored video data. The analysis process examines the video frame by frame, using an AI algorithm to detect specific events and anomalies within the video. This AI algorithm utilizes object detection technology and records event timestamps.

[0038] Detected events are automatically translated into natural language by the server. This translation makes the data more human-readable. In addition, if necessary, speech synthesis technology is used to convert the data into speech.

[0039] Subsequently, the server indexes the relevant portion of the video based on the verbalized event information. This indexing process makes it possible to easily search for and play specific events.

[0040] The completed index data is provided as an interface for users to review and edit. Through this interface, users can view event details and, if necessary, modify descriptions or add additional notes.

[0041] Finally, the server automatically generates a report based on the edited data. The report can be exported in formats such as PDF and HTML and shared on other digital platforms. This allows all stakeholders to efficiently share information and address issues quickly.

[0042] As a concrete example, consider a video taken on a manufacturing line. When a user uploads this video to the system, the AI ​​automatically detects errors in handling parts and generates a text index such as "Incorrect placement of part A." The user can review this information through the interface and add further comments. When creating a report, a detailed report containing these events is automatically generated and shared with the relevant departments.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] The user uploads work video data from their terminal to the server. At this time, the system performs user authentication to confirm that the user has legitimate access rights.

[0046] Step 2:

[0047] The server saves the received video data to storage. During saving, it analyzes metadata related to the video (such as the date and time of shooting and camera location information) and registers it in the database.

[0048] Step 3:

[0049] The server divides the stored video data into frames and prepares them for analysis. This division is based on the video's frame rate and optimized for subsequent analysis.

[0050] Step 4:

[0051] The server uses an AI model to process each frame and detect specific events or anomalies within the video. For example, it might use an object detection algorithm to track the movement of an object of interest.

[0052] Step 5:

[0053] The server converts detected events into natural language text. Here, predefined rules and contextual analysis are used to express the information in a language that is easily understandable to humans.

[0054] Step 6:

[0055] The server indexes each event location in the video based on the verbalized event information. This process records detailed information, including the time and location of the event.

[0056] Step 7:

[0057] Users can access the server through their terminal and view an indexed list of events from the interface. Users can then view the list, select a specific event, and edit or add details.

[0058] Step 8:

[0059] The server automatically generates reports based on information edited by the user. The report's structure is determined, and event information is incorporated in a predetermined format.

[0060] Step 9:

[0061] The server exports the final report in PDF or HTML format and distributes it to the designated stakeholders. This output is then integrated into the company's workflow as needed.

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] This invention aims to automate the process of analyzing and reporting video information, thereby reducing the burden on supervisors and enabling efficient information sharing. Conventional manual video analysis and report creation are time-consuming, labor-intensive, and prone to human error. Therefore, this invention provides a system to address these issues.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes means for receiving video information and storing it in a storage device, means for analyzing the stored video information and detecting specific events, and means for converting the detected events into natural language. This enables the rapid and accurate detection of events in the video, allows users to easily review and edit the information, and efficiently outputs automatically generated reports.

[0067] "Visual information" refers to dynamic or static visual data acquired through cameras or other acquisition devices.

[0068] A "storage device" is a piece of hardware or digital platform capable of storing data for extended periods.

[0069] "Analysis" refers to the process of examining data and extracting specific patterns or information.

[0070] An "event" refers to a specific event or anomaly that occurs within video information.

[0071] "Natural language" refers to the forms of expression that humans use in everyday life, and is used to describe data in a way that is easy to decipher.

[0072] "Indexing" is the process of assigning specific markers or tags to information to make it easier to search and refer to.

[0073] A "user" is a person or organization that has the authority to operate the system and to view and edit the results.

[0074] An "operation screen" is a display and input area for a user to interface with the system.

[0075] A "report" is a formatted document that compiles analyzed information and user comments.

[0076] "Digital document format" refers to a file format for documents that can be stored and shared electronically.

[0077] An "object recognition algorithm" is a computational method for detecting specific objects or patterns within video information.

[0078] A "sharing platform" is a system for sharing digital data online with multiple users or organizations.

[0079] This invention is a system that efficiently analyzes video information, verbalizes and indexes specific events, thereby reducing the burden on supervisors and automating report creation. An example of the implementation of this system is as follows:

[0080] First, the user uploads the work video from their device to the server. The device needs to have software capable of processing common video formats (e.g., MP4, AVI) and an internet connection. When the server receives this video information, it immediately saves it to its storage device. For storage devices, SSDs or HDDs are used to stably hold large amounts of data.

[0081] Next, the server analyzes the stored video information. This analysis requires computing resources to run AI algorithms, and object recognition algorithms such as YOLO (You Only Look Once) are used. The server inspects the video frame by frame, detects specific events, and records timestamps.

[0082] The server translates detected events into natural language. This process utilizes natural language processing (NLP) techniques to transform the data into a human-readable format. If necessary, speech synthesis technologies such as Google® Text-to-Speech can be used to convert the generated text into speech.

[0083] Subsequently, the server indexes the relevant sections of the video based on the transcribed information, allowing users to easily search and play them. The index information, along with timestamps, is recorded and can be viewed and edited by users through the user interface. The interface also allows users to access event information and add comments.

[0084] Finally, the server automatically generates a report based on the edited information. The report is output in digital document formats such as PDF and HTML and provided on the company's shared platform. This allows all stakeholders to quickly share information and address issues efficiently.

[0085] One concrete example involves using video footage of a manufacturing line. When a user uploads this video to the system, an AI algorithm automatically detects errors in handling parts and generates linguistic data such as "Incorrect placement of part A." The user can then review this information on the operation screen and add comments as needed. An example of a prompt to the generating AI model in this case might be, "Can you detect and index errors in handling parts from a video of a manufacturing line?"

[0086] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0087] Step 1:

[0088] The user uploads the work video from their terminal to the server. The input is the video file stored on the terminal, and the output is a notification that the file transfer to the server is complete. At this stage, the user operates a dedicated upload screen, selects the target video file, and presses the transfer button to start the process.

[0089] Step 2:

[0090] The server stores the received video information in its storage device. The input is the moved video file, and the output is the video metadata (file format, size, resolution, etc.). This metadata is stored in a database for data management.

[0091] Step 3:

[0092] The server analyzes the stored video information. The input is a video file stored in a memory device, and the output is timestamp information related to a specific event. The server uses an object recognition algorithm (e.g., YOLO) to process the data frame by frame and detect events.

[0093] Step 4:

[0094] The server converts detected events into natural language. The input is timestamp information about the events, and the output is text data expressed in natural language. The server uses natural language processing (NLP) to generate the text. If necessary, the server may also consider using speech synthesis technology to convert the generated text into speech.

[0095] Step 5:

[0096] The server indexes the relevant parts of the video based on the transcribed information. The input consists of transcribed text and timestamp information, and the output is indexed data. This allows the user to view detailed event information using the interface.

[0097] Step 6:

[0098] Users view and edit indexed information through the user interface. Input is indexed data provided by the server, and output includes user comments and modifications. Users can add comments or modify content for events displayed on the screen.

[0099] Step 7:

[0100] The server automatically generates reports using the edited information. The input is user-edited data, and the output is a report in PDF or HTML format. This report is shared among stakeholders to help resolve issues quickly.

[0101] (Application Example 1)

[0102] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0103] In factories and workplaces, there is a need to quickly detect equipment malfunctions and quality defects, enabling immediate responses to workers, thereby improving production efficiency and reducing the burden on supervisors. However, conventional methods have the problem of requiring a great deal of time and effort to detect and report abnormalities.

[0104] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0105] In this invention, the server includes means for receiving video data and storing it in a data storage device, means for analyzing the stored video data and detecting specific events within the video, and means for converting the detected events into natural language. This makes it possible to detect abnormal operations in factories and work sites in real time and notify workers by voice.

[0106] "Video data" refers to digital information that records visual information acquired by a camera or camera.

[0107] A "data storage device" is an electronic device for securely storing video data over long periods of time, and it functions as part of a server.

[0108] An "event" refers to a specific action or change in state that is detected within the video data.

[0109] "Natural language" refers to the language that humans use on a daily basis and is used to express analyzed information.

[0110] "Indexing" is the process of assigning identifiers to digital data in order to facilitate the retrieval of information.

[0111] An "information display device" is a device that provides an interface for users to view and edit digital data and event content.

[0112] "Voice notification" is a means of conveying information to the user through voice about a specific event that has been detected.

[0113] An "object recognition algorithm" is a computational method for automatically identifying objects and patterns in video data.

[0114] "Electronic format" is a general term for file formats that can be displayed on a computer as digital information.

[0115] The embodiments for carrying out this invention are as follows.

[0116] The server first stores the video data transmitted by the user in a data storage device. During this storage process, the server adds metadata to streamline data management. Next, the server analyzes the stored video data. In this analysis, an AI model is used to detect various events within the video and convert them into natural language. In this process, an object recognition algorithm is used to accurately recognize and verbalize specific events.

[0117] The information, translated into natural language, is indexed by the server and managed in a user-friendly format. Users can view and edit this data using the provided information display device. Particularly important is the server's ability to communicate any abnormal activity detected in real time to the worker via voice notification, facilitating prompt action. Furthermore, based on the edited data, the server can automatically generate documents and export them electronically.

[0118] A concrete example is monitoring machine operation on a factory production line. If an operator is wearing smart glasses, they will receive an immediate voice notification if there is any abnormal movement, enabling early detection and rapid response to problems.

[0119] Examples of prompts to input into a generative AI model:

[0120] "The AI ​​will detect abnormal machine operation in the factory in real time and notify workers. Please propose specific methods to improve the efficiency of this process."

[0121] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0122] Step 1:

[0123] The server receives video data sent by the user and stores it in the data storage device. During this process, the raw data, as input, is stored along with metadata for data identification. The output is the storage location within the storage device associated with the metadata. Specifically, the server processes the file upload and creates a directory for data storage.

[0124] Step 2:

[0125] The server sequentially reads the stored video data and analyzes it using an object recognition AI model. The input is the stored video data, and the output is digital data that identifies specific events contained within it. The AI ​​model evaluates each frame and identifies specific actions or states. Specifically, the AI ​​generates data about the position and movement of objects and detects anomalies.

[0126] Step 3:

[0127] The server converts events detected by the AI ​​into natural language. The input is digital data representing the results of the event analysis, and the output is text in natural language that is understandable to humans. A natural language processing algorithm is used for this conversion, and the result is generated as a text file.

[0128] Step 4:

[0129] The server indexes events based on the converted natural language text. The input is natural language text, and the output is index information organized for easy searching and access. Specific operations include adding timestamps to the text and creating indexes using specific keywords.

[0130] Step 5:

[0131] The server transmits the indexed information to an information display device available to the user. The user uses this device to review the content and edit it as needed. The input is indexed information, and the output is display data that can be viewed and edited. Specifically, it provides editing and commenting functions to the user interface.

[0132] Step 6:

[0133] The server automatically generates documents based on edited information and exports them in electronic format. The input is edited index information, and the output is an electronic document such as PDF or HTML. Specifically, it uses a document generation engine to format the document and save it in the selected file format.

[0134] Step 7:

[0135] The server notifies workers of abnormal events detected in real time using speech synthesis technology. The input is real-time abnormality detection information, and the output is a notification in the form of a voice message. Specifically, the server generates synthesized speech and transmits it to smart glasses via the network.

[0136] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0137] This invention combines an emotion engine with a video analysis system to reflect the user's emotional state within the system, thereby improving the analysis and reporting process. This enhances the user experience and provides deeper insights.

[0138] First, the user uploads the work video to the server via their terminal. During this process, the system authenticates the user and ensures data integrity. The video file is recorded in storage by the server, and associated metadata is also saved simultaneously.

[0139] Next, the server analyzes the video data stored in storage frame by frame. Using AI algorithms, it inspects each frame to detect events within the video. During this process, object detection technology identifies specific actions or anomalies, and assigns corresponding timestamps.

[0140] The detected events are converted into natural language on the server. The conversion process includes rule-based sentence generation, and the output text is in a format that is easily understandable to humans. Furthermore, the events in the video are indexed based on this.

[0141] Here, a unique emotion engine is incorporated into the present invention. When a user uses the interface to view events, emotion data is collected through the camera and microphone. The emotion engine analyzes this data to identify the user's emotional state. This emotion information is used to optimize the interface's display content and feedback functions, providing the user with the optimal operating environment.

[0142] The server also associates detected events with the user's emotional state and incorporates this into the report. The report becomes more visual and detailed as it reflects the user's reactions and opinions as a result of the sentiment analysis.

[0143] As a concrete example, consider a quality control work video. When a user inputs video into the system, AI automatically detects specific events, such as defective products. Meanwhile, an emotion engine analyzes the user's stress and level of comfort, providing a pleasant UI / UX. Finally, a report is created based on this information, serving as a guideline for quality improvement.

[0144] The following describes the processing flow.

[0145] Step 1:

[0146] Users upload work video data to the server using their terminals. During this process, the system automatically authenticates the user and verifies access rights. Data uploads are only permitted if the user is legitimate.

[0147] Step 2:

[0148] The server stores the received video data in storage and simultaneously records related metadata (such as the date and time of shooting and camera location information). This information is used for reference in subsequent processes.

[0149] Step 3:

[0150] The server divides the stored video data into individual frames and prepares each frame for analysis by sending it to an AI algorithm. This division improves the accuracy of the analysis and enables the rapid detection of specific events.

[0151] Step 4:

[0152] The server uses AI algorithms to detect specific events and anomalies in each frame. Object detection technology is employed, and relevant events are timestamped. This process identifies events of interest within the video.

[0153] Step 5:

[0154] The server converts detected events into natural language and creates an index. The converted data is presented in an easy-to-read text format, and by simultaneously indexing which part of the video it corresponds to, consistent management becomes possible.

[0155] Step 6:

[0156] Users can view and edit indexed events through an interface via their device. Here, user sentiment data is collected using the camera and microphone, and a sentiment engine analyzes this data in real time.

[0157] Step 7:

[0158] The server optimizes the interface display and feedback by considering the user's emotional state, as analyzed by the emotion engine. This adjustment aims to improve the user experience and provide an efficient operating environment.

[0159] Step 8:

[0160] The server automatically generates a detailed work report by combining the results of editing and sentiment analysis. The report includes an event summary, index information, and analysis results based on user sentiment.

[0161] Step 9:

[0162] The server exports the generated report in PDF or HTML format and distributes it to the designated stakeholders. This format is easily integrated with other digital platforms, facilitating smooth report sharing.

[0163] (Example 2)

[0164] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0165] Conventional video analysis systems had the functionality to detect specific situations from video data and automatically generate report documents, but they did not optimize the interface or feedback to reflect the user's emotional state. As a result, the user experience was not sufficiently improved, and the insights provided by the system were limited. Furthermore, because it was not possible to create report documents that integrated the situation and the user's emotions, there was a lack of visual and detailed reporting.

[0166] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0167] In this invention, the server includes means for receiving video data and storing it in a storage device, means for analyzing the video data and detecting specific situations, means for converting the detected situations into natural language, means for acquiring and analyzing the user's emotional state, means for optimizing the interface and feedback based on the emotional state, and means for integrating the detected situations and emotional states and reflecting them in a report document. This enables deeper insights by reflecting the user's emotional state in the video analysis, and realizes the automatic generation of visually detailed report documents.

[0168] "Video data" refers to visual information acquired by cameras or other recording devices and stored in digital format.

[0169] A "storage device" is a device used to store data for a long or short period of time, and includes hard disk drives and solid-state drives.

[0170] "Analysis" is the process of extracting useful information from data, and it is carried out using algorithms and machine learning techniques.

[0171] "Situation" refers to a specific event or action that occurs within the video, and video analysis technology is used to detect it.

[0172] "Natural language" refers to the language that humans normally use, and it is sometimes applied to documents and texts generated by machines.

[0173] "Indexing" is the process of assigning keys and tags to data to make it easier to manage, facilitating searching and organization.

[0174] "User's emotional state" refers to the psychological or emotional response a user shows to a particular situation, and is often obtained through facial recognition or voice analysis.

[0175] "Integration" is the process of combining multiple data and pieces of information into a single system or report, enabling a more comprehensive understanding.

[0176] A "report document" is a document format that compiles analyzed data and information, and is generated according to a predetermined format.

[0177] This invention is a system that combines video analysis technology and emotion analysis technology, and generates a detailed report document using video data provided by the user. A specific embodiment of the system is described below.

[0178] The server receives video data and stores it in storage. High-speed, high-capacity storage is desirable for this process. Next, the server analyzes the video data frame by frame using AI algorithms such as TENSORFLOW® and OpenCV. When a specific situation is detected, object detection algorithms are applied to identify anomalies or unusual behavior and pinpoint the relevant area.

[0179] The device provides an interface for the user to review the analyzed situation and provide feedback. Emotional data, such as facial expressions and voice tone, is captured through the user's camera and microphone. The server uses an emotion engine to analyze this data and identify the user's emotional state. Based on this information, the interface and feedback functions are optimized.

[0180] Ultimately, the server integrates the detected situation and the user's emotional state, automatically generating a detailed, visually easy-to-understand report. This report can be exported in electronic file format or markup language format.

[0181] As a concrete example, consider quality control videos in the manufacturing industry. When a user inputs video footage into the system, the AI ​​automatically analyzes specific situations, such as detecting defective products. Simultaneously, an emotion engine assesses the user's stress and sense of security, and reflects this in the report. This allows for deeper insights and improvement measures to be provided.

[0182] An example of a prompt message when using a generative AI model is, "Perform frame analysis on the uploaded video and create a detailed report including user sentiment analysis."

[0183] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0184] Step 1:

[0185] The user uploads video data to the server via their device. The input consists of the user's file selection and authentication information, which the server uses to verify data integrity and save it to storage. This step prepares the video data and associated metadata (e.g., timestamp, user ID) in storage.

[0186] Step 2:

[0187] The server analyzes video data stored in storage frame by frame. The input is video data stored in memory, and AI algorithms (e.g., TensorFlow, OpenCV) are used to analyze each frame. This detects specific events or actions within the video, and outputs timestamps indicating the situation.

[0188] Step 3:

[0189] The server translates the detected situation into natural language. The input is the timestamped situation data detected in step 2, which is then converted into an easy-to-read text format using a natural language generation engine. Through this step, the output is text in a format that is easily understandable to humans.

[0190] Step 4:

[0191] Users access the interface through their device to view the analyzed situation. The input is text information converted into natural language, which is used to optimize the user interface for feedback. The user's camera and microphone are activated to collect emotional data, which is then analyzed by the server.

[0192] Step 5:

[0193] The server analyzes the collected emotional data using an emotion engine to identify the emotional state. The input is emotional data obtained from the user, and based on this, it outputs metrics to optimize the interface display content and feedback functions.

[0194] Step 6:

[0195] The server integrates detected situations and user emotional states to automatically generate a report. Inputs include naturally translated situational data and sentiment analysis results, and a visually detailed report is exported as an electronic file. This report provides meaningful insights for users and other stakeholders.

[0196] (Application Example 2)

[0197] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0198] Conventional video analysis systems detect and report specific events by analyzing video data, but they struggled to optimize the user experience and enrich reports without considering the user's emotional state. Furthermore, in quality control in factory settings, optimizing the environment while considering the emotions of workers is required, but solutions to this problem were lacking. Moreover, the generated reports were merely lists of facts and did not reflect the emotions or reactions of workers.

[0199] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0200] In this invention, the server includes means for receiving video data and storing it in an information storage device, means for analyzing the stored video data and detecting specific events within the video, means for converting the detected events into natural language, and emotion analysis means for collecting and analyzing user emotion data. This makes it possible to optimize the user interface considering the user's emotional state and to create reports that reflect the emotion analysis results.

[0201] "Video data" refers to a collection of visual information recorded using a camera or other recording device.

[0202] An "information storage device" is hardware or media used to store data temporarily or long-term.

[0203] "Means of analysis" refers to techniques or methods for processing given information and converting its content into a format that is easy to understand.

[0204] A "specific event" refers to a particular action or state identified within the video footage.

[0205] "Natural language" refers to the language that humans use on a daily basis, and is the form it takes before being converted into a language that can be interpreted by machines.

[0206] "Emotional analysis means" refers to technologies or methods for digitizing and analyzing human emotional states.

[0207] A "user interface" is a collection of visual and functional elements that allow a system and a user to interact with each other.

[0208] A "report" is a document generated by the system that includes analysis results and other supplementary information.

[0209] "Optimization" is the act of improving a system or process to achieve a specific objective, thereby increasing its efficiency and effectiveness.

[0210] To implement this invention, coordination between robots and servers in a factory monitoring system is crucial. The server first receives video data from cameras and other sources and stores it in an information storage device. The stored data is analyzed, and specific events such as defective products on the production line are detected using AI. Software such as TensorFlow and OpenCV are used for this purpose.

[0211] Next, the detected events are converted into natural language. This indexes the events in a human-readable format, making them easier to review and edit later.

[0212] This is where a unique emotion analysis method is incorporated. The robot collects emotional data from nearby workers through cameras and microphones and analyzes their emotional state using Azure® Emotion API and Google Cloud Vision. This allows for the quantification of workers' stress levels and sense of security.

[0213] As users view events through the interface, this sentiment data is analyzed, and the user interface is optimized for optimal user experience. In this way, the server integrates the results of the sentiment analysis into a report, generating a document with more visual and detailed information.

[0214] As a concrete example, consider a factory with a new product manufacturing line. Factory robots perform video analysis to automatically detect minute defects in the products, while simultaneously adjusting the UI if nearby workers are experiencing stress. As a result, problem identification and optimization of the work environment are achieved simultaneously. An example of a prompt might be: "Explain how to detect defective products from manufacturing line video data and optimize the UI / UX based on the results of worker sentiment analysis."

[0215] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0216] Step 1:

[0217] The server receives video data from the factory cameras. The input is a real-time video stream captured by the cameras, which is taken into the server as data. The server saves this video to its information storage device. The output is the video data stored in the storage device.

[0218] Step 2:

[0219] The server analyzes the stored video data using an AI algorithm. The input is the video data saved in step 1. Using an AI algorithm (such as TensorFlow or OpenCV), object detection and event identification are performed on a frame-by-frame basis, and the output includes a list of detected events.

[0220] Step 3:

[0221] The server converts detected events into natural language. The input is the events detected in step 2. In this process, natural language generation technology is used to convert the extracted events into text format. The output is text data of the events as language.

[0222] Step 4:

[0223] The server indexes events within the video in a user-accessible format based on the verbalized event data. The input is the text data from step 3, and the indexing program generates an efficiently searchable index as needed. The output is event data with index information added.

[0224] Step 5:

[0225] As a user works around the robot, the robot collects emotional data from the user through its camera and microphone. The input consists of real-time video and audio data. Based on this, an emotional analysis algorithm (such as Azure Emotion API or Google Cloud Vision) is used to analyze the user's emotional state. The output is analytical data indicating the user's emotional state.

[0226] Step 6:

[0227] The server applies the emotion analysis results to provide a user-optimized interface. The input is the emotion state data from step 5. Based on the emotion data, the UI / UX components are dynamically adjusted to provide an interface that is easy for the user to use. The output is the optimized user interface.

[0228] Step 7:

[0229] The server automatically generates a report based on the combined data of detected events and sentiment analysis results. The input is the data from steps 4 and 5. Using the report generation program, a document conforming to the report format is created, and the output is compiled as a detailed, visual report.

[0230] Step 8:

[0231] The server exports the generated report in PDF or HTML format. The input is the report data generated in step 7. A format conversion program is used to convert the report into an electronic document format, which is then saved or sent as needed. The output is a document file in the specified format.

[0232] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0233] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0234] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0235] [Second Embodiment]

[0236] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0237] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0238] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0239] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0240] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0241] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0242] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0243] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0244] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0245] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0246] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0247] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0248] This invention is a system that efficiently analyzes video data, verbalizes and indexes specific events, thereby reducing the burden on supervisors and automating report creation. The system operates as follows:

[0249] First, the user uploads a video of their work from their device to the server. The uploaded video is automatically saved to storage by the server. The storage system maintains metadata about the data and is managed to ensure that subsequent processing proceeds smoothly.

[0250] Next, the server sequentially analyzes the stored video data. The analysis process examines the video frame by frame, using an AI algorithm to detect specific events and anomalies within the video. This AI algorithm utilizes object detection technology and records event timestamps.

[0251] Detected events are automatically translated into natural language by the server. This translation makes the data more human-readable. In addition, if necessary, speech synthesis technology is used to convert the data into speech.

[0252] Subsequently, the server indexes the relevant portion of the video based on the verbalized event information. This indexing process makes it possible to easily search for and play specific events.

[0253] The completed index data is provided as an interface for users to review and edit. Through this interface, users can view event details and, if necessary, modify descriptions or add additional notes.

[0254] Finally, the server automatically generates a report based on the edited data. The report can be exported in formats such as PDF and HTML and shared on other digital platforms. This allows all stakeholders to efficiently share information and address issues quickly.

[0255] As a concrete example, consider a video taken on a manufacturing line. When a user uploads this video to the system, the AI ​​automatically detects errors in handling parts and generates a text index such as "Incorrect placement of part A." The user can review this information through the interface and add further comments. When creating a report, a detailed report containing these events is automatically generated and shared with the relevant departments.

[0256] The following describes the processing flow.

[0257] Step 1:

[0258] The user uploads work video data from their terminal to the server. At this time, the system performs user authentication to confirm that the user has legitimate access rights.

[0259] Step 2:

[0260] The server saves the received video data to storage. During saving, it analyzes metadata related to the video (such as the date and time of shooting and camera location information) and registers it in the database.

[0261] Step 3:

[0262] The server divides the stored video data into frames and prepares them for analysis. This division is based on the video's frame rate and optimized for subsequent analysis.

[0263] Step 4:

[0264] The server uses an AI model to process each frame and detect specific events or anomalies within the video. For example, it might use an object detection algorithm to track the movement of an object of interest.

[0265] Step 5:

[0266] The server converts detected events into natural language text. Here, predefined rules and contextual analysis are used to express the information in a language that is easily understandable to humans.

[0267] Step 6:

[0268] The server indexes each event location in the video based on the verbalized event information. This process records detailed information, including the time and location of the event.

[0269] Step 7:

[0270] Users can access the server through their terminal and view an indexed list of events from the interface. Users can then view the list, select a specific event, and edit or add details.

[0271] Step 8:

[0272] The server automatically generates reports based on information edited by the user. The report's structure is determined, and event information is incorporated in a predetermined format.

[0273] Step 9:

[0274] The server exports the final report in PDF or HTML format and distributes it to the designated stakeholders. This output is then integrated into the company's workflow as needed.

[0275] (Example 1)

[0276] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0277] This invention aims to automate the process of analyzing and reporting video information, thereby reducing the burden on supervisors and enabling efficient information sharing. Conventional manual video analysis and report creation are time-consuming, labor-intensive, and prone to human error. Therefore, this invention provides a system to address these issues.

[0278] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0279] In this invention, the server includes means for receiving video information and storing it in a storage device, means for analyzing the stored video information to detect specific events, and means for converting the detected events into natural language. As a result, events in the video can be detected quickly and accurately, allowing users to easily view and edit information, and enabling the efficient output of automatically generated reports.

[0280] "Video information" refers to dynamic or static visual data acquired through a camera or other acquisition device.

[0281] "Storage device" refers to hardware or a digital platform capable of storing data over a long period of time.

[0282] "Analysis" refers to the process of examining data and extracting specific patterns or information.

[0283] "Event" refers to a specific event or anomaly occurring in the video information.

[0284] "Natural language" refers to the form of expression as a language that humans use in daily life, and is used to describe data in an easily interpretable form.

[0285] "Indexing" refers to the process of attaching specific markers or tags to information to facilitate searching and referencing.

[0286] "User" refers to a person or group having the authority to operate the system and view and edit the results.

[0287] "Operation screen" refers to the display and input area for the user to interface with the system.

[0288] "Report" refers to a formatted document that summarizes the analyzed information and the user's comments.

[0289] "Digital document format" refers to a file format for documents that can be stored and shared electronically.

[0290] An "object recognition algorithm" is a computational method for detecting specific objects or patterns within video information.

[0291] A "sharing platform" is a system for sharing digital data online with multiple users or organizations.

[0292] This invention is a system that efficiently analyzes video information, verbalizes and indexes specific events, thereby reducing the burden on supervisors and automating report creation. An example of the implementation of this system is as follows:

[0293] First, the user uploads the work video from their device to the server. The device needs to have software capable of processing common video formats (e.g., MP4, AVI) and an internet connection. When the server receives this video information, it immediately saves it to its storage device. For storage devices, SSDs or HDDs are used to stably hold large amounts of data.

[0294] Next, the server analyzes the stored video information. This analysis requires computing resources to run AI algorithms, and object recognition algorithms such as YOLO (You Only Look Once) are used. The server inspects the video frame by frame, detects specific events, and records timestamps.

[0295] The server translates detected events into natural language. This process utilizes natural language processing (NLP) techniques to transform the data into a human-readable format. If necessary, speech synthesis technologies such as Google Text-to-Speech can be used to convert the generated text into speech.

[0296] Subsequently, the server indexes the relevant sections of the video based on the transcribed information, allowing users to easily search and play them. The index information, along with timestamps, is recorded and can be viewed and edited by users through the user interface. The interface also allows users to access event information and add comments.

[0297] Finally, the server automatically generates a report based on the edited information. The report is output in digital document formats such as PDF and HTML and provided on the company's shared platform. This allows all stakeholders to quickly share information and address issues efficiently.

[0298] One concrete example involves using video footage of a manufacturing line. When a user uploads this video to the system, an AI algorithm automatically detects errors in handling parts and generates linguistic data such as "Incorrect placement of part A." The user can then review this information on the operation screen and add comments as needed. An example of a prompt to the generating AI model in this case might be, "Can you detect and index errors in handling parts from a video of a manufacturing line?"

[0299] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0300] Step 1:

[0301] The user uploads the work video from their terminal to the server. The input is the video file stored on the terminal, and the output is a notification that the file transfer to the server is complete. At this stage, the user operates a dedicated upload screen, selects the target video file, and presses the transfer button to start the process.

[0302] Step 2:

[0303] The server stores the received video information in a storage device. The input is the moved video file, and as output, it stores the meta information of the video (file format, size, resolution, etc.). This meta information is stored in a database for data management.

[0304] Step 3:

[0305] The server analyzes the stored video information. The input is the video file stored in the storage device, and as output, timestamp information regarding a specific event is obtained. The server uses an object recognition algorithm (e.g., YOLO), processes the data frame by frame, and detects events.

[0306] Step 4:

[0307] The server converts the detected event into natural language. The input is the timestamp information regarding the event, and as output, text data expressed in natural language is obtained. The server utilizes natural language processing (NLP) to generate text. At this time, it is also considered to vocalize the generated text using speech synthesis technology as necessary.

[0308] Step 5:

[0309] Based on the verbalized information, the server indexes the corresponding parts within the video. The input is the text in natural language and the timestamp information, and as output, indexed data is obtained. This enables the user to confirm the details of the event using the interface.

[0310] Step 6:

[0311] The user checks and edits the indexed information through the operation screen. The input is the indexed data provided by the server, and as output, the user's comments and correction information are generated. The user can add necessary comments or modify the content for the events displayed on the screen.

[0312] Step 7:

[0313] The server automatically generates reports using the edited information. The input is user-edited data, and the output is a report in PDF or HTML format. This report is shared among stakeholders to help resolve issues quickly.

[0314] (Application Example 1)

[0315] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0316] In factories and workplaces, there is a need to quickly detect equipment malfunctions and quality defects, enabling immediate responses to workers, thereby improving production efficiency and reducing the burden on supervisors. However, conventional methods have the problem of requiring a great deal of time and effort to detect and report abnormalities.

[0317] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0318] In this invention, the server includes means for receiving video data and storing it in a data storage device, means for analyzing the stored video data and detecting specific events within the video, and means for converting the detected events into natural language. This makes it possible to detect abnormal operations in factories and work sites in real time and notify workers by voice.

[0319] "Video data" refers to digital information that records visual information acquired by a camera or camera.

[0320] A "data storage device" is an electronic device for securely storing video data over long periods of time, and it functions as part of a server.

[0321] An "event" refers to a specific action or change in state that is detected within the video data.

[0322] "Natural language" refers to the language that humans use on a daily basis and is used to express analyzed information.

[0323] "Indexing" is the process of assigning identifiers to digital data in order to facilitate the retrieval of information.

[0324] An "information display device" is a device that provides an interface for users to view and edit digital data and event content.

[0325] "Voice notification" is a means of conveying information to the user through voice about a specific event that has been detected.

[0326] An "object recognition algorithm" is a computational method for automatically identifying objects and patterns in video data.

[0327] "Electronic format" is a general term for file formats that can be displayed on a computer as digital information.

[0328] The embodiments for carrying out this invention are as follows.

[0329] The server first stores the video data transmitted by the user in a data storage device. During this storage process, the server adds metadata to streamline data management. Next, the server analyzes the stored video data. In this analysis, an AI model is used to detect various events within the video and convert them into natural language. In this process, an object recognition algorithm is used to accurately recognize and verbalize specific events.

[0330] The information, translated into natural language, is indexed by the server and managed in a user-friendly format. Users can view and edit this data using the provided information display device. Particularly important is the server's ability to communicate any abnormal activity detected in real time to the worker via voice notification, facilitating prompt action. Furthermore, based on the edited data, the server can automatically generate documents and export them electronically.

[0331] A concrete example is monitoring machine operation on a factory production line. If an operator is wearing smart glasses, they will receive an immediate voice notification if there is any abnormal movement, enabling early detection and rapid response to problems.

[0332] Examples of prompts to input into a generative AI model:

[0333] "The AI ​​will detect abnormal machine operation in the factory in real time and notify workers. Please propose specific methods to improve the efficiency of this process."

[0334] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0335] Step 1:

[0336] The server receives video data sent by the user and stores it in the data storage device. During this process, the raw data, as input, is stored along with metadata for data identification. The output is the storage location within the storage device associated with the metadata. Specifically, the server processes the file upload and creates a directory for data storage.

[0337] Step 2:

[0338] The server sequentially reads the stored video data and analyzes it using an object recognition AI model. The input is the stored video data, and the output is digital data that identifies specific events contained within it. The AI ​​model evaluates each frame and identifies specific actions or states. Specifically, the AI ​​generates data about the position and movement of objects and detects anomalies.

[0339] Step 3:

[0340] The server converts events detected by the AI ​​into natural language. The input is digital data representing the results of the event analysis, and the output is text in natural language that is understandable to humans. A natural language processing algorithm is used for this conversion, and the result is generated as a text file.

[0341] Step 4:

[0342] The server indexes events based on the converted natural language text. The input is natural language text, and the output is index information organized for easy searching and access. Specific operations include adding timestamps to the text and creating indexes using specific keywords.

[0343] Step 5:

[0344] The server transmits the indexed information to an information display device available to the user. The user uses this device to review the content and edit it as needed. The input is indexed information, and the output is display data that can be viewed and edited. Specifically, it provides editing and commenting functions to the user interface.

[0345] Step 6:

[0346] The server automatically generates documents based on edited information and exports them in electronic format. The input is edited index information, and the output is an electronic document such as PDF or HTML. Specifically, it uses a document generation engine to format the document and save it in the selected file format.

[0347] Step 7:

[0348] The server notifies workers of abnormal events detected in real time using speech synthesis technology. The input is real-time abnormality detection information, and the output is a notification in the form of a voice message. Specifically, the server generates synthesized speech and transmits it to smart glasses via the network.

[0349] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0350] This invention combines an emotion engine with a video analysis system to reflect the user's emotional state within the system, thereby improving the analysis and reporting process. This enhances the user experience and provides deeper insights.

[0351] First, the user uploads the work video to the server via their terminal. During this process, the system authenticates the user and ensures data integrity. The video file is recorded in storage by the server, and associated metadata is also saved simultaneously.

[0352] Next, the server analyzes the video data stored in storage frame by frame. Using AI algorithms, it inspects each frame to detect events within the video. During this process, object detection technology identifies specific actions or anomalies, and assigns corresponding timestamps.

[0353] The detected events are converted into natural language on the server. The conversion process includes rule-based sentence generation, and the output text is in a format that is easily understandable to humans. Furthermore, the events in the video are indexed based on this.

[0354] Here, a unique emotion engine is incorporated into the present invention. When a user uses the interface to view events, emotion data is collected through the camera and microphone. The emotion engine analyzes this data to identify the user's emotional state. This emotion information is used to optimize the interface's display content and feedback functions, providing the user with the optimal operating environment.

[0355] The server also associates detected events with the user's emotional state and incorporates this into the report. The report becomes more visual and detailed as it reflects the user's reactions and opinions as a result of the sentiment analysis.

[0356] As a concrete example, consider a quality control work video. When a user inputs video into the system, AI automatically detects specific events, such as defective products. Meanwhile, an emotion engine analyzes the user's stress and level of comfort, providing a pleasant UI / UX. Finally, a report is created based on this information, serving as a guideline for quality improvement.

[0357] The following describes the processing flow.

[0358] Step 1:

[0359] Users upload work video data to the server using their terminals. During this process, the system automatically authenticates the user and verifies access rights. Data uploads are only permitted if the user is legitimate.

[0360] Step 2:

[0361] The server stores the received video data in storage and simultaneously records related metadata (such as the date and time of shooting and camera location information). This information is used for reference in subsequent processes.

[0362] Step 3:

[0363] The server divides the stored video data into individual frames and prepares each frame for analysis by sending it to an AI algorithm. This division improves the accuracy of the analysis and enables the rapid detection of specific events.

[0364] Step 4:

[0365] The server uses AI algorithms to detect specific events and anomalies in each frame. Object detection technology is employed, and relevant events are timestamped. This process identifies events of interest within the video.

[0366] Step 5:

[0367] The server converts detected events into natural language and creates an index. The converted data is presented in an easy-to-read text format, and by simultaneously indexing which part of the video it corresponds to, consistent management becomes possible.

[0368] Step 6:

[0369] Users can view and edit indexed events through an interface via their device. Here, user sentiment data is collected using the camera and microphone, and a sentiment engine analyzes this data in real time.

[0370] Step 7:

[0371] The server optimizes the interface display and feedback by considering the user's emotional state, as analyzed by the emotion engine. This adjustment aims to improve the user experience and provide an efficient operating environment.

[0372] Step 8:

[0373] The server automatically generates a detailed work report by combining the results of editing and sentiment analysis. The report includes an event summary, index information, and analysis results based on user sentiment.

[0374] Step 9:

[0375] The server exports the generated report in PDF or HTML format and distributes it to the designated stakeholders. This format is easily integrated with other digital platforms, facilitating smooth report sharing.

[0376] (Example 2)

[0377] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0378] Conventional video analysis systems had the functionality to detect specific situations from video data and automatically generate report documents, but they did not optimize the interface or feedback to reflect the user's emotional state. As a result, the user experience was not sufficiently improved, and the insights provided by the system were limited. Furthermore, because it was not possible to create report documents that integrated the situation and the user's emotions, there was a lack of visual and detailed reporting.

[0379] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0380] In this invention, the server includes means for receiving video data and storing it in a storage device, means for analyzing the video data and detecting specific situations, means for converting the detected situations into natural language, means for acquiring and analyzing the user's emotional state, means for optimizing the interface and feedback based on the emotional state, and means for integrating the detected situations and emotional states and reflecting them in a report document. This enables deeper insights by reflecting the user's emotional state in the video analysis, and realizes the automatic generation of visually detailed report documents.

[0381] "Video data" refers to visual information acquired by cameras or other recording devices and stored in digital format.

[0382] A "storage device" is a device used to store data for a long or short period of time, and includes hard disk drives and solid-state drives.

[0383] "Analysis" is the process of extracting useful information from data, and it is carried out using algorithms and machine learning techniques.

[0384] "Situation" refers to a specific event or action that occurs within the video, and video analysis technology is used to detect it.

[0385] "Natural language" refers to the language that humans normally use, and it is sometimes applied to documents and texts generated by machines.

[0386] "Indexing" is the process of assigning keys and tags to data to make it easier to manage, facilitating searching and organization.

[0387] "User's emotional state" refers to the psychological or emotional response a user shows to a particular situation, and is often obtained through facial recognition or voice analysis.

[0388] "Integration" is the process of combining multiple data and pieces of information into a single system or report, enabling a more comprehensive understanding.

[0389] A "report document" is a document format that compiles analyzed data and information, and is generated according to a predetermined format.

[0390] This invention is a system that combines video analysis technology and emotion analysis technology, and generates a detailed report document using video data provided by the user. A specific embodiment of the system is described below.

[0391] The server receives video data and stores it in storage. High-speed, high-capacity storage is desirable for this process. Next, the server uses AI algorithms such as TensorFlow and OpenCV to analyze the video data frame by frame. When a specific situation is detected, object detection algorithms are applied to identify anomalies or unusual behavior and pinpoint the relevant parts.

[0392] The device provides an interface for the user to review the analyzed situation and provide feedback. Emotional data, such as facial expressions and voice tone, is captured through the user's camera and microphone. The server uses an emotion engine to analyze this data and identify the user's emotional state. Based on this information, the interface and feedback functions are optimized.

[0393] Ultimately, the server integrates the detected situation and the user's emotional state, automatically generating a detailed, visually easy-to-understand report. This report can be exported in electronic file format or markup language format.

[0394] As a concrete example, consider quality control videos in the manufacturing industry. When a user inputs video footage into the system, the AI ​​automatically analyzes specific situations, such as detecting defective products. Simultaneously, an emotion engine assesses the user's stress and sense of security, and reflects this in the report. This allows for deeper insights and improvement measures to be provided.

[0395] An example of a prompt message when using a generative AI model is, "Perform frame analysis on the uploaded video and create a detailed report including user sentiment analysis."

[0396] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0397] Step 1:

[0398] The user uploads video data to the server via their device. The input consists of the user's file selection and authentication information, which the server uses to verify data integrity and save it to storage. This step prepares the video data and associated metadata (e.g., timestamp, user ID) in storage.

[0399] Step 2:

[0400] The server analyzes video data stored in storage frame by frame. The input is video data stored in memory, and AI algorithms (e.g., TensorFlow, OpenCV) are used to analyze each frame. This detects specific events or actions within the video, and outputs timestamps indicating the situation.

[0401] Step 3:

[0402] The server translates the detected situation into natural language. The input is the timestamped situation data detected in step 2, which is then converted into an easy-to-read text format using a natural language generation engine. Through this step, the output is text in a format that is easily understandable to humans.

[0403] Step 4:

[0404] Users access the interface through their device to view the analyzed situation. The input is text information converted into natural language, which is used to optimize the user interface for feedback. The user's camera and microphone are activated to collect emotional data, which is then analyzed by the server.

[0405] Step 5:

[0406] The server analyzes the collected emotional data using an emotion engine to identify the emotional state. The input is emotional data obtained from the user, and based on this, it outputs metrics to optimize the interface display content and feedback functions.

[0407] Step 6:

[0408] The server integrates detected situations and user emotional states to automatically generate a report. Inputs include naturally translated situational data and sentiment analysis results, and a visually detailed report is exported as an electronic file. This report provides meaningful insights for users and other stakeholders.

[0409] (Application Example 2)

[0410] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0411] Conventional video analysis systems detect and report specific events by analyzing video data, but they struggled to optimize the user experience and enrich reports without considering the user's emotional state. Furthermore, in quality control in factory settings, optimizing the environment while considering the emotions of workers is required, but solutions to this problem were lacking. Moreover, the generated reports were merely lists of facts and did not reflect the emotions or reactions of workers.

[0412] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0413] In this invention, the server includes means for receiving video data and storing it in an information storage device, means for analyzing the stored video data and detecting specific events within the video, means for converting the detected events into natural language, and emotion analysis means for collecting and analyzing user emotion data. This makes it possible to optimize the user interface considering the user's emotional state and to create reports that reflect the emotion analysis results.

[0414] "Video data" refers to a collection of visual information recorded using a camera or other recording device.

[0415] An "information storage device" is hardware or media used to store data temporarily or long-term.

[0416] "Means of analysis" refers to techniques or methods for processing given information and converting its content into a format that is easy to understand.

[0417] A "specific event" refers to a particular action or state identified within the video footage.

[0418] "Natural language" refers to the language that humans use on a daily basis, and is the form it takes before being converted into a language that can be interpreted by machines.

[0419] "Emotional analysis means" refers to technologies or methods for digitizing and analyzing human emotional states.

[0420] A "user interface" is a collection of visual and functional elements that allow a system and a user to interact with each other.

[0421] A "report" is a document generated by the system that includes analysis results and other supplementary information.

[0422] "Optimization" is the act of improving a system or process to achieve a specific objective, thereby increasing its efficiency and effectiveness.

[0423] To implement this invention, coordination between robots and servers in a factory monitoring system is crucial. The server first receives video data from cameras and other sources and stores it in an information storage device. The stored data is analyzed, and specific events such as defective products on the production line are detected using AI. Software such as TensorFlow and OpenCV are used for this purpose.

[0424] Next, the detected events are converted into natural language. This indexes the events in a human-readable format, making them easier to review and edit later.

[0425] This is where a unique emotion analysis method is incorporated. The robot collects emotional data from nearby workers through its camera and microphone, and analyzes their emotional state using the Azure Emotion API and Google Cloud Vision. This allows for the quantification of workers' stress levels and sense of security.

[0426] As users view events through the interface, this sentiment data is analyzed, and the user interface is optimized for optimal user experience. In this way, the server integrates the results of the sentiment analysis into a report, generating a document with more visual and detailed information.

[0427] As a concrete example, consider a factory with a new product manufacturing line. Factory robots perform video analysis to automatically detect minute defects in the products, while simultaneously adjusting the UI if nearby workers are experiencing stress. As a result, problem identification and optimization of the work environment are achieved simultaneously. An example of a prompt might be: "Explain how to detect defective products from manufacturing line video data and optimize the UI / UX based on the results of worker sentiment analysis."

[0428] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0429] Step 1:

[0430] The server receives video data from the factory cameras. The input is a real-time video stream captured by the cameras, which is taken into the server as data. The server saves this video to its information storage device. The output is the video data stored in the storage device.

[0431] Step 2:

[0432] The server analyzes the stored video data using an AI algorithm. The input is the video data saved in step 1. Using an AI algorithm (such as TensorFlow or OpenCV), object detection and event identification are performed on a frame-by-frame basis, and the output includes a list of detected events.

[0433] Step 3:

[0434] The server converts detected events into natural language. The input is the events detected in step 2. In this process, natural language generation technology is used to convert the extracted events into text format. The output is text data of the events as language.

[0435] Step 4:

[0436] The server indexes events within the video in a user-accessible format based on the verbalized event data. The input is the text data from step 3, and the indexing program generates an efficiently searchable index as needed. The output is event data with index information added.

[0437] Step 5:

[0438] As a user works around the robot, the robot collects emotional data from the user through its camera and microphone. The input consists of real-time video and audio data. Based on this, an emotional analysis algorithm (such as Azure Emotion API or Google Cloud Vision) is used to analyze the user's emotional state. The output is analytical data indicating the user's emotional state.

[0439] Step 6:

[0440] The server applies the emotion analysis results to provide a user-optimized interface. The input is the emotion state data from step 5. Based on the emotion data, the UI / UX components are dynamically adjusted to provide an interface that is easy for the user to use. The output is the optimized user interface.

[0441] Step 7:

[0442] The server automatically generates a report based on the combined data of detected events and sentiment analysis results. The input is the data from steps 4 and 5. Using the report generation program, a document conforming to the report format is created, and the output is compiled as a detailed, visual report.

[0443] Step 8:

[0444] The server exports the generated report in PDF or HTML format. The input is the report data generated in step 7. A format conversion program is used to convert the report into an electronic document format, which is then saved or sent as needed. The output is a document file in the specified format.

[0445] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0446] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0447] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0448] [Third Embodiment]

[0449] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0450] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0451] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0452] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0453] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0454] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0455] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0456] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0457] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0458] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0459] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0460] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0461] This invention is a system that efficiently analyzes video data, verbalizes and indexes specific events, thereby reducing the burden on supervisors and automating report creation. The system operates as follows:

[0462] First, the user uploads a video of their work from their device to the server. The uploaded video is automatically saved to storage by the server. The storage system maintains metadata about the data and is managed to ensure that subsequent processing proceeds smoothly.

[0463] Next, the server sequentially analyzes the stored video data. The analysis process examines the video frame by frame, using an AI algorithm to detect specific events and anomalies within the video. This AI algorithm utilizes object detection technology and records event timestamps.

[0464] Detected events are automatically translated into natural language by the server. This translation makes the data more human-readable. In addition, if necessary, speech synthesis technology is used to convert the data into speech.

[0465] Subsequently, the server indexes the relevant portion of the video based on the verbalized event information. This indexing process makes it possible to easily search for and play specific events.

[0466] The completed index data is provided as an interface for users to review and edit. Through this interface, users can view event details and, if necessary, modify descriptions or add additional notes.

[0467] Finally, the server automatically generates a report based on the edited data. The report can be exported in formats such as PDF and HTML and shared on other digital platforms. This allows all stakeholders to efficiently share information and address issues quickly.

[0468] As a concrete example, consider a video taken on a manufacturing line. When a user uploads this video to the system, the AI ​​automatically detects errors in handling parts and generates a text index such as "Incorrect placement of part A." The user can review this information through the interface and add further comments. When creating a report, a detailed report containing these events is automatically generated and shared with the relevant departments.

[0469] The following describes the processing flow.

[0470] Step 1:

[0471] The user uploads work video data from their terminal to the server. At this time, the system performs user authentication to confirm that the user has legitimate access rights.

[0472] Step 2:

[0473] The server saves the received video data to storage. During saving, it analyzes metadata related to the video (such as the date and time of shooting and camera location information) and registers it in the database.

[0474] Step 3:

[0475] The server divides the stored video data into frames and prepares them for analysis. This division is based on the video's frame rate and optimized for subsequent analysis.

[0476] Step 4:

[0477] The server uses an AI model to process each frame and detect specific events or anomalies within the video. For example, it might use an object detection algorithm to track the movement of an object of interest.

[0478] Step 5:

[0479] The server converts detected events into natural language text. Here, predefined rules and contextual analysis are used to express the information in a language that is easily understandable to humans.

[0480] Step 6:

[0481] The server indexes each event location in the video based on the verbalized event information. This process records detailed information, including the time and location of the event.

[0482] Step 7:

[0483] Users can access the server through their terminal and view an indexed list of events from the interface. Users can then view the list, select a specific event, and edit or add details.

[0484] Step 8:

[0485] The server automatically generates reports based on information edited by the user. The report's structure is determined, and event information is incorporated in a predetermined format.

[0486] Step 9:

[0487] The server exports the final report in PDF or HTML format and distributes it to the designated stakeholders. This output is then integrated into the company's workflow as needed.

[0488] (Example 1)

[0489] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0490] This invention aims to automate the process of analyzing and reporting video information, thereby reducing the burden on supervisors and enabling efficient information sharing. Conventional manual video analysis and report creation are time-consuming, labor-intensive, and prone to human error. Therefore, this invention provides a system to address these issues.

[0491] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0492] In this invention, the server includes means for receiving video information and storing it in a storage device, means for analyzing the stored video information and detecting specific events, and means for converting the detected events into natural language. This enables the rapid and accurate detection of events in the video, allows users to easily review and edit the information, and efficiently outputs automatically generated reports.

[0493] "Visual information" refers to dynamic or static visual data acquired through cameras or other acquisition devices.

[0494] A "storage device" is a piece of hardware or digital platform capable of storing data for extended periods.

[0495] "Analysis" refers to the process of examining data and extracting specific patterns or information.

[0496] An "event" refers to a specific event or anomaly that occurs within video information.

[0497] "Natural language" refers to the forms of expression that humans use in everyday life, and is used to describe data in a way that is easy to decipher.

[0498] "Indexing" is the process of assigning specific markers or tags to information to make it easier to search and refer to.

[0499] A "user" is a person or organization that has the authority to operate the system and to view and edit the results.

[0500] An "operation screen" is a display and input area for a user to interface with the system.

[0501] A "report" is a formatted document that compiles analyzed information and user comments.

[0502] "Digital document format" refers to a file format for documents that can be stored and shared electronically.

[0503] An "object recognition algorithm" is a computational method for detecting specific objects or patterns within video information.

[0504] A "sharing platform" is a system for sharing digital data online with multiple users or organizations.

[0505] This invention is a system that efficiently analyzes video information, verbalizes and indexes specific events, thereby reducing the burden on supervisors and automating report creation. An example of the implementation of this system is as follows:

[0506] First, the user uploads the work video from their device to the server. The device needs to have software capable of processing common video formats (e.g., MP4, AVI) and an internet connection. When the server receives this video information, it immediately saves it to its storage device. For storage devices, SSDs or HDDs are used to stably hold large amounts of data.

[0507] Next, the server analyzes the stored video information. This analysis requires computing resources to run AI algorithms, and object recognition algorithms such as YOLO (You Only Look Once) are used. The server inspects the video frame by frame, detects specific events, and records timestamps.

[0508] The server translates detected events into natural language. This process utilizes natural language processing (NLP) techniques to transform the data into a human-readable format. If necessary, speech synthesis technologies such as Google Text-to-Speech can be used to convert the generated text into speech.

[0509] Subsequently, the server indexes the relevant sections of the video based on the transcribed information, allowing users to easily search and play them. The index information, along with timestamps, is recorded and can be viewed and edited by users through the user interface. The interface also allows users to access event information and add comments.

[0510] Finally, the server automatically generates a report based on the edited information. The report is output in digital document formats such as PDF and HTML and provided on the company's shared platform. This allows all stakeholders to quickly share information and address issues efficiently.

[0511] One concrete example involves using video footage of a manufacturing line. When a user uploads this video to the system, an AI algorithm automatically detects errors in handling parts and generates linguistic data such as "Incorrect placement of part A." The user can then review this information on the operation screen and add comments as needed. An example of a prompt to the generating AI model in this case might be, "Can you detect and index errors in handling parts from a video of a manufacturing line?"

[0512] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0513] Step 1:

[0514] The user uploads the work video from their terminal to the server. The input is the video file stored on the terminal, and the output is a notification that the file transfer to the server is complete. At this stage, the user operates a dedicated upload screen, selects the target video file, and presses the transfer button to start the process.

[0515] Step 2:

[0516] The server stores the received video information in its storage device. The input is the moved video file, and the output is the video metadata (file format, size, resolution, etc.). This metadata is stored in a database for data management.

[0517] Step 3:

[0518] The server analyzes the stored video information. The input is a video file stored in a memory device, and the output is timestamp information related to a specific event. The server uses an object recognition algorithm (e.g., YOLO) to process the data frame by frame and detect events.

[0519] Step 4:

[0520] The server converts detected events into natural language. The input is timestamp information about the events, and the output is text data expressed in natural language. The server uses natural language processing (NLP) to generate the text. If necessary, the server may also consider using speech synthesis technology to convert the generated text into speech.

[0521] Step 5:

[0522] The server indexes the relevant parts of the video based on the transcribed information. The input consists of transcribed text and timestamp information, and the output is indexed data. This allows the user to view detailed event information using the interface.

[0523] Step 6:

[0524] Users view and edit indexed information through the user interface. Input is indexed data provided by the server, and output includes user comments and modifications. Users can add comments or modify content for events displayed on the screen.

[0525] Step 7:

[0526] The server automatically generates reports using the edited information. The input is user-edited data, and the output is a report in PDF or HTML format. This report is shared among stakeholders to help resolve issues quickly.

[0527] (Application Example 1)

[0528] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0529] In factories and workplaces, there is a need to quickly detect equipment malfunctions and quality defects, enabling immediate responses to workers, thereby improving production efficiency and reducing the burden on supervisors. However, conventional methods have the problem of requiring a great deal of time and effort to detect and report abnormalities.

[0530] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0531] In this invention, the server includes means for receiving video data and storing it in a data storage device, means for analyzing the stored video data and detecting specific events within the video, and means for converting the detected events into natural language. This makes it possible to detect abnormal operations in factories and work sites in real time and notify workers by voice.

[0532] "Video data" refers to digital information that records visual information acquired by a camera or camera.

[0533] A "data storage device" is an electronic device for securely storing video data over long periods of time, and it functions as part of a server.

[0534] An "event" refers to a specific action or change in state that is detected within the video data.

[0535] "Natural language" refers to the language that humans use on a daily basis and is used to express analyzed information.

[0536] "Indexing" is the process of assigning identifiers to digital data in order to facilitate the retrieval of information.

[0537] An "information display device" is a device that provides an interface for users to view and edit digital data and event content.

[0538] "Voice notification" is a means of conveying information to the user through voice about a specific event that has been detected.

[0539] An "object recognition algorithm" is a computational method for automatically identifying objects and patterns in video data.

[0540] "Electronic format" is a general term for file formats that can be displayed on a computer as digital information.

[0541] The embodiments for carrying out this invention are as follows.

[0542] The server first stores the video data transmitted by the user in a data storage device. During this storage process, the server adds metadata to streamline data management. Next, the server analyzes the stored video data. In this analysis, an AI model is used to detect various events within the video and convert them into natural language. In this process, an object recognition algorithm is used to accurately recognize and verbalize specific events.

[0543] The information, translated into natural language, is indexed by the server and managed in a user-friendly format. Users can view and edit this data using the provided information display device. Particularly important is the server's ability to communicate any abnormal activity detected in real time to the worker via voice notification, facilitating prompt action. Furthermore, based on the edited data, the server can automatically generate documents and export them electronically.

[0544] A concrete example is monitoring machine operation on a factory production line. If an operator is wearing smart glasses, they will receive an immediate voice notification if there is any abnormal movement, enabling early detection and rapid response to problems.

[0545] Examples of prompts to input into a generative AI model:

[0546] "The AI ​​will detect abnormal machine operation in the factory in real time and notify workers. Please propose specific methods to improve the efficiency of this process."

[0547] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0548] Step 1:

[0549] The server receives video data sent by the user and stores it in the data storage device. During this process, the raw data, as input, is stored along with metadata for data identification. The output is the storage location within the storage device associated with the metadata. Specifically, the server processes the file upload and creates a directory for data storage.

[0550] Step 2:

[0551] The server sequentially reads the stored video data and analyzes it using an object recognition AI model. The input is the stored video data, and the output is digital data that identifies specific events contained within it. The AI ​​model evaluates each frame and identifies specific actions or states. Specifically, the AI ​​generates data about the position and movement of objects and detects anomalies.

[0552] Step 3:

[0553] The server converts events detected by the AI ​​into natural language. The input is digital data representing the results of the event analysis, and the output is text in natural language that is understandable to humans. A natural language processing algorithm is used for this conversion, and the result is generated as a text file.

[0554] Step 4:

[0555] The server indexes events based on the converted natural language text. The input is natural language text, and the output is index information organized for easy searching and access. Specific operations include adding timestamps to the text and creating indexes using specific keywords.

[0556] Step 5:

[0557] The server transmits the indexed information to an information display device available to the user. The user uses this device to review the content and edit it as needed. The input is indexed information, and the output is display data that can be viewed and edited. Specifically, it provides editing and commenting functions to the user interface.

[0558] Step 6:

[0559] The server automatically generates documents based on edited information and exports them in electronic format. The input is edited index information, and the output is an electronic document such as PDF or HTML. Specifically, it uses a document generation engine to format the document and save it in the selected file format.

[0560] Step 7:

[0561] The server notifies workers of abnormal events detected in real time using speech synthesis technology. The input is real-time abnormality detection information, and the output is a notification in the form of a voice message. Specifically, the server generates synthesized speech and transmits it to smart glasses via the network.

[0562] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0563] This invention combines an emotion engine with a video analysis system to reflect the user's emotional state within the system, thereby improving the analysis and reporting process. This enhances the user experience and provides deeper insights.

[0564] First, the user uploads the work video to the server via their terminal. During this process, the system authenticates the user and ensures data integrity. The video file is recorded in storage by the server, and associated metadata is also saved simultaneously.

[0565] Next, the server analyzes the video data stored in storage frame by frame. Using AI algorithms, it inspects each frame to detect events within the video. During this process, object detection technology identifies specific actions or anomalies, and assigns corresponding timestamps.

[0566] The detected events are converted into natural language on the server. The conversion process includes rule-based sentence generation, and the output text is in a format that is easily understandable to humans. Furthermore, the events in the video are indexed based on this.

[0567] Here, a unique emotion engine is incorporated into the present invention. When a user uses the interface to view events, emotion data is collected through the camera and microphone. The emotion engine analyzes this data to identify the user's emotional state. This emotion information is used to optimize the interface's display content and feedback functions, providing the user with the optimal operating environment.

[0568] The server also associates detected events with the user's emotional state and incorporates this into the report. The report becomes more visual and detailed as it reflects the user's reactions and opinions as a result of the sentiment analysis.

[0569] As a concrete example, consider a quality control work video. When a user inputs video into the system, AI automatically detects specific events, such as defective products. Meanwhile, an emotion engine analyzes the user's stress and level of comfort, providing a pleasant UI / UX. Finally, a report is created based on this information, serving as a guideline for quality improvement.

[0570] The following describes the processing flow.

[0571] Step 1:

[0572] Users upload work video data to the server using their terminals. During this process, the system automatically authenticates the user and verifies access rights. Data uploads are only permitted if the user is legitimate.

[0573] Step 2:

[0574] The server stores the received video data in storage and simultaneously records related metadata (such as the date and time of shooting and camera location information). This information is used for reference in subsequent processes.

[0575] Step 3:

[0576] The server divides the stored video data into individual frames and prepares each frame for analysis by sending it to an AI algorithm. This division improves the accuracy of the analysis and enables the rapid detection of specific events.

[0577] Step 4:

[0578] The server uses AI algorithms to detect specific events and anomalies in each frame. Object detection technology is employed, and relevant events are timestamped. This process identifies events of interest within the video.

[0579] Step 5:

[0580] The server converts detected events into natural language and creates an index. The converted data is presented in an easy-to-read text format, and by simultaneously indexing which part of the video it corresponds to, consistent management becomes possible.

[0581] Step 6:

[0582] Users can view and edit indexed events through an interface via their device. Here, user sentiment data is collected using the camera and microphone, and a sentiment engine analyzes this data in real time.

[0583] Step 7:

[0584] The server optimizes the interface display and feedback by considering the user's emotional state, as analyzed by the emotion engine. This adjustment aims to improve the user experience and provide an efficient operating environment.

[0585] Step 8:

[0586] The server automatically generates a detailed work report by combining the results of editing and sentiment analysis. The report includes an event summary, index information, and analysis results based on user sentiment.

[0587] Step 9:

[0588] The server exports the generated report in PDF or HTML format and distributes it to the designated stakeholders. This format is easily integrated with other digital platforms, facilitating smooth report sharing.

[0589] (Example 2)

[0590] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0591] Conventional video analysis systems had the functionality to detect specific situations from video data and automatically generate report documents, but they did not optimize the interface or feedback to reflect the user's emotional state. As a result, the user experience was not sufficiently improved, and the insights provided by the system were limited. Furthermore, because it was not possible to create report documents that integrated the situation and the user's emotions, there was a lack of visual and detailed reporting.

[0592] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0593] In this invention, the server includes means for receiving video data and storing it in a storage device, means for analyzing the video data and detecting specific situations, means for converting the detected situations into natural language, means for acquiring and analyzing the user's emotional state, means for optimizing the interface and feedback based on the emotional state, and means for integrating the detected situations and emotional states and reflecting them in a report document. This enables deeper insights by reflecting the user's emotional state in the video analysis, and realizes the automatic generation of visually detailed report documents.

[0594] "Video data" refers to visual information acquired by cameras or other recording devices and stored in digital format.

[0595] A "storage device" is a device used to store data for a long or short period of time, and includes hard disk drives and solid-state drives.

[0596] "Analysis" is the process of extracting useful information from data, and it is carried out using algorithms and machine learning techniques.

[0597] "Situation" refers to a specific event or action that occurs within the video, and video analysis technology is used to detect it.

[0598] "Natural language" refers to the language that humans normally use, and it is sometimes applied to documents and texts generated by machines.

[0599] "Indexing" is the process of assigning keys and tags to data to make it easier to manage, facilitating searching and organization.

[0600] "User's emotional state" refers to the psychological or emotional response a user shows to a particular situation, and is often obtained through facial recognition or voice analysis.

[0601] "Integration" is the process of combining multiple data and pieces of information into a single system or report, enabling a more comprehensive understanding.

[0602] A "report document" is a document format that compiles analyzed data and information, and is generated according to a predetermined format.

[0603] This invention is a system that combines video analysis technology and emotion analysis technology, and generates a detailed report document using video data provided by the user. A specific embodiment of the system is described below.

[0604] The server receives video data and stores it in storage. High-speed, high-capacity storage is desirable for this process. Next, the server uses AI algorithms such as TensorFlow and OpenCV to analyze the video data frame by frame. When a specific situation is detected, object detection algorithms are applied to identify anomalies or unusual behavior and pinpoint the relevant parts.

[0605] The device provides an interface for the user to review the analyzed situation and provide feedback. Emotional data, such as facial expressions and voice tone, is captured through the user's camera and microphone. The server uses an emotion engine to analyze this data and identify the user's emotional state. Based on this information, the interface and feedback functions are optimized.

[0606] Ultimately, the server integrates the detected situation and the user's emotional state, automatically generating a detailed, visually easy-to-understand report. This report can be exported in electronic file format or markup language format.

[0607] As a concrete example, consider quality control videos in the manufacturing industry. When a user inputs video footage into the system, the AI ​​automatically analyzes specific situations, such as detecting defective products. Simultaneously, an emotion engine assesses the user's stress and sense of security, and reflects this in the report. This allows for deeper insights and improvement measures to be provided.

[0608] An example of a prompt message when using a generative AI model is, "Perform frame analysis on the uploaded video and create a detailed report including user sentiment analysis."

[0609] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0610] Step 1:

[0611] The user uploads video data to the server via their device. The input consists of the user's file selection and authentication information, which the server uses to verify data integrity and save it to storage. This step prepares the video data and associated metadata (e.g., timestamp, user ID) in storage.

[0612] Step 2:

[0613] The server analyzes video data stored in storage frame by frame. The input is video data stored in memory, and AI algorithms (e.g., TensorFlow, OpenCV) are used to analyze each frame. This detects specific events or actions within the video, and outputs timestamps indicating the situation.

[0614] Step 3:

[0615] The server translates the detected situation into natural language. The input is the timestamped situation data detected in step 2, which is then converted into an easy-to-read text format using a natural language generation engine. Through this step, the output is text in a format that is easily understandable to humans.

[0616] Step 4:

[0617] Users access the interface through their device to view the analyzed situation. The input is text information converted into natural language, which is used to optimize the user interface for feedback. The user's camera and microphone are activated to collect emotional data, which is then analyzed by the server.

[0618] Step 5:

[0619] The server analyzes the collected emotional data using an emotion engine to identify the emotional state. The input is emotional data obtained from the user, and based on this, it outputs metrics to optimize the interface display content and feedback functions.

[0620] Step 6:

[0621] The server integrates detected situations and user emotional states to automatically generate a report. Inputs include naturally translated situational data and sentiment analysis results, and a visually detailed report is exported as an electronic file. This report provides meaningful insights for users and other stakeholders.

[0622] (Application Example 2)

[0623] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0624] Conventional video analysis systems detect and report specific events by analyzing video data, but they struggled to optimize the user experience and enrich reports without considering the user's emotional state. Furthermore, in quality control in factory settings, optimizing the environment while considering the emotions of workers is required, but solutions to this problem were lacking. Moreover, the generated reports were merely lists of facts and did not reflect the emotions or reactions of workers.

[0625] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0626] In this invention, the server includes means for receiving video data and storing it in an information storage device, means for analyzing the stored video data and detecting specific events within the video, means for converting the detected events into natural language, and emotion analysis means for collecting and analyzing user emotion data. This makes it possible to optimize the user interface considering the user's emotional state and to create reports that reflect the emotion analysis results.

[0627] "Video data" refers to a collection of visual information recorded using a camera or other recording device.

[0628] An "information storage device" is hardware or media used to store data temporarily or long-term.

[0629] "Means of analysis" refers to techniques or methods for processing given information and converting its content into a format that is easy to understand.

[0630] A "specific event" refers to a particular action or state identified within the video footage.

[0631] "Natural language" refers to the language that humans use on a daily basis, and is the form it takes before being converted into a language that can be interpreted by machines.

[0632] "Emotional analysis means" refers to technologies or methods for digitizing and analyzing human emotional states.

[0633] A "user interface" is a collection of visual and functional elements that allow a system and a user to interact with each other.

[0634] A "report" is a document generated by the system that includes analysis results and other supplementary information.

[0635] "Optimization" is the act of improving a system or process to achieve a specific objective, thereby increasing its efficiency and effectiveness.

[0636] To implement this invention, coordination between robots and servers in a factory monitoring system is crucial. The server first receives video data from cameras and other sources and stores it in an information storage device. The stored data is analyzed, and specific events such as defective products on the production line are detected using AI. Software such as TensorFlow and OpenCV are used for this purpose.

[0637] Next, the detected events are converted into natural language. This indexes the events in a human-readable format, making them easier to review and edit later.

[0638] This is where a unique emotion analysis method is incorporated. The robot collects emotional data from nearby workers through its camera and microphone, and analyzes their emotional state using the Azure Emotion API and Google Cloud Vision. This allows for the quantification of workers' stress levels and sense of security.

[0639] As users view events through the interface, this sentiment data is analyzed, and the user interface is optimized for optimal user experience. In this way, the server integrates the results of the sentiment analysis into a report, generating a document with more visual and detailed information.

[0640] As a concrete example, consider a factory with a new product manufacturing line. Factory robots perform video analysis to automatically detect minute defects in the products, while simultaneously adjusting the UI if nearby workers are experiencing stress. As a result, problem identification and optimization of the work environment are achieved simultaneously. An example of a prompt might be: "Explain how to detect defective products from manufacturing line video data and optimize the UI / UX based on the results of worker sentiment analysis."

[0641] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0642] Step 1:

[0643] The server receives video data from the factory cameras. The input is a real-time video stream captured by the cameras, which is taken into the server as data. The server saves this video to its information storage device. The output is the video data stored in the storage device.

[0644] Step 2:

[0645] The server analyzes the stored video data using an AI algorithm. The input is the video data saved in step 1. Using an AI algorithm (such as TensorFlow or OpenCV), object detection and event identification are performed on a frame-by-frame basis, and the output includes a list of detected events.

[0646] Step 3:

[0647] The server converts detected events into natural language. The input is the events detected in step 2. In this process, natural language generation technology is used to convert the extracted events into text format. The output is text data of the events as language.

[0648] Step 4:

[0649] The server indexes events within the video in a user-accessible format based on the verbalized event data. The input is the text data from step 3, and the indexing program generates an efficiently searchable index as needed. The output is event data with index information added.

[0650] Step 5:

[0651] As a user works around the robot, the robot collects emotional data from the user through its camera and microphone. The input consists of real-time video and audio data. Based on this, an emotional analysis algorithm (such as Azure Emotion API or Google Cloud Vision) is used to analyze the user's emotional state. The output is analytical data indicating the user's emotional state.

[0652] Step 6:

[0653] The server applies the emotion analysis results to provide a user-optimized interface. The input is the emotion state data from step 5. Based on the emotion data, the UI / UX components are dynamically adjusted to provide an interface that is easy for the user to use. The output is the optimized user interface.

[0654] Step 7:

[0655] The server automatically generates a report based on the combined data of detected events and sentiment analysis results. The input is the data from steps 4 and 5. Using the report generation program, a document conforming to the report format is created, and the output is compiled as a detailed, visual report.

[0656] Step 8:

[0657] The server exports the generated report in PDF or HTML format. The input is the report data generated in step 7. A format conversion program is used to convert the report into an electronic document format, which is then saved or sent as needed. The output is a document file in the specified format.

[0658] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0659] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0660] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0661] [Fourth Embodiment]

[0662] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0663] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0664] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0665] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0666] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0667] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0668] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0669] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors in the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0670] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0671] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0672] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0673] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0674] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0675] This invention is a system that efficiently analyzes video data, verbalizes and indexes specific events, thereby reducing the burden on supervisors and automating report creation. The system operates as follows:

[0676] First, the user uploads a video of their work from their device to the server. The uploaded video is automatically saved to storage by the server. The storage system maintains metadata about the data and is managed to ensure that subsequent processing proceeds smoothly.

[0677] Next, the server sequentially analyzes the stored video data. The analysis process examines the video frame by frame, using an AI algorithm to detect specific events and anomalies within the video. This AI algorithm utilizes object detection technology and records event timestamps.

[0678] Detected events are automatically translated into natural language by the server. This translation makes the data more human-readable. In addition, if necessary, speech synthesis technology is used to convert the data into speech.

[0679] Subsequently, the server indexes the relevant portion of the video based on the verbalized event information. This indexing process makes it possible to easily search for and play specific events.

[0680] The completed index data is provided as an interface for users to review and edit. Through this interface, users can view event details and, if necessary, modify descriptions or add additional notes.

[0681] Finally, the server automatically generates a report based on the edited data. The report can be exported in formats such as PDF and HTML and shared on other digital platforms. This allows all stakeholders to efficiently share information and address issues quickly.

[0682] As a concrete example, consider a video taken on a manufacturing line. When a user uploads this video to the system, the AI ​​automatically detects errors in handling parts and generates a text index such as "Incorrect placement of part A." The user can review this information through the interface and add further comments. When creating a report, a detailed report containing these events is automatically generated and shared with the relevant departments.

[0683] The following describes the processing flow.

[0684] Step 1:

[0685] The user uploads work video data from their terminal to the server. At this time, the system performs user authentication to confirm that the user has legitimate access rights.

[0686] Step 2:

[0687] The server saves the received video data to storage. During saving, it analyzes metadata related to the video (such as the date and time of shooting and camera location information) and registers it in the database.

[0688] Step 3:

[0689] The server divides the stored video data into frames and prepares them for analysis. This division is based on the video's frame rate and optimized for subsequent analysis.

[0690] Step 4:

[0691] The server uses an AI model to process each frame and detect specific events or anomalies within the video. For example, it might use an object detection algorithm to track the movement of an object of interest.

[0692] Step 5:

[0693] The server converts detected events into natural language text. Here, predefined rules and contextual analysis are used to express the information in a language that is easily understandable to humans.

[0694] Step 6:

[0695] The server indexes each event location in the video based on the verbalized event information. This process records detailed information, including the time and location of the event.

[0696] Step 7:

[0697] Users can access the server through their terminal and view an indexed list of events from the interface. Users can then view the list, select a specific event, and edit or add details.

[0698] Step 8:

[0699] The server automatically generates reports based on information edited by the user. The report's structure is determined, and event information is incorporated in a predetermined format.

[0700] Step 9:

[0701] The server exports the final report in PDF or HTML format and distributes it to the designated stakeholders. This output is then integrated into the company's workflow as needed.

[0702] (Example 1)

[0703] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0704] This invention aims to automate the process of analyzing and reporting video information, thereby reducing the burden on supervisors and enabling efficient information sharing. Conventional manual video analysis and report creation are time-consuming, labor-intensive, and prone to human error. Therefore, this invention provides a system to address these issues.

[0705] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0706] In this invention, the server includes means for receiving video information and storing it in a storage device, means for analyzing the stored video information and detecting specific events, and means for converting the detected events into natural language. This enables the rapid and accurate detection of events in the video, allows users to easily review and edit the information, and efficiently outputs automatically generated reports.

[0707] "Visual information" refers to dynamic or static visual data acquired through cameras or other acquisition devices.

[0708] A "storage device" is a piece of hardware or digital platform capable of storing data for extended periods.

[0709] "Analysis" refers to the process of examining data and extracting specific patterns or information.

[0710] An "event" refers to a specific event or anomaly that occurs within video information.

[0711] "Natural language" refers to the forms of expression that humans use in everyday life, and is used to describe data in a way that is easy to decipher.

[0712] "Indexing" is the process of assigning specific markers or tags to information to make it easier to search and refer to.

[0713] A "user" is a person or organization that has the authority to operate the system and to view and edit the results.

[0714] An "operation screen" is a display and input area for a user to interface with the system.

[0715] A "report" is a formatted document that compiles analyzed information and user comments.

[0716] "Digital document format" refers to a file format for documents that can be stored and shared electronically.

[0717] An "object recognition algorithm" is a computational method for detecting specific objects or patterns within video information.

[0718] A "sharing platform" is a system for sharing digital data online with multiple users or organizations.

[0719] This invention is a system that efficiently analyzes video information, verbalizes and indexes specific events, thereby reducing the burden on supervisors and automating report creation. An example of the implementation of this system is as follows:

[0720] First, the user uploads the work video from their device to the server. The device needs to have software capable of processing common video formats (e.g., MP4, AVI) and an internet connection. When the server receives this video information, it immediately saves it to its storage device. For storage devices, SSDs or HDDs are used to stably hold large amounts of data.

[0721] Next, the server analyzes the stored video information. This analysis requires computing resources to run AI algorithms, and object recognition algorithms such as YOLO (You Only Look Once) are used. The server inspects the video frame by frame, detects specific events, and records timestamps.

[0722] The server translates detected events into natural language. This process utilizes natural language processing (NLP) techniques to transform the data into a human-readable format. If necessary, speech synthesis technologies such as Google Text-to-Speech can be used to convert the generated text into speech.

[0723] Subsequently, the server indexes the relevant sections of the video based on the transcribed information, allowing users to easily search and play them. The index information, along with timestamps, is recorded and can be viewed and edited by users through the user interface. The interface also allows users to access event information and add comments.

[0724] Finally, the server automatically generates a report based on the edited information. The report is output in digital document formats such as PDF and HTML and provided on the company's shared platform. This allows all stakeholders to quickly share information and address issues efficiently.

[0725] One concrete example involves using video footage of a manufacturing line. When a user uploads this video to the system, an AI algorithm automatically detects errors in handling parts and generates linguistic data such as "Incorrect placement of part A." The user can then review this information on the operation screen and add comments as needed. An example of a prompt to the generating AI model in this case might be, "Can you detect and index errors in handling parts from a video of a manufacturing line?"

[0726] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0727] Step 1:

[0728] The user uploads the work video from their terminal to the server. The input is the video file stored on the terminal, and the output is a notification that the file transfer to the server is complete. At this stage, the user operates a dedicated upload screen, selects the target video file, and presses the transfer button to start the process.

[0729] Step 2:

[0730] The server stores the received video information in its storage device. The input is the moved video file, and the output is the video metadata (file format, size, resolution, etc.). This metadata is stored in a database for data management.

[0731] Step 3:

[0732] The server analyzes the stored video information. The input is a video file stored in a memory device, and the output is timestamp information related to a specific event. The server uses an object recognition algorithm (e.g., YOLO) to process the data frame by frame and detect events.

[0733] Step 4:

[0734] The server converts detected events into natural language. The input is timestamp information about the events, and the output is text data expressed in natural language. The server uses natural language processing (NLP) to generate the text. If necessary, the server may also consider using speech synthesis technology to convert the generated text into speech.

[0735] Step 5:

[0736] The server indexes the relevant parts of the video based on the transcribed information. The input consists of transcribed text and timestamp information, and the output is indexed data. This allows the user to view detailed event information using the interface.

[0737] Step 6:

[0738] Users view and edit indexed information through the user interface. Input is indexed data provided by the server, and output includes user comments and modifications. Users can add comments or modify content for events displayed on the screen.

[0739] Step 7:

[0740] The server automatically generates reports using the edited information. The input is user-edited data, and the output is a report in PDF or HTML format. This report is shared among stakeholders to help resolve issues quickly.

[0741] (Application Example 1)

[0742] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0743] In factories and workplaces, there is a need to quickly detect equipment malfunctions and quality defects, enabling immediate responses to workers, thereby improving production efficiency and reducing the burden on supervisors. However, conventional methods have the problem of requiring a great deal of time and effort to detect and report abnormalities.

[0744] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0745] In this invention, the server includes means for receiving video data and storing it in a data storage device, means for analyzing the stored video data and detecting specific events within the video, and means for converting the detected events into natural language. This makes it possible to detect abnormal operations in factories and work sites in real time and notify workers by voice.

[0746] "Video data" refers to digital information that records visual information acquired by a camera or camera.

[0747] A "data storage device" is an electronic device for securely storing video data over long periods of time, and it functions as part of a server.

[0748] An "event" refers to a specific action or change in state that is detected within the video data.

[0749] "Natural language" refers to the language that humans use on a daily basis and is used to express analyzed information.

[0750] "Indexing" is the process of assigning identifiers to digital data in order to facilitate the retrieval of information.

[0751] An "information display device" is a device that provides an interface for users to view and edit digital data and event content.

[0752] "Voice notification" is a means of conveying information to the user through voice about a specific event that has been detected.

[0753] An "object recognition algorithm" is a computational method for automatically identifying objects and patterns in video data.

[0754] "Electronic format" is a general term for file formats that can be displayed on a computer as digital information.

[0755] The embodiments for carrying out this invention are as follows.

[0756] The server first stores the video data transmitted by the user in a data storage device. During this storage process, the server adds metadata to streamline data management. Next, the server analyzes the stored video data. In this analysis, an AI model is used to detect various events within the video and convert them into natural language. In this process, an object recognition algorithm is used to accurately recognize and verbalize specific events.

[0757] The information, translated into natural language, is indexed by the server and managed in a user-friendly format. Users can view and edit this data using the provided information display device. Particularly important is the server's ability to communicate any abnormal activity detected in real time to the worker via voice notification, facilitating prompt action. Furthermore, based on the edited data, the server can automatically generate documents and export them electronically.

[0758] A concrete example is monitoring machine operation on a factory production line. If an operator is wearing smart glasses, they will receive an immediate voice notification if there is any abnormal movement, enabling early detection and rapid response to problems.

[0759] Examples of prompts to input into a generative AI model:

[0760] "The AI ​​will detect abnormal machine operation in the factory in real time and notify workers. Please propose specific methods to improve the efficiency of this process."

[0761] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0762] Step 1:

[0763] The server receives video data sent by the user and stores it in the data storage device. During this process, the raw data, as input, is stored along with metadata for data identification. The output is the storage location within the storage device associated with the metadata. Specifically, the server processes the file upload and creates a directory for data storage.

[0764] Step 2:

[0765] The server sequentially reads the stored video data and analyzes it using an object recognition AI model. The input is the stored video data, and the output is digital data that identifies specific events contained within it. The AI ​​model evaluates each frame and identifies specific actions or states. Specifically, the AI ​​generates data about the position and movement of objects and detects anomalies.

[0766] Step 3:

[0767] The server converts events detected by the AI ​​into natural language. The input is digital data representing the results of the event analysis, and the output is text in natural language that is understandable to humans. A natural language processing algorithm is used for this conversion, and the result is generated as a text file.

[0768] Step 4:

[0769] The server indexes events based on the converted natural language text. The input is natural language text, and the output is index information organized for easy searching and access. Specific operations include adding timestamps to the text and creating indexes using specific keywords.

[0770] Step 5:

[0771] The server transmits the indexed information to an information display device available to the user. The user uses this device to review the content and edit it as needed. The input is indexed information, and the output is display data that can be viewed and edited. Specifically, it provides editing and commenting functions to the user interface.

[0772] Step 6:

[0773] The server automatically generates documents based on edited information and exports them in electronic format. The input is edited index information, and the output is an electronic document such as PDF or HTML. Specifically, it uses a document generation engine to format the document and save it in the selected file format.

[0774] Step 7:

[0775] The server notifies workers of abnormal events detected in real time using speech synthesis technology. The input is real-time abnormality detection information, and the output is a notification in the form of a voice message. Specifically, the server generates synthesized speech and transmits it to smart glasses via the network.

[0776] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0777] This invention combines an emotion engine with a video analysis system to reflect the user's emotional state within the system, thereby improving the analysis and reporting process. This enhances the user experience and provides deeper insights.

[0778] First, the user uploads the work video to the server via their terminal. During this process, the system authenticates the user and ensures data integrity. The video file is recorded in storage by the server, and associated metadata is also saved simultaneously.

[0779] Next, the server analyzes the video data stored in storage frame by frame. Using AI algorithms, it inspects each frame to detect events within the video. During this process, object detection technology identifies specific actions or anomalies, and assigns corresponding timestamps.

[0780] The detected events are converted into natural language on the server. The conversion process includes rule-based sentence generation, and the output text is in a format that is easily understandable to humans. Furthermore, the events in the video are indexed based on this.

[0781] Here, a unique emotion engine is incorporated into the present invention. When a user uses the interface to view events, emotion data is collected through the camera and microphone. The emotion engine analyzes this data to identify the user's emotional state. This emotion information is used to optimize the interface's display content and feedback functions, providing the user with the optimal operating environment.

[0782] The server also associates detected events with the user's emotional state and incorporates this into the report. The report becomes more visual and detailed as it reflects the user's reactions and opinions as a result of the sentiment analysis.

[0783] As a concrete example, consider a quality control work video. When a user inputs video into the system, AI automatically detects specific events, such as defective products. Meanwhile, an emotion engine analyzes the user's stress and level of comfort, providing a pleasant UI / UX. Finally, a report is created based on this information, serving as a guideline for quality improvement.

[0784] The following describes the processing flow.

[0785] Step 1:

[0786] Users upload work video data to the server using their terminals. During this process, the system automatically authenticates the user and verifies access rights. Data uploads are only permitted if the user is legitimate.

[0787] Step 2:

[0788] The server stores the received video data in storage and simultaneously records related metadata (such as the date and time of shooting and camera location information). This information is used for reference in subsequent processes.

[0789] Step 3:

[0790] The server divides the stored video data into individual frames and prepares each frame for analysis by sending it to an AI algorithm. This division improves the accuracy of the analysis and enables the rapid detection of specific events.

[0791] Step 4:

[0792] The server uses AI algorithms to detect specific events and anomalies in each frame. Object detection technology is employed, and relevant events are timestamped. This process identifies events of interest within the video.

[0793] Step 5:

[0794] The server converts detected events into natural language and creates an index. The converted data is presented in an easy-to-read text format, and by simultaneously indexing which part of the video it corresponds to, consistent management becomes possible.

[0795] Step 6:

[0796] Users can view and edit indexed events through an interface via their device. Here, user sentiment data is collected using the camera and microphone, and a sentiment engine analyzes this data in real time.

[0797] Step 7:

[0798] The server optimizes the interface display and feedback by considering the user's emotional state, as analyzed by the emotion engine. This adjustment aims to improve the user experience and provide an efficient operating environment.

[0799] Step 8:

[0800] The server automatically generates a detailed work report by combining the results of editing and sentiment analysis. The report includes an event summary, index information, and analysis results based on user sentiment.

[0801] Step 9:

[0802] The server exports the generated report in PDF or HTML format and distributes it to the designated stakeholders. This format is easily integrated with other digital platforms, facilitating smooth report sharing.

[0803] (Example 2)

[0804] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0805] Conventional video analysis systems had the functionality to detect specific situations from video data and automatically generate report documents, but they did not optimize the interface or feedback to reflect the user's emotional state. As a result, the user experience was not sufficiently improved, and the insights provided by the system were limited. Furthermore, because it was not possible to create report documents that integrated the situation and the user's emotions, there was a lack of visual and detailed reporting.

[0806] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0807] In this invention, the server includes means for receiving video data and storing it in a storage device, means for analyzing the video data and detecting specific situations, means for converting the detected situations into natural language, means for acquiring and analyzing the user's emotional state, means for optimizing the interface and feedback based on the emotional state, and means for integrating the detected situations and emotional states and reflecting them in a report document. This enables deeper insights by reflecting the user's emotional state in the video analysis, and realizes the automatic generation of visually detailed report documents.

[0808] "Video data" refers to visual information acquired by cameras or other recording devices and stored in digital format.

[0809] A "storage device" is a device used to store data for a long or short period of time, and includes hard disk drives and solid-state drives.

[0810] "Analysis" is the process of extracting useful information from data, and it is carried out using algorithms and machine learning techniques.

[0811] "Situation" refers to a specific event or action that occurs within the video, and video analysis technology is used to detect it.

[0812] "Natural language" refers to the language that humans normally use, and it is sometimes applied to documents and texts generated by machines.

[0813] "Indexing" is the process of assigning keys and tags to data to make it easier to manage, facilitating searching and organization.

[0814] "User's emotional state" refers to the psychological or emotional response a user shows to a particular situation, and is often obtained through facial recognition or voice analysis.

[0815] "Integration" is the process of combining multiple data and pieces of information into a single system or report, enabling a more comprehensive understanding.

[0816] A "report document" is a document format that compiles analyzed data and information, and is generated according to a predetermined format.

[0817] This invention is a system that combines video analysis technology and emotion analysis technology, and generates a detailed report document using video data provided by the user. A specific embodiment of the system is described below.

[0818] The server receives video data and stores it in storage. High-speed, high-capacity storage is desirable for this process. Next, the server uses AI algorithms such as TensorFlow and OpenCV to analyze the video data frame by frame. When a specific situation is detected, object detection algorithms are applied to identify anomalies or unusual behavior and pinpoint the relevant parts.

[0819] The device provides an interface for the user to review the analyzed situation and provide feedback. Emotional data, such as facial expressions and voice tone, is captured through the user's camera and microphone. The server uses an emotion engine to analyze this data and identify the user's emotional state. Based on this information, the interface and feedback functions are optimized.

[0820] Ultimately, the server integrates the detected situation and the user's emotional state, automatically generating a detailed, visually easy-to-understand report. This report can be exported in electronic file format or markup language format.

[0821] As a concrete example, consider quality control videos in the manufacturing industry. When a user inputs video footage into the system, the AI ​​automatically analyzes specific situations, such as detecting defective products. Simultaneously, an emotion engine assesses the user's stress and sense of security, and reflects this in the report. This allows for deeper insights and improvement measures to be provided.

[0822] An example of a prompt message when using a generative AI model is, "Perform frame analysis on the uploaded video and create a detailed report including user sentiment analysis."

[0823] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0824] Step 1:

[0825] The user uploads video data to the server via their device. The input consists of the user's file selection and authentication information, which the server uses to verify data integrity and save it to storage. This step prepares the video data and associated metadata (e.g., timestamp, user ID) in storage.

[0826] Step 2:

[0827] The server analyzes video data stored in storage frame by frame. The input is video data stored in memory, and AI algorithms (e.g., TensorFlow, OpenCV) are used to analyze each frame. This detects specific events or actions within the video, and outputs timestamps indicating the situation.

[0828] Step 3:

[0829] The server translates the detected situation into natural language. The input is the timestamped situation data detected in step 2, which is then converted into an easy-to-read text format using a natural language generation engine. Through this step, the output is text in a format that is easily understandable to humans.

[0830] Step 4:

[0831] Users access the interface through their device to view the analyzed situation. The input is text information converted into natural language, which is used to optimize the user interface for feedback. The user's camera and microphone are activated to collect emotional data, which is then analyzed by the server.

[0832] Step 5:

[0833] The server analyzes the collected emotional data using an emotion engine to identify the emotional state. The input is emotional data obtained from the user, and based on this, it outputs metrics to optimize the interface display content and feedback functions.

[0834] Step 6:

[0835] The server integrates detected situations and user emotional states to automatically generate a report. Inputs include naturally translated situational data and sentiment analysis results, and a visually detailed report is exported as an electronic file. This report provides meaningful insights for users and other stakeholders.

[0836] (Application Example 2)

[0837] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0838] Conventional video analysis systems detect and report specific events by analyzing video data, but they struggled to optimize the user experience and enrich reports without considering the user's emotional state. Furthermore, in quality control in factory settings, optimizing the environment while considering the emotions of workers is required, but solutions to this problem were lacking. Moreover, the generated reports were merely lists of facts and did not reflect the emotions or reactions of workers.

[0839] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0840] In this invention, the server includes means for receiving video data and storing it in an information storage device, means for analyzing the stored video data and detecting specific events within the video, means for converting the detected events into natural language, and emotion analysis means for collecting and analyzing user emotion data. This makes it possible to optimize the user interface considering the user's emotional state and to create reports that reflect the emotion analysis results.

[0841] "Video data" refers to a collection of visual information recorded using a camera or other recording device.

[0842] An "information storage device" is hardware or media used to store data temporarily or long-term.

[0843] "Means of analysis" refers to techniques or methods for processing given information and converting its content into a format that is easy to understand.

[0844] A "specific event" refers to a particular action or state identified within the video footage.

[0845] "Natural language" refers to the language that humans use on a daily basis, and is the form it takes before being converted into a language that can be interpreted by machines.

[0846] "Emotional analysis means" refers to technologies or methods for digitizing and analyzing human emotional states.

[0847] A "user interface" is a collection of visual and functional elements that allow a system and a user to interact with each other.

[0848] A "report" is a document generated by the system that includes analysis results and other supplementary information.

[0849] "Optimization" is the act of improving a system or process to achieve a specific objective, thereby increasing its efficiency and effectiveness.

[0850] To implement this invention, coordination between robots and servers in a factory monitoring system is crucial. The server first receives video data from cameras and other sources and stores it in an information storage device. The stored data is analyzed, and specific events such as defective products on the production line are detected using AI. Software such as TensorFlow and OpenCV are used for this purpose.

[0851] Next, the detected events are converted into natural language. This indexes the events in a human-readable format, making them easier to review and edit later.

[0852] This is where a unique emotion analysis method is incorporated. The robot collects emotional data from nearby workers through its camera and microphone, and analyzes their emotional state using the Azure Emotion API and Google Cloud Vision. This allows for the quantification of workers' stress levels and sense of security.

[0853] As users view events through the interface, this sentiment data is analyzed, and the user interface is optimized for optimal user experience. In this way, the server integrates the results of the sentiment analysis into a report, generating a document with more visual and detailed information.

[0854] As a concrete example, consider a factory with a new product manufacturing line. Factory robots perform video analysis to automatically detect minute defects in the products, while simultaneously adjusting the UI if nearby workers are experiencing stress. As a result, problem identification and optimization of the work environment are achieved simultaneously. An example of a prompt might be: "Explain how to detect defective products from manufacturing line video data and optimize the UI / UX based on the results of worker sentiment analysis."

[0855] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0856] Step 1:

[0857] The server receives video data from the factory cameras. The input is a real-time video stream captured by the cameras, which is taken into the server as data. The server saves this video to its information storage device. The output is the video data stored in the storage device.

[0858] Step 2:

[0859] The server analyzes the stored video data using an AI algorithm. The input is the video data saved in step 1. Using an AI algorithm (such as TensorFlow or OpenCV), object detection and event identification are performed on a frame-by-frame basis, and the output includes a list of detected events.

[0860] Step 3:

[0861] The server converts detected events into natural language. The input is the events detected in step 2. In this process, natural language generation technology is used to convert the extracted events into text format. The output is text data of the events as language.

[0862] Step 4:

[0863] The server indexes events within the video in a user-accessible format based on the verbalized event data. The input is the text data from step 3, and the indexing program generates an efficiently searchable index as needed. The output is event data with index information added.

[0864] Step 5:

[0865] As a user works around the robot, the robot collects emotional data from the user through its camera and microphone. The input consists of real-time video and audio data. Based on this, an emotional analysis algorithm (such as Azure Emotion API or Google Cloud Vision) is used to analyze the user's emotional state. The output is analytical data indicating the user's emotional state.

[0866] Step 6:

[0867] The server applies the emotion analysis results to provide a user-optimized interface. The input is the emotion state data from step 5. Based on the emotion data, the UI / UX components are dynamically adjusted to provide an interface that is easy for the user to use. The output is the optimized user interface.

[0868] Step 7:

[0869] The server automatically generates a report based on the combined data of detected events and sentiment analysis results. The input is the data from steps 4 and 5. Using the report generation program, a document conforming to the report format is created, and the output is compiled as a detailed, visual report.

[0870] Step 8:

[0871] The server exports the generated report in PDF or HTML format. The input is the report data generated in step 7. A format conversion program is used to convert the report into an electronic document format, which is then saved or sent as needed. The output is a document file in the specified format.

[0872] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0873] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0874] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0875] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0876] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0877] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0878] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0879] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0880] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0881] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0882] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0883] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0884] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0885] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0886] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0887] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0888] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0889] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0890] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0891] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0892] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0893] The following is further disclosed regarding the embodiments described above.

[0894] (Claim 1)

[0895] A means of receiving video data and saving it to storage,

[0896] A means for analyzing stored video data and detecting specific events within the video,

[0897] A means of converting detected events into natural language,

[0898] A method for indexing events within a video based on verbalized information,

[0899] A means of providing an interface for users to view and edit indexed events,

[0900] A means of automatically generating reports using edited information,

[0901] A system that includes this.

[0902] (Claim 2)

[0903] The system according to claim 1, wherein an object detection algorithm is used in event detection.

[0904] (Claim 3)

[0905] The system according to claim 1, which exports a report in PDF or HTML format.

[0906] "Example 1"

[0907] (Claim 1)

[0908] A means for receiving video information and storing it in a storage device,

[0909] A means for analyzing stored video information and detecting specific events within the video,

[0910] A means of converting detected events into natural language,

[0911] A means of indexing events in a video based on verbalized information,

[0912] A means for providing a user interface that allows users to view and edit indexed events,

[0913] A means of automatically generating reports using edited information,

[0914] A means of outputting the generated report in digital document format,

[0915] A system that includes this.

[0916] (Claim 2)

[0917] The system according to claim 1, wherein an object recognition algorithm is used in event detection.

[0918] (Claim 3)

[0919] The system according to claim 1, which provides reports on a shared platform.

[0920] "Application Example 1"

[0921] (Claim 1)

[0922] A means for receiving video data and storing it in a data storage device,

[0923] A means for analyzing stored video data and detecting specific events within the video,

[0924] A means of converting detected events into natural language,

[0925] A means of indexing events in a video based on verbalized information,

[0926] A means for providing an information display device that allows users to view and edit indexed events,

[0927] A means of automatically generating documents using edited information,

[0928] A means of notifying workers of abnormal operations detected in real time via voice,

[0929] A system that includes this.

[0930] (Claim 2)

[0931] The system according to claim 1, wherein an object recognition algorithm is used in event detection.

[0932] (Claim 3)

[0933] The system according to claim 1 for exporting documents in electronic format.

[0934] "Example 2 of combining an emotion engine"

[0935] (Claim 1)

[0936] A means for receiving video data and saving it to a storage device,

[0937] A means for analyzing stored video data and detecting specific situations within the video,

[0938] A means of converting the detected situation into natural language,

[0939] A means of indexing the situation in a video based on verbalized information,

[0940] A means of providing an interface for users to view and edit indexed status,

[0941] A means of automatically generating a report document using edited information,

[0942] A means of acquiring and analyzing the emotional state of users,

[0943] Means for optimizing interfaces and feedback based on emotional states,

[0944] A means of integrating detected situations and emotional states and reflecting them in the report document,

[0945] A system that includes this.

[0946] (Claim 2)

[0947] The system according to claim 1, wherein an object detection algorithm is used in situation detection.

[0948] (Claim 3)

[0949] The system according to claim 1, which exports a report document in an electronic file format or a markup language format.

[0950] "Application example 2 when combining with an emotional engine"

[0951] (Claim 1)

[0952] A means for receiving video data and storing it in an information storage device,

[0953] A means for analyzing stored video data and detecting specific events within the video,

[0954] A means of converting detected events into natural language,

[0955] A means of indexing events within a video based on verbalized information,

[0956] A means of providing a user interface for users to view and edit indexed events,

[0957] A means of automatically generating reports using edited information,

[0958] A sentiment analysis method for collecting and analyzing user sentiment data,

[0959] A means for optimizing the display of the user interface based on the results of sentiment analysis,

[0960] Methods for incorporating sentiment analysis results into reports,

[0961] A system that includes this.

[0962] (Claim 2)

[0963] The system according to claim 1, which uses an object detection algorithm in event detection.

[0964] (Claim 3)

[0965] The system according to claim 1 for exporting a report in digital document format. [Explanation of Symbols]

[0966] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of receiving video data and saving it to storage, A means for analyzing stored video data and detecting specific events within the video, A means of converting detected events into natural language, A method for indexing events within a video based on verbalized information, A means of providing an interface for users to view and edit indexed events, A means of automatically generating reports using edited information, A system that includes this.

2. The system according to claim 1, wherein an object detection algorithm is used in event detection.

3. The system according to claim 1, which exports a report in PDF or HTML format.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A