System

The system addresses the challenge of analyzing and presenting diverse data formats by preprocessing and using generative AI and AGI to provide strategic insights in an easily understandable format, facilitating quick and effective decision-making.

JP2026019077APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024120486
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Modern systems struggle to analyze and present strategic advice effectively from diverse data formats such as text, audio, images, and video in a unified manner, lacking the ability to handle multiple formats and present insights in an easily understandable format.

Method used

A system that includes means for receiving, preprocessing, analyzing, and presenting data using generative AI and artificial general intelligence, applying different preprocessing methods based on data type and visualizing results for intuitive understanding.

Benefits of technology

Enables comprehensive analysis of various data formats to generate and present strategic advice quickly and intuitively, supporting efficient decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019077000001_ABST
    Figure 2026019077000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system, comprising: means for receiving text, audio, image, and video; means for pre-processing the received text, audio, image, and video; means for analyzing the pre-processed text, audio, image, and video using a generative AI and artificial universal intelligence; means for generating strategic advice based on the analysis; and means for presenting the generated advice to a user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Modern companies and individual investors face challenges in gaining useful insights from vast amounts of data. Furthermore, this data exists in a variety of formats (text, audio, images, and video), requiring technology to appropriately analyze each data format. However, many current systems specialize in a single data format, and are not adequately equipped to handle a variety of data formats in a unified manner and generate strategic advice. Another challenge is the difficulty of presenting the generated advice to users in an easy-to-understand manner. [Means for solving the problem]

[0005] The present invention solves the above-mentioned problems by providing a system including a means for receiving text, audio, image, and video data, a means for preprocessing the received data, a means for analyzing the preprocessed data using generative AI and artificial general intelligence, a means for generating strategic advice based on the analysis results, and a means for presenting the generated advice to a user. In particular, the system includes a technology for performing different preprocessing depending on the type of data and a technology for visualizing and presenting the analysis results to a user. This allows a user to comprehensively analyze information obtained from various data formats and quickly obtain useful strategic advice.

[0006] "Text" refers to information that includes character data.

[0007] "Audio" is sound data transmitted via sound waves.

[0008] "Image" refers to still image data that contains visual information.

[0009] "Moving images" are dynamic data of visual information that includes the passage of time.

[0010] "Data" is a collection of information for a computer to process.

[0011] The "receiving means" is a means for acquiring data sent from a user.

[0012] The "preprocessing means" is a means for processing received data so that it is easy to analyze.

[0013] "Means for analysis" refers to means for analyzing data using generative AI and artificial general intelligence.

[0014] "Generative AI" is artificial intelligence used to analyze data and generate advice.

[0015] "Artificial general intelligence (AGI)" is an advanced artificial intelligence that can perform a wide variety of tasks at the same level as humans.

[0016] "Strategic advice" is specific guidance provided to users based on the analysis results.

[0017] The "presentation means" is a means for displaying the generated advice to the user in an easy-to-understand manner.

[0018] "System" refers to a computer-based technical device that includes the above means and functions as a whole. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11]FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0027] [First embodiment]

[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0040] This invention relates to a system that analyzes text, audio, image, and video data to generate and present strategic advice. This system is composed of components such as a server, a terminal, and a user, and each component fulfills its role to realize the overall function of the invention.

[0041] Specific system configuration and operation

[0042] 1. Data Collection

[0043] Users upload data in various formats to the system, such as market analysis reports and interview videos.

[0044] The server receives the uploaded data and stores it in a specific directory.

[0045] 2. Data Preprocessing

[0046] The server processes the received data using different pre-processing methods: text data is formatted using natural language processing (NLP), audio data is converted to text using speech recognition technology, and image and video data is analyzed using image recognition technology.

[0047] For example, it extracts text from PDFs of market analysis reports uploaded by users and converts audio from interview videos into text.

[0048] 3. Data Analysis

[0049] The server then analyzes the pre-processed data using generative AI and artificial general intelligence (AGI) to discover important patterns, trends, and potential opportunities.

[0050] For example, market trends can be identified from extracted text data, and consumer opinions can be analyzed based on the content of interviews.

[0051] 4. Generating strategic advice

[0052] The server generates specific strategic advice based on the analysis results, including strategy suggestions based on market trends and the construction of marketing messages.

[0053] For example, the analysis results can be used to propose new market segments and generate effective messages for target customers.

[0054] 5. Presentation to the User

[0055] The server provides a means to visually present the generated advice to the user, who can then access this information using their terminal and view the results in an easy to understand format.

[0056] For example, users can view suggested market strategies and analysis results on the dashboard.

[0057] Specific examples

[0058] Let's say a startup company is planning to launch a new product into the market. This company needs to understand the competitive situation and develop an effective marketing strategy. To do this, they use this system.

[0059] 1. Users upload market analysis reports (PDF) and consumer interview videos.

[0060] 2. The server receives this data and performs preprocessing by extracting text from the PDF and converting audio from the video to text.

[0061] 3. The server uses generative AI and AGI to analyze the data and analyse the competitive landscape and consumer opinion.

[0062] 4. The server uses the analysis results to develop new market segments and generate effective marketing messages.

[0063] 5. Users access the dashboard from their devices, review the proposed strategies and analysis results, and make decisions based on the results.

[0064] As described above, the present invention provides a system that processes a variety of data formats in an integrated manner and provides useful strategic insights. Users can easily upload data and obtain analytical results in an intuitive format, supporting quick and appropriate decision-making.

[0065] The processing flow will be explained below.

[0066] Step 1:

[0067] A user uploads data such as text, audio, images, and video to the system.

[0068] Example of specific operation: A user uploads a market analysis report (PDF) and an interview video.

[0069] Step 2:

[0070] The server receives the uploaded data and stores it in a specific directory.

[0071] Example of specific operation: The server stores PDF files and video files in a designated folder.

[0072] Step 3:

[0073] The server initiates pre-processing based on the type of data.

[0074] Examples of how it works: Extracting text from PDF files and separating audio from video files.

[0075] Step 4:

[0076] The server uses speech recognition technology to convert the voice data into text.

[0077] Example of specific operation: Use a speech recognition engine to convert the audio of an interview video into text.

[0078] Step 5:

[0079] The server uses image recognition technology to extract important features from image and video data.

[0080] Example of specific operation: Extract key frames from a video file and analyze them using an image recognition model.

[0081] Step 6:

[0082] The server analyzes the pre-processed data using generative AI and artificial general intelligence.

[0083] Specific operation example: Apply natural language processing (NLP) technology to text data to perform keyword extraction and topic modeling.

[0084] Step 7:

[0085] The server synthesizes the analysis results and derives relevant insights.

[0086] Example of how it works: Integrate data from market analysis reports with text data from interviews to identify common trends and patterns.

[0087] Step 8:

[0088] The server generates strategic advice based on the analysis results.

[0089] Example of specific operation: Propose ways to enter new market segments based on market trends.

[0090] Step 9:

[0091] The server visualizes the generated advice and presents it to the user.

[0092] Example of how it works: Displays strategic advice and analytical results in graphs and charts in a dashboard format.

[0093] Step 10:

[0094] The terminal displays a dashboard to the user, allowing the user to review the results.

[0095] Example of how it works: A user opens a web browser and views the dashboard with suggested market segments and measures.

[0096] Example 1

[0097] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0098] There is a need for a method to process diverse data formats (text, audio, images, video) in a unified manner and obtain information that will enable users to make strategic decisions quickly and accurately, but current technology lacks an efficient and effective way to achieve this. Therefore, there is a need for a system that can analyze diverse data formats and provide effective strategic advice.

[0099] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0100] In this invention, the server includes means for receiving data, means for preprocessing the received data based on different formats, means for analyzing the preprocessed data using generative artificial intelligence and general artificial intelligence, means for generating strategic advice based on the analysis results, and means for visually presenting the generated advice to the user. This makes it possible to extract useful information from data in various formats and provide strategic advice in a form that the user can intuitively understand.

[0101] The "means for receiving data" is a function for importing various data formats such as text, audio, images, and video provided by the user into the server.

[0102] The "means for preprocessing based on different formats" is a function for applying appropriate natural language processing technology, speech recognition technology, and image recognition technology depending on the format of the received data, and converting the data into an analyzable state.

[0103] "Generative AI" refers to algorithms and their implementations that learn large amounts of data and perform highly accurate predictions and generation for a variety of tasks.

[0104] "General artificial intelligence" is a general-purpose artificial intelligence technology that does not depend on specific tasks and can autonomously solve a variety of problems.

[0105] "Means for analysis" refers to processing functions that extract patterns and identify trends from preprocessed data to derive useful information contained in the data.

[0106] The "means for generating strategic advice" is a function for specifically proposing feasible strategies and policies to users based on the analysis results.

[0107] "Visual presentation means" refers to a function for displaying the generated strategic advice in a form that can be intuitively understood by the user, including in the form of graphs or dashboards.

[0108] This invention relates to a system that analyzes text, audio, image, and video data to generate and present strategic advice. This system is composed of components such as a server, a terminal, and a user, and each component fulfills its role to realize the overall function of the invention.

[0109] Data collection

[0110] Users can upload various types of data, such as market analysis reports and interview videos, to the system by using the file selection function on the browser to upload the data files.

[0111] The server receives these data and saves them in the specified directory. Uploaded files are sent to the server via HTTP POST request and saved in a specific folder on the server (e.g., / uploaded_data).

[0112] Data Preprocessing

[0113] The server performs preprocessing on the received data. Text data is formatted using natural language processing (NLP) technology, voice data is converted to text using voice recognition technology, and image and video data is analyzed using image recognition technology. For example, the following technologies are used:

[0114] Extract text from PDFs: Uses Python's pdfplumber library.

[0115] Converts voice data to text: Uses Google's Speech-to-Text API.

[0116] Image and video analysis: Uses OpenCV and TensorFlow models.

[0117] Data analysis

[0118] The server uses generative artificial intelligence (generative AI) and artificial general intelligence (AGI) to analyze the pre-processed data. This analysis uncovers important patterns, trends, and potential opportunities. Specifically, it uses the following techniques:

[0119] BERT model: Used for keyword extraction and sentiment analysis of preprocessed text data.

[0120] GPT-4 model: Used to analyze the context of the interview content.

[0121] Strategic Advice Generation

[0122] The server generates specific strategic advice based on the analysis results, including strategy proposals and marketing message construction based on market trends. For example, it uses GPT-4 to generate strategy proposals based on detected trends. It can also take a context-based approach by taking into account the user's past success stories.

[0123] Presenting to the user

[0124] The server provides a means to visually present the generated advice to the user, who can then access this information using their device and view the results in an easy-to-understand format, for example by displaying the results in a dashboard using a JavaScript library (e.g., D3.js or Chart.js).

[0125] Specific examples

[0126] Consider a startup company planning to launch a new product in the market. The company needs to understand the competitive landscape and develop an effective marketing strategy.

[0127] 1. The user uploads the market analysis report (PDF) and consumer interview video to the system. This is done using the file selection function on the browser.

[0128] 2. The server receives these data, extracts text from PDFs, and converts audio from videos to text. For example, the pdfplumber library extracts text from PDFs and Google's Speech-to-Text API converts audio to text.

[0129] 3. The server uses generative AI and AGI to analyze the competitive landscape and consumer opinions, using the BERT model for market analysis and the GPT-4 model for detailed analysis of interview content.

[0130] 4. The server uses the analysis results to develop strategies for entering new market segments and generating effective marketing messages. It uses GPT-4 to generate strategic recommendations based on the trends detected.

[0131] 5. Users access the dashboard from their device to view the proposed strategies and analysis results. The dashboard displays data visualized using D3.js and Chart.js.

[0132] Based on this example, here are some example prompts:

[0133] "Analyze market analysis reports and consumer interview videos regarding the launch of new products, and propose effective marketing strategies based on the situation of competitors and consumer opinions."

[0134] Through the above process, useful information can be extracted from a variety of data formats, enabling users to make quick and appropriate decisions.

[0135] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0136] Step 1: Upload your data

[0137] Subject: User

[0138] Input: Market analysis report (PDF), consumer interview video (MP4, etc.)

[0139] Output: Data file uploaded to the server

[0140] Specific operation: The user uses the file selection function on the browser to select the required data file and upload it to the system, where the data is sent to the server and received.

[0141] Step 2: Receiving and storing data

[0142] Subject: Server

[0143] Input: User uploaded data file

[0144] Output: Data files saved in a specific directory

[0145] Specific operation: The server receives the HTTP POST request and saves the uploaded file in a pre-specified folder (e.g., / uploaded_data), allowing data to be managed centrally.

[0146] Step 3: Data preprocessing (text data)

[0147] Subject: Server

[0148] Input: Saved text data file (PDF)

[0149] Output: Preprocessed text data

[0150] What it does: It uses the Python pdfplumber library to extract text from PDF files and saves the extracted text in a text file or database.

[0151] Step 4: Data preprocessing (audio data)

[0152] Subject: Server

[0153] Input: Saved audio data file (MP4)

[0154] Output: Preprocessed text data

[0155] What it does: The audio data is processed using Google's Speech-to-Text API, which converts the audio into text, which is then saved in a text file or database.

[0156] Step 5: Data preprocessing (image and video data)

[0157] Subject: Server

[0158] Input: Saved image or video data files (JPG, MP4, etc.)

[0159] Output: Preprocessed image recognition data

[0160] Specific operation: Analyzes images and videos using OpenCV and TensorFlow models, for example, performs tasks such as object detection and face recognition, and saves the analysis results.

[0161] Step 6: Data analysis (text data)

[0162] Subject: Server

[0163] Input: Preprocessed text data

[0164] Output: Analyzed patterns and trends

[0165] Specific operation: Using the BERT model, keyword extraction and sentiment analysis are performed. The analysis results are stored in a database.

[0166] Step 7: Data analysis (audio data)

[0167] Subject: Server

[0168] Input: Preprocessed text data (converted from audio)

[0169] Output: Parsed opinions and feedback

[0170] How it works: Using the GPT-4 model, the voice data is analyzed in detail to extract consumer opinions and feedback. The analysis results are stored in a database.

[0171] Step 8: Generate strategic advice

[0172] Subject: Server

[0173] Input: Analysis results (text data, audio data, image and video data)

[0174] Output: Strategic advice

[0175] How it works: Generative AI is used to generate strategic advice based on the analysis results. For example, GPT-4 is used to create strategic proposals based on market trends. The generated advice is then stored in a database.

[0176] Step 9: Present to the user

[0177] Subject: Server

[0178] Input: Generated strategic advice

[0179] Output: Visualized data (dashboard)

[0180] Specific behavior: Strategic advice is visualized and displayed on a dashboard using JavaScript libraries (e.g., D3.js, Chart.js). Users can access the dashboard from their devices and check the visualized results.

[0181] (Application example 1)

[0182] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0183] In autonomous vehicles, it is difficult to efficiently analyze large amounts of diverse data, such as operational data, voice instructions, and camera footage, and provide safe and optimal driving routes. In particular, it is necessary to generate and present strategic advice in real time, which will help prevent accidents and ensure efficient route selection.

[0184] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0185] In this invention, the server includes means for receiving text, audio, image, and video data, means for preprocessing the received data, means for analyzing the preprocessed data using generative AI and artificial general intelligence, means for generating strategic advice based on the analysis results, means for presenting the generated advice to a user, means for receiving camera footage, sensor data, and voice instructions from an autonomous vehicle, means for analyzing road conditions using the received data, and means for proposing a safe driving route based on the analysis results, thereby enabling the autonomous vehicle to select a safe and optimal driving route in real time, prevent accidents, and drive efficiently.

[0186] "Text" is a data format that expresses information using characters and symbols.

[0187] "Voice" refers to data relating to human speech or acoustics transmitted through sound waves.

[0188] An "image" is a still image that represents visual information in pixels.

[0189] "Video" refers to dynamic visual data that expresses movement using successive image frames.

[0190] "Preprocessing" refers to the initial processing of received data to convert it into a form suitable for analysis.

[0191] "Generative AI" is an artificial intelligence technology that learns from large datasets and generates new data.

[0192] "Artificial general intelligence" refers to advanced artificial intelligence that has the ability to solve a wide range of problems, not just specific tasks.

[0193] "Strategic advice" refers to recommendations based on data analysis that provide users with optimal courses of action.

[0194] An "autonomous vehicle" is a vehicle that has the ability to operate autonomously using artificial intelligence and various sensors.

[0195] "Camera footage" refers to video data captured by a camera and recorded as visual information.

[0196] "Sensor data" refers to data about physical phenomena measured by various sensors.

[0197] "Voice instructions" refer to commands or instructions given using voice.

[0198] "Road conditions" refers to environmental information that affects the safety and efficiency of driving, such as road congestion and road surface conditions.

[0199] A "travel route" is the optimal route to a destination.

[0200] The present invention is a system for analyzing operational data of autonomous vehicles to support safe and efficient operation. This system analyzes text, audio, image, and video data to generate and present strategic advice. Specific embodiments of this system are described below.

[0201] System configuration and operation

[0202] This system is mainly composed of components such as servers, terminals, and users, each of which fulfills its respective role to realize the overall function.

[0203] Server Roles

[0204] The server is the core of this system and has the following functions:

[0205] Data reception and preprocessing:

[0206] The server receives text, audio, image, and video data and pre-processes this data: for example, camera footage is processed frame by frame using OpenCV, and audio data is converted to text using the SpeechRecognition library.

[0207] Data Analysis:

[0208] The pre-processed data is then analyzed using deep learning models (generative AI and artificial general intelligence), for example, to detect road conditions and vehicle anomalies using deep learning frameworks such as TensorFlow and Keras.

[0209] Strategic advice generation:

[0210] Based on the analysis results, strategic advice is generated, such as suggesting safe driving routes and generating warning messages for drivers.

[0211] Present your proposal:

[0212] The server sends the generated advice to the terminal and presents it visually to the user. Operation information and advice are displayed in real time in a dashboard format.

[0213] Device Role

[0214] A terminal is a device that is directly operated by the user, such as a smartphone or tablet. A terminal has the following functions:

[0215] Data entry and collection:

[0216] It sends camera footage, sensor data, and voice instructions uploaded by the user to the server.

[0217] View the advice offered:

[0218] The terminal visually displays strategic advice sent from the server, allowing users to see driving routes and warning messages in real time.

[0219] User Roles

[0220] The user is the entity that uses the system and performs the following operations:

[0221] Upload data:

[0222] Users upload a variety of data to the system, including market analysis reports and interview videos.

[0223] Review the suggested advice:

[0224] Users can view the strategic advice and driving routes generated through a dashboard on their device and make decisions based on the results.

[0225] Examples of specific examples and prompts

[0226] When an autonomous vehicle is driving in a city and analyzing camera footage to detect obstacles and pedestrians ahead, the server uses the following prompt sentence:

[0227] Camera video analysis prompt: "Analyze the video data obtained from the vehicle's front camera to detect obstacles on the road."

[0228] When the vehicle selects a particular route based on the driver's voice instructions, the server uses the following prompt sentence:

[0229] Voice command analysis prompt: "Please convert the driver's current voice commands into text and analyze their actions."

[0230] This will enable autonomous vehicles to select safe and optimal routes in real time, prevent accidents, and operate efficiently.

[0231] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0232] Step 1:

[0233] The server receives camera footage, sensor data, and audio instructions from the user.

[0234] Input: Camera video files, sensor data files, audio files

[0235] Specific operation: Captures camera images as frame data, receives sensor data as numerical data, and recognizes audio files as audio data.

[0236] Step 2:

[0237] The server pre-processes the received data.

[0238] Input: Camera image frame data, sensor data values, audio data

[0239] Specific operations: Camera image frames are resized and preprocessed using OpenCV, sensor data is normalized to an appropriate range, and audio data is converted to text using the SpeechRecognition library.

[0240] Output: Preprocessed video data, normalized sensor data, and transcribed audio data

[0241] Step 3:

[0242] The server analyzes the pre-processed data using generative AI models and artificial general intelligence.

[0243] Input: Preprocessed video data, normalized sensor data, and transcribed audio data

[0244] Specific operation: Preprocessed video data is input into deep learning models (TensorFlow and Keras) to analyze road conditions, sensor data is analyzed, anomaly detection algorithms are applied, and voice data is analyzed to determine instructions.

[0245] Output: Analyzed road conditions, abnormal sensor data, voice instructions

[0246] Step 4:

[0247] The server generates strategic advice based on the analysis results.

[0248] Input: Analyzed road conditions, abnormal sensor data, voice instructions

[0249] Specific operation: Based on the analysis results, a strategic algorithm is applied to generate safe driving routes and warning messages for drivers.

[0250] Output: Generated route suggestions and warning messages for the driver

[0251] Step 5:

[0252] The server transmits the generated advice to the terminal.

[0253] Input: Generated route suggestions, warning messages for drivers

[0254] Specific operation: The generated advice is sent to the terminal and displayed on the terminal in real time.

[0255] Output: Route suggestions displayed on the terminal, warning messages to the driver

[0256] Step 6:

[0257] The terminal visually presents the generated advice to the user.

[0258] Input: Route suggestions displayed on the device, warning messages to the driver

[0259] Specific operation: Routes and warning messages are presented to the user in real time using a dashboard-style UI.

[0260] Output: User-visible route and warning messages

[0261] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0262] This invention relates to a system that generates and presents strategic advice by analyzing text, voice, image, and video data and combining it with an emotion engine that recognizes the user's emotions. This system is composed of components such as a server, a terminal, and a user, and each component fulfills its role to realize the overall function of the invention.

[0263] Specific system configuration and operation

[0264] 1. Data Collection

[0265] A user uploads data such as text, audio, images, and video to the system.

[0266] The server receives the uploaded data and stores it in a specific directory.

[0267] 2. Data Preprocessing

[0268] The server processes the received data using different pre-processing methods: for example, text data is formatted using natural language processing (NLP), audio data is converted to text using speech recognition technology, and image and video data is analyzed using image recognition technology.

[0269] 3. Data Analysis

[0270] The server then analyzes the pre-processed data using generative AI and artificial general intelligence (AGI) to discover important patterns, trends, and potential opportunities.

[0271] 4. Emotional Recognition

[0272] The server uses an emotion engine to recognize and analyze the user's emotions, for example, identifying the user's emotional state (e.g., joy, sadness, anger) from the user's voice data and text data.

[0273] 5. Generating strategic advice

[0274] The server generates specific strategic advice based on the analysis results, and provides advice that also takes into account the results of the user's sentiment analysis.

[0275] 6. Presentation to the User

[0276] The server provides a means to visually present the generated advice to the user, who can then access this information using their terminal and view the results in an easy to understand format.

[0277] Specific examples

[0278] For example, suppose a startup company is planning to launch a new product on the market. This company needs to understand the competitive situation and develop an effective marketing strategy. To do this, they use this system.

[0279] 1. The user uploads the market analysis report (PDF) and video of the consumer interview to the system.

[0280] 2. The server receives these data, extracts the text data, and converts the audio from the video to text.

[0281] 3. The server uses generative AI and AGI to analyze the data and analyse the competitive landscape and consumer opinion.

[0282] 4. The server uses an emotion engine to recognize and analyze consumer emotions from the interview audio data, for example, identifying which products consumers have positive emotions about.

[0283] 5. Based on the analysis, the server generates strategic advice that takes into account the user's emotional state, including how to enter new market segments and effective marketing messages to target customers.

[0284] 6. Users access the dashboard using their devices, review the proposed market strategies and analysis results, and make decisions based on the results.

[0285] In this way, the present invention provides a system that processes a variety of data formats in an integrated manner and provides useful strategic insights that take user sentiment into account. Users can easily upload data and receive analysis results in an intuitively understandable format, supporting quick and appropriate decision-making.

[0286] The processing flow will be explained below.

[0287] Step 1:

[0288] A user uploads data such as text, audio, images, and video to the system.

[0289] Example of specific operation: A user uploads a market analysis report (PDF) and a consumer interview video.

[0290] Step 2:

[0291] The server receives the uploaded data and stores it in a specific directory.

[0292] Example of specific operation: The server stores PDF files and video files in a designated folder.

[0293] Step 3:

[0294] The server initiates pre-processing based on the type of data.

[0295] Examples of how it works: Extracting text from PDF files and separating audio from video files.

[0296] Step 4:

[0297] The server uses speech recognition technology to convert the voice data into text.

[0298] Example of specific operation: Use a speech recognition engine to convert the audio of an interview video into text.

[0299] Step 5:

[0300] The server uses image recognition technology to extract important features from image and video data.

[0301] Example of specific operation: Extract key frames from a video file and analyze them using an image recognition model.

[0302] Step 6:

[0303] The server analyzes the pre-processed data using generative AI and artificial general intelligence.

[0304] Specific operation example: Apply natural language processing (NLP) technology to text data to perform keyword extraction and topic modeling.

[0305] Step 7:

[0306] The server uses an emotion engine to recognize and analyze the user's emotions.

[0307] Specific operation example: Identify the user's emotional state from voice data and text data, and determine joy, sadness, anger, etc.

[0308] Step 8:

[0309] The server synthesizes the analysis results and derives relevant insights.

[0310] Example of how it works: Integrate data from market analysis reports with text data from interviews to identify common trends and patterns.

[0311] Step 9:

[0312] The server generates strategic advice based on the analysis results.

[0313] Example of specific operation: Based on market trends, suggest ways to enter new market segments and create marketing messages that take into account the user's emotional state.

[0314] Step 10:

[0315] The server visualizes the generated advice and presents it to the user.

[0316] Example of how it works: Displays strategic advice and analytical results in graphs and charts in a dashboard format.

[0317] Step 11:

[0318] The terminal displays a dashboard to the user, allowing the user to review the results.

[0319] Example of how it works: A user opens a web browser and views the dashboard with suggested market segments and measures.

[0320] Example 2

[0321] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0322] Conventional data analysis systems lack the ability to process multiple data formats (text, audio, images, and videos) in a unified manner, making it particularly difficult to generate strategic advice that takes into account user sentiment analysis. Furthermore, they lacked a means to provide analysis results to users in an easy-to-understand format, making it difficult for users to make appropriate decisions.

[0323] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0324] In this invention, the server includes means for receiving text, audio, image, and video data, means for preprocessing the received data according to the type of data, means for analyzing the preprocessed data using a generation AI and an intelligent system, means for recognizing and analyzing a user's emotions, means for generating strategic advice based on the analysis results and emotion recognition results, and means for visually presenting the generated advice to the user. This enables unified analysis of multiple data formats and the provision of strategic advice that takes the user's emotions into consideration. Furthermore, visual presentation of the analysis results allows the user to easily understand and make appropriate decisions.

[0325] "Text data" is information expressed in the form of characters and sentences.

[0326] "Audio data" is information that records human voices and sounds in digital format.

[0327] "Image data" is visual information such as photographs and illustrations recorded in digital format.

[0328] "Video data" means a digital recording that combines moving visual and audio information.

[0329] "Means for receiving" refers to a technique or device for capturing data sent from a user within the system.

[0330] A "preprocessing means" is a technique or device for shaping or transforming collected data so that it is easier to analyze.

[0331] "Generative AI" refers to algorithms or systems that use artificial intelligence techniques to generate new information from data.

[0332] An "intelligent system" is an artificial intelligence system that has general knowledge and uses it to solve problems and analyze data.

[0333] "Emotion recognition and analysis means" refers to a technique or device for identifying and analyzing a user's emotional state from data provided by the user.

[0334] The "means for generating strategic advice" is a technology or device for generating guidelines or suggestions that are useful to the user based on the analysis results and emotion recognition results.

[0335] "Means for visually presenting to the user" refers to a technique or device for providing the analysis results and advice to the user in a visual format such as a graph or chart.

[0336] This invention relates to a system that generates and presents strategic advice by analyzing text, voice, image, and video data and combining it with an emotion engine that recognizes the user's emotions. This system is composed of components such as a server, a terminal, and a user, each of which fulfills a specific role to realize the overall function of the invention.

[0337] Data collection

[0338] Users upload data such as text, audio, images, and videos to the system. An interface for uploading this data is provided on the user's terminal. For example, a file can be selected by dragging and dropping using the file upload function of a web browser.

[0339] Data Preprocessing

[0340] The server receives the uploaded data and stores it in a specific directory. The received data is classified by type and the following preprocessing is performed depending on the data format:

[0341] Text data: A natural language processing (NLP) engine (e.g., SpaCy) is used to check grammar and filter out unwanted words.

[0342] Voice Data: We use a speech recognition engine (e.g., Google Speech-to-Text) to convert your voice into text.

[0343] Image data: An image recognition system (e.g., OpenCV) is used to identify objects in the image and generate metadata.

[0344] Video data: The video is broken down into frames, and each frame is analyzed using image recognition technology.

[0345] Data analysis

[0346] The server uses generative AI models (e.g., GPT-4) and intelligent systems to analyze the pre-processed data, allowing it to discover important patterns, trends, and potential opportunities, such as extracting competitor trends and identifying market niches based on market analysis reports.

[0347] Emotion recognition

[0348] The server uses an emotion recognition engine (e.g., IBM Watson Natural Language Understanding) to recognize and analyze emotions from the data provided by the user. It can identify a range of emotional states (such as joy, sadness, and anger) from text and audio data. For example, it can analyze the emotional responses from audio data of consumer interview videos to determine which products have the most positive opinions.

[0349] Strategic Advice Generation

[0350] The server generates strategic advice for users based on the analysis results and emotion recognition results. Specific suggestions are made using a generative AI model (e.g., GPT-4). For example, it can generate a suggestion such as, "Targeting a new product specifically for women in their 20s is likely to increase market share."

[0351] Presenting to the user

[0352] The server provides a dashboard to visually present the generated advice to the user. Users can access this dashboard from their own devices and view the analysis results in an intuitive, easy-to-understand format. The dashboard contains visual elements such as graphs, tables, and heat maps, and users can click on results to access more detailed information.

[0353] Specific examples

[0354] For example, suppose a startup company is planning to launch a new product. This company needs to understand the competitive situation and develop an effective marketing strategy. To do this, they use this system.

[0355] 1. The user uploads the market analysis report (PDF) and video of the consumer interview to the system.

[0356] 2. The server receives these data, extracts the text data, and converts the audio from the video to text.

[0357] 3. The server uses generative AI and intelligence systems to analyze the data and analyse competitors' situations and consumer opinions.

[0358] 4. The server uses an emotion recognition engine to recognize and analyze the consumer's emotions from the interview audio data, for example, to identify which products the consumer has positive feelings about.

[0359] 5. Based on the analysis, the server generates strategic advice that takes into account the user's emotional state, including how to enter new market segments and effective marketing messages to target customers.

[0360] 6. Users access the dashboard using their devices, review the proposed market strategies and analysis results, and make decisions based on the results.

[0361] Prompt Sentence Examples

[0362] "Based on this PDF report and video interviews, you can develop a market strategy and analyze consumer sentiment."

[0363] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0364] Step 1: Collect data

[0365] Input: User-uploaded text, audio, image, and video data

[0366] Specific operation: A user logs in to the system from their own terminal and uses the file upload function to upload data such as text, audio, images, and videos. Files are selected by drag-and-drop operation.

[0367] Output: Data stored on the server

[0368] Step 2: Save the data to the appropriate directory

[0369] Input: User uploaded data

[0370] Specific operation: The server receives the data uploaded by the user and saves it in a specific directory in real time. At this time, the data is temporarily stored in a buffer and then transferred to permanent storage after verification.

[0371] Output: Data stored in a specific directory on the server

[0372] Step 3: Preprocessing the data

[0373] Input: Raw data stored on the server

[0374] Specific operation: The server classifies the received data by type and performs appropriate preprocessing for each type. For example:

[0375] Text data: Use a natural language processing (NLP) engine (e.g., SpaCy) to check grammar and filter unnecessary words.

[0376] Voice Data: Converts speech to text using a speech recognition engine (e.g., Google Speech-to-Text).

[0377] Image data: An image recognition system (e.g., OpenCV) is used to identify objects in the image and generate metadata.

[0378] Video data: The video is broken down into frames, and each frame is analyzed using image recognition technology.

[0379] Output: Preprocessed data

[0380] Step 4: Data analysis

[0381] Input: Preprocessed data

[0382] What it does: The server analyzes the pre-processed data using generative AI models (e.g., GPT-4) and intelligent systems to discover important patterns, trends, and potential opportunities in the data.

[0383] Output: Analysis results (patterns, trends, potential opportunities)

[0384] Step 5: Recognize emotions

[0385] Input: Preprocessed text and audio data

[0386] Specific operation: The server uses an emotion recognition engine (e.g., IBM Watson Natural Language Understanding) to recognize and analyze emotions from the data provided by the user, and identifies a range of emotional states from the voice and text data.

[0387] Output: Emotion recognition results (happiness, sadness, anger, etc.)

[0388] Step 6: Generate strategic advice

[0389] Input: Analysis results and emotion recognition results

[0390] Specific operation: The server uses the generative AI model to generate specific strategic advice based on the analysis and emotion recognition results. For example, it generates a suggestion such as, "Targeting a new product specifically for women in their 20s is likely to increase market share."

[0391] Output: Strategic advice

[0392] Step 7: Present to the user

[0393] Input: Generated strategic advice

[0394] Specific operation: The server provides a dashboard to visually present the generated advice to the user. The user accesses this dashboard using their own device and understands the analysis results through visual elements such as graphs and charts.

[0395] Output: Visual analysis results and advice on a user-accessible dashboard

[0396] (Application example 2)

[0397] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0398] In typical brick-and-mortar stores, it is difficult for staff to accurately grasp customer emotions in real time and provide optimal product recommendations and services accordingly. Furthermore, staff need training and experience to deal with a wide variety of customers individually, and a lack of this training can lead to a decline in customer satisfaction. For this reason, there is a demand for support tools that enable staff to deal with customers quickly and accurately.

[0399] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving text, audio, image, and video data, means for preprocessing the received data, means for analyzing the preprocessed data using generative AI and artificial general intelligence, means for generating strategic advice based on the analysis results, means for presenting the generated advice to the user, means for recognizing emotions from the user's voice and video, and means for presenting advice generated in real time taking into account the emotion recognition results. This enables staff at physical stores to grasp customer emotions in real time and quickly provide optimal product suggestions and services based on that.

[0400] definition statement

[0401] "Text data" is data made up of character information.

[0402] "Audio data" refers to information collected as sound waves recorded in digital format.

[0403] "Image data" is data that stores visual information as a still image.

[0404] "Moving image data" is data that expresses dynamic visual information by combining successive still images along a time axis.

[0405] "Preprocessing" is the initial data processing to convert raw data into an analyzable format.

[0406] "Generative AI" is an artificial intelligence technology that generates new data and information based on input data.

[0407] "Artificial general intelligence" is artificial intelligence with a wide range of capabilities that can handle a wide variety of tasks.

[0408] An "emotion engine" is a technology that recognizes and analyzes a user's emotional state from voice, facial expressions, etc.

[0409] "Strategic advice" refers to specific and effective suggestions and instructions based on the results of analysis.

[0410] "Real-time presentation" means that analysis and advice are performed instantly, and information is provided to the user immediately.

[0411] MODE FOR CARRYING OUT THE INVENTION

[0412] The following describes an embodiment of the present invention.

[0413] The system includes means for receiving text, audio, image, and video data, means for preprocessing the received data, means for analyzing the preprocessed data using generative AI and artificial general intelligence, means for generating strategic advice based on the analysis results, means for presenting the generated advice to a user, means for recognizing emotions from the user's voice and video, and means for presenting the generated advice in real time, taking into account the emotion recognition results.

[0414] Data collection

[0415] Users use their smartphones to collect video and audio data of customers in physical stores via a camera and microphone, and these video and audio data are converted into text and image data.

[0416] Data Preprocessing

[0417] The server preprocesses the received video and audio data. Video data is converted into still images using an image recognition algorithm, and audio data is converted into text using speech recognition technology. Libraries and APIs such as OpenCV, Google Speech Recognition, and Emotion Recognition are used for preprocessing.

[0418] Data analysis

[0419] The pre-processed data is then analyzed using generative AI and artificial general intelligence to identify the products customers are interested in and their emotional state (interest, delight, doubt, etc.). The analysis is performed using OpenAI's API.

[0420] Emotion recognition

[0421] The user's emotional state is recognized from audio and video data. Emotion Recognition technology identifies the customer's emotional state and complements strategic advice based on this information.

[0422] Strategic Advice Generation

[0423] The server generates optimal strategic advice for the customer based on the analysis results and emotion recognition results, including product suggestions and service delivery methods that take into account the customer's emotions.

[0424] Presenting to the user

[0425] Strategic advice is presented to users in real time via a smartphone application, allowing wait staff to take immediate and appropriate action.

[0426] Specific examples

[0427] For example, if a customer asks for details about a particular product, the system might:

[0428] Customer question: "What are the features of this camera?"

[0429] Emotions recognized from customer facial expressions: Interest

[0430] Strategic advice: "This camera takes high-resolution photos and is waterproof. Plus, other customers often buy it."

[0431] Prompt Sentence Examples

[0432] Generate advice based on the user's statements and feelings below.

[0433] Say: "What is special about this camera?"

[0434] Emotion: "Interest"

[0435] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0436] Program processing steps

[0437] Step 1:

[0438] Users use the camera and microphone on their smartphones to collect video and audio data of customers in physical stores. The collected data is input as raw video frames and audio waveform data.

[0439] Step 2:

[0440] The terminal transmits the collected data to a server, which receives the video frames and audio waveform data, stores the video data as an image file, and stores the audio data as an audio file.

[0441] Step 3:

[0442] The server pre-processes the received data: video data is converted into still images using the OpenCV library, and audio data is converted into text using the Google Speech Recognition API. This pre-processing converts video data into an image format and audio data into a text format.

[0443] Step 4:

[0444] The server analyzes the preprocessed data using generative AI and artificial general intelligence (AGI). For the analysis, OpenAI's API is used to extract visual features from image data and customer interests and concerns from text data. This results in data feature quantities.

[0445] Step 5:

[0446] The server uses Emotion Recognition technology to recognize emotions from the user's video and audio data. It uses a facial expression recognition algorithm (e.g., OpenFace) from the video data and a voice emotion recognition algorithm from the audio data. This recognition yields the user's emotional state (interest, joy, etc.).

[0447] Step 6:

[0448] The server generates strategic advice based on the analysis results and emotion recognition results. It uses OpenAI's API to integrate this data and generate prompts. Based on these prompts, the server automatically generates advice that includes optimal product recommendations and service delivery methods for the customer.

[0449] Step 7:

[0450] The server then sends the generated strategic advice to the terminal, which then displays the received advice in real time on the smartphone screen so that the customer service staff can immediately check it.

[0451] Step 8:

[0452] The user (customer service staff) can then provide appropriate product recommendations and services to customers based on the strategic advice provided, which will enable quick and accurate customer service and is expected to improve customer satisfaction.

[0453] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0454] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0455] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0456] [Second embodiment]

[0457] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0458] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0459] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0460] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0461] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0462] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0463] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0464] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0465] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0466] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0467] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0468] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0469] This invention relates to a system that analyzes text, audio, image, and video data to generate and present strategic advice. This system is composed of components such as a server, a terminal, and a user, and each component fulfills its role to realize the overall function of the invention.

[0470] Specific system configuration and operation

[0471] 1. Data Collection

[0472] Users upload data in various formats to the system, such as market analysis reports and interview videos.

[0473] The server receives the uploaded data and stores it in a specific directory.

[0474] 2. Data Preprocessing

[0475] The server processes the received data using different pre-processing methods: text data is formatted using natural language processing (NLP), audio data is converted to text using speech recognition technology, and image and video data is analyzed using image recognition technology.

[0476] For example, it extracts text from PDFs of market analysis reports uploaded by users and converts audio from interview videos into text.

[0477] 3. Data Analysis

[0478] The server then analyzes the pre-processed data using generative AI and artificial general intelligence (AGI) to discover important patterns, trends, and potential opportunities.

[0479] For example, market trends can be identified from extracted text data, and consumer opinions can be analyzed based on the content of interviews.

[0480] 4. Generating strategic advice

[0481] The server generates specific strategic advice based on the analysis results, including strategy suggestions based on market trends and the construction of marketing messages.

[0482] For example, the analysis results can be used to propose new market segments and generate effective messages for target customers.

[0483] 5. Presentation to the User

[0484] The server provides a means to visually present the generated advice to the user, who can then access this information using their terminal and view the results in an easy to understand format.

[0485] For example, users can view suggested market strategies and analysis results on the dashboard.

[0486] Specific examples

[0487] Let's say a startup company is planning to launch a new product into the market. This company needs to understand the competitive situation and develop an effective marketing strategy. To do this, they use this system.

[0488] 1. Users upload market analysis reports (PDF) and consumer interview videos.

[0489] 2. The server receives this data and performs preprocessing by extracting text from the PDF and converting audio from the video to text.

[0490] 3. The server uses generative AI and AGI to analyze the data and analyse the competitive landscape and consumer opinion.

[0491] 4. The server uses the analysis results to develop new market segments and generate effective marketing messages.

[0492] 5. Users access the dashboard from their devices, review the proposed strategies and analysis results, and make decisions based on the results.

[0493] As described above, the present invention provides a system that processes a variety of data formats in an integrated manner and provides useful strategic insights. Users can easily upload data and obtain analytical results in an intuitive format, supporting quick and appropriate decision-making.

[0494] The processing flow will be explained below.

[0495] Step 1:

[0496] A user uploads data such as text, audio, images, and video to the system.

[0497] Example of specific operation: A user uploads a market analysis report (PDF) and an interview video.

[0498] Step 2:

[0499] The server receives the uploaded data and stores it in a specific directory.

[0500] Example of specific operation: The server stores PDF files and video files in a designated folder.

[0501] Step 3:

[0502] The server initiates pre-processing based on the type of data.

[0503] Examples of how it works: Extracting text from PDF files and separating audio from video files.

[0504] Step 4:

[0505] The server uses speech recognition technology to convert the voice data into text.

[0506] Example of specific operation: Use a speech recognition engine to convert the audio of an interview video into text.

[0507] Step 5:

[0508] The server uses image recognition technology to extract important features from image and video data.

[0509] Example of specific operation: Extract key frames from a video file and analyze them using an image recognition model.

[0510] Step 6:

[0511] The server analyzes the pre-processed data using generative AI and artificial general intelligence.

[0512] Specific operation example: Apply natural language processing (NLP) technology to text data to perform keyword extraction and topic modeling.

[0513] Step 7:

[0514] The server synthesizes the analysis results and derives relevant insights.

[0515] Example of how it works: Integrate data from market analysis reports with text data from interviews to identify common trends and patterns.

[0516] Step 8:

[0517] The server generates strategic advice based on the analysis results.

[0518] Example of specific operation: Propose ways to enter new market segments based on market trends.

[0519] Step 9:

[0520] The server visualizes the generated advice and presents it to the user.

[0521] Example of how it works: Displays strategic advice and analytical results in graphs and charts in a dashboard format.

[0522] Step 10:

[0523] The terminal displays a dashboard to the user, allowing the user to review the results.

[0524] Example of how it works: A user opens a web browser and views the dashboard with suggested market segments and measures.

[0525] Example 1

[0526] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0527] There is a need for a method to process diverse data formats (text, audio, images, video) in a unified manner and obtain information that will enable users to make strategic decisions quickly and accurately, but current technology lacks an efficient and effective way to achieve this. Therefore, there is a need for a system that can analyze diverse data formats and provide effective strategic advice.

[0528] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0529] In this invention, the server includes means for receiving data, means for preprocessing the received data based on different formats, means for analyzing the preprocessed data using generative artificial intelligence and general artificial intelligence, means for generating strategic advice based on the analysis results, and means for visually presenting the generated advice to the user. This makes it possible to extract useful information from data in various formats and provide strategic advice in a form that the user can intuitively understand.

[0530] The "means for receiving data" is a function for importing various data formats such as text, audio, images, and video provided by the user into the server.

[0531] The "means for preprocessing based on different formats" is a function for applying appropriate natural language processing technology, speech recognition technology, and image recognition technology depending on the format of the received data, and converting the data into an analyzable state.

[0532] "Generative AI" refers to algorithms and their implementations that learn large amounts of data and perform highly accurate predictions and generation for a variety of tasks.

[0533] "General artificial intelligence" is a general-purpose artificial intelligence technology that does not depend on specific tasks and can autonomously solve a variety of problems.

[0534] "Means for analysis" refers to processing functions that extract patterns and identify trends from preprocessed data to derive useful information contained in the data.

[0535] The "means for generating strategic advice" is a function for specifically proposing feasible strategies and policies to users based on the analysis results.

[0536] "Visual presentation means" refers to a function for displaying the generated strategic advice in a form that can be intuitively understood by the user, including in the form of graphs or dashboards.

[0537] This invention relates to a system that analyzes text, audio, image, and video data to generate and present strategic advice. This system is composed of components such as a server, a terminal, and a user, and each component fulfills its role to realize the overall function of the invention.

[0538] Data collection

[0539] Users can upload various types of data, such as market analysis reports and interview videos, to the system by using the file selection function on the browser to upload the data files.

[0540] The server receives these data and saves them in the specified directory. Uploaded files are sent to the server via HTTP POST request and saved in a specific folder on the server (e.g., / uploaded_data).

[0541] Data Preprocessing

[0542] The server performs preprocessing on the received data. Text data is formatted using natural language processing (NLP) technology, voice data is converted to text using voice recognition technology, and image and video data is analyzed using image recognition technology. For example, the following technologies are used:

[0543] Extract text from PDFs: Uses Python's pdfplumber library.

[0544] Converts voice data to text: Uses Google's Speech-to-Text API.

[0545] Image and video analysis: Uses OpenCV and TensorFlow models.

[0546] Data analysis

[0547] The server uses generative artificial intelligence (generative AI) and artificial general intelligence (AGI) to analyze the pre-processed data. This analysis uncovers important patterns, trends, and potential opportunities. Specifically, it uses the following techniques:

[0548] BERT model: Used for keyword extraction and sentiment analysis of preprocessed text data.

[0549] GPT-4 model: Used to analyze the context of the interview content.

[0550] Strategic Advice Generation

[0551] The server generates specific strategic advice based on the analysis results, including strategy proposals and marketing message construction based on market trends. For example, it uses GPT-4 to generate strategy proposals based on detected trends. It can also take a context-based approach by taking into account the user's past success stories.

[0552] Presenting to the user

[0553] The server provides a means to visually present the generated advice to the user, who can then access this information using their device and view the results in an easy-to-understand format, for example by displaying the results in a dashboard using a JavaScript library (e.g., D3.js or Chart.js).

[0554] Specific examples

[0555] Consider a startup company planning to launch a new product in the market. The company needs to understand the competitive landscape and develop an effective marketing strategy.

[0556] 1. The user uploads the market analysis report (PDF) and consumer interview video to the system. This is done using the file selection function on the browser.

[0557] 2. The server receives these data, extracts text from PDFs, and converts audio from videos to text. For example, the pdfplumber library extracts text from PDFs and Google's Speech-to-Text API converts audio to text.

[0558] 3. The server uses generative AI and AGI to analyze the competitive landscape and consumer opinions, using the BERT model for market analysis and the GPT-4 model for detailed analysis of interview content.

[0559] 4. The server uses the analysis results to develop strategies for entering new market segments and generating effective marketing messages. It uses GPT-4 to generate strategic recommendations based on the trends detected.

[0560] 5. Users access the dashboard from their device to view the proposed strategies and analysis results. The dashboard displays data visualized using D3.js and Chart.js.

[0561] Based on this example, here are some example prompts:

[0562] "Analyze market analysis reports and consumer interview videos regarding the launch of new products, and propose effective marketing strategies based on the situation of competitors and consumer opinions."

[0563] Through the above process, useful information can be extracted from a variety of data formats, enabling users to make quick and appropriate decisions.

[0564] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0565] Step 1: Upload your data

[0566] Subject: User

[0567] Input: Market analysis report (PDF), consumer interview video (MP4, etc.)

[0568] Output: Data file uploaded to the server

[0569] Specific operation: The user uses the file selection function on the browser to select the required data file and upload it to the system, where the data is sent to the server and received.

[0570] Step 2: Receiving and storing data

[0571] Subject: Server

[0572] Input: User uploaded data file

[0573] Output: Data files saved in a specific directory

[0574] Specific operation: The server receives the HTTP POST request and saves the uploaded file in a pre-specified folder (e.g., / uploaded_data), allowing data to be managed centrally.

[0575] Step 3: Data preprocessing (text data)

[0576] Subject: Server

[0577] Input: Saved text data file (PDF)

[0578] Output: Preprocessed text data

[0579] What it does: It uses the Python pdfplumber library to extract text from PDF files and saves the extracted text in a text file or database.

[0580] Step 4: Data preprocessing (audio data)

[0581] Subject: Server

[0582] Input: Saved audio data file (MP4)

[0583] Output: Preprocessed text data

[0584] What it does: The audio data is processed using Google's Speech-to-Text API, which converts the audio into text, which is then saved in a text file or database.

[0585] Step 5: Data preprocessing (image and video data)

[0586] Subject: Server

[0587] Input: Saved image or video data files (JPG, MP4, etc.)

[0588] Output: Preprocessed image recognition data

[0589] Specific operation: Analyzes images and videos using OpenCV and TensorFlow models, for example, performs tasks such as object detection and face recognition, and saves the analysis results.

[0590] Step 6: Data analysis (text data)

[0591] Subject: Server

[0592] Input: Preprocessed text data

[0593] Output: Analyzed patterns and trends

[0594] Specific operation: Using the BERT model, keyword extraction and sentiment analysis are performed. The analysis results are stored in a database.

[0595] Step 7: Data analysis (audio data)

[0596] Subject: Server

[0597] Input: Preprocessed text data (converted from audio)

[0598] Output: Parsed opinions and feedback

[0599] How it works: Using the GPT-4 model, the voice data is analyzed in detail to extract consumer opinions and feedback. The analysis results are stored in a database.

[0600] Step 8: Generate strategic advice

[0601] Subject: Server

[0602] Input: Analysis results (text data, audio data, image and video data)

[0603] Output: Strategic advice

[0604] How it works: Generative AI is used to generate strategic advice based on the analysis results. For example, GPT-4 is used to create strategic proposals based on market trends. The generated advice is then stored in a database.

[0605] Step 9: Present to the user

[0606] Subject: Server

[0607] Input: Generated strategic advice

[0608] Output: Visualized data (dashboard)

[0609] Specific behavior: Strategic advice is visualized and displayed on a dashboard using JavaScript libraries (e.g., D3.js, Chart.js). Users can access the dashboard from their devices and check the visualized results.

[0610] (Application example 1)

[0611] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0612] In autonomous vehicles, it is difficult to efficiently analyze large amounts of diverse data, such as operational data, voice instructions, and camera footage, and provide safe and optimal driving routes. In particular, it is necessary to generate and present strategic advice in real time, which will help prevent accidents and ensure efficient route selection.

[0613] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0614] In this invention, the server includes means for receiving text, audio, image, and video data, means for preprocessing the received data, means for analyzing the preprocessed data using generative AI and artificial general intelligence, means for generating strategic advice based on the analysis results, means for presenting the generated advice to a user, means for receiving camera footage, sensor data, and voice instructions from an autonomous vehicle, means for analyzing road conditions using the received data, and means for proposing a safe driving route based on the analysis results, thereby enabling the autonomous vehicle to select a safe and optimal driving route in real time, prevent accidents, and drive efficiently.

[0615] "Text" is a data format that expresses information using characters and symbols.

[0616] "Voice" refers to data relating to human speech or acoustics transmitted through sound waves.

[0617] An "image" is a still image that represents visual information in pixels.

[0618] "Video" refers to dynamic visual data that expresses movement using successive image frames.

[0619] "Preprocessing" refers to the initial processing of received data to convert it into a form suitable for analysis.

[0620] "Generative AI" is an artificial intelligence technology that learns from large datasets and generates new data.

[0621] "Artificial general intelligence" refers to advanced artificial intelligence that has the ability to solve a wide range of problems, not just specific tasks.

[0622] "Strategic advice" refers to recommendations based on data analysis that provide users with optimal courses of action.

[0623] An "autonomous vehicle" is a vehicle that has the ability to operate autonomously using artificial intelligence and various sensors.

[0624] "Camera footage" refers to video data captured by a camera and recorded as visual information.

[0625] "Sensor data" refers to data about physical phenomena measured by various sensors.

[0626] "Voice instructions" refer to commands or instructions given using voice.

[0627] "Road conditions" refers to environmental information that affects the safety and efficiency of driving, such as road congestion and road surface conditions.

[0628] A "travel route" is the optimal route to a destination.

[0629] The present invention is a system for analyzing operational data of autonomous vehicles to support safe and efficient operation. This system analyzes text, audio, image, and video data to generate and present strategic advice. Specific embodiments of this system are described below.

[0630] System configuration and operation

[0631] This system is mainly composed of components such as servers, terminals, and users, each of which fulfills its respective role to realize the overall function.

[0632] Server Roles

[0633] The server is the core of this system and has the following functions:

[0634] Data reception and preprocessing:

[0635] The server receives text, audio, image, and video data and pre-processes this data: for example, camera footage is processed frame by frame using OpenCV, and audio data is converted to text using the SpeechRecognition library.

[0636] Data Analysis:

[0637] The pre-processed data is then analyzed using deep learning models (generative AI and artificial general intelligence), for example, to detect road conditions and vehicle anomalies using deep learning frameworks such as TensorFlow and Keras.

[0638] Strategic advice generation:

[0639] Based on the analysis results, strategic advice is generated, such as suggesting safe driving routes and generating warning messages for drivers.

[0640] Present your proposal:

[0641] The server sends the generated advice to the terminal and presents it visually to the user. Operation information and advice are displayed in real time in a dashboard format.

[0642] Device Role

[0643] A terminal is a device that is directly operated by the user, such as a smartphone or tablet. A terminal has the following functions:

[0644] Data entry and collection:

[0645] It sends camera footage, sensor data, and voice instructions uploaded by the user to the server.

[0646] View the advice offered:

[0647] The terminal visually displays strategic advice sent from the server, allowing users to see driving routes and warning messages in real time.

[0648] User Roles

[0649] The user is the entity that uses the system and performs the following operations:

[0650] Upload data:

[0651] Users upload a variety of data to the system, including market analysis reports and interview videos.

[0652] Review the suggested advice:

[0653] Users can view the strategic advice and driving routes generated through a dashboard on their device and make decisions based on the results.

[0654] Examples of specific examples and prompts

[0655] When an autonomous vehicle is driving in a city and analyzing camera footage to detect obstacles and pedestrians ahead, the server uses the following prompt sentence:

[0656] Camera video analysis prompt: "Analyze the video data obtained from the vehicle's front camera to detect obstacles on the road."

[0657] When the vehicle selects a particular route based on the driver's voice instructions, the server uses the following prompt sentence:

[0658] Voice command analysis prompt: "Please convert the driver's current voice commands into text and analyze their actions."

[0659] This will enable autonomous vehicles to select safe and optimal routes in real time, prevent accidents, and operate efficiently.

[0660] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0661] Step 1:

[0662] The server receives camera footage, sensor data, and audio instructions from the user.

[0663] Input: Camera video files, sensor data files, audio files

[0664] Specific operation: Captures camera images as frame data, receives sensor data as numerical data, and recognizes audio files as audio data.

[0665] Step 2:

[0666] The server pre-processes the received data.

[0667] Input: Camera image frame data, sensor data values, audio data

[0668] Specific operations: Camera image frames are resized and preprocessed using OpenCV, sensor data is normalized to an appropriate range, and audio data is converted to text using the SpeechRecognition library.

[0669] Output: Preprocessed video data, normalized sensor data, and transcribed audio data

[0670] Step 3:

[0671] The server analyzes the pre-processed data using generative AI models and artificial general intelligence.

[0672] Input: Preprocessed video data, normalized sensor data, and transcribed audio data

[0673] Specific operation: Preprocessed video data is input into deep learning models (TensorFlow and Keras) to analyze road conditions, sensor data is analyzed, anomaly detection algorithms are applied, and voice data is analyzed to determine instructions.

[0674] Output: Analyzed road conditions, abnormal sensor data, voice instructions

[0675] Step 4:

[0676] The server generates strategic advice based on the analysis results.

[0677] Input: Analyzed road conditions, abnormal sensor data, voice instructions

[0678] Specific operation: Based on the analysis results, a strategic algorithm is applied to generate safe driving routes and warning messages for drivers.

[0679] Output: Generated route suggestions and warning messages for the driver

[0680] Step 5:

[0681] The server transmits the generated advice to the terminal.

[0682] Input: Generated route suggestions, warning messages for drivers

[0683] Specific operation: The generated advice is sent to the terminal and displayed on the terminal in real time.

[0684] Output: Route suggestions displayed on the terminal, warning messages to the driver

[0685] Step 6:

[0686] The terminal visually presents the generated advice to the user.

[0687] Input: Route suggestions displayed on the device, warning messages to the driver

[0688] Specific operation: Routes and warning messages are presented to the user in real time using a dashboard-style UI.

[0689] Output: User-visible route and warning messages

[0690] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0691] This invention relates to a system that generates and presents strategic advice by analyzing text, voice, image, and video data and combining it with an emotion engine that recognizes the user's emotions. This system is composed of components such as a server, a terminal, and a user, and each component fulfills its role to realize the overall function of the invention.

[0692] Specific system configuration and operation

[0693] 1. Data Collection

[0694] A user uploads data such as text, audio, images, and video to the system.

[0695] The server receives the uploaded data and stores it in a specific directory.

[0696] 2. Data Preprocessing

[0697] The server processes the received data using different pre-processing methods: for example, text data is formatted using natural language processing (NLP), audio data is converted to text using speech recognition technology, and image and video data is analyzed using image recognition technology.

[0698] 3. Data Analysis

[0699] The server then analyzes the pre-processed data using generative AI and artificial general intelligence (AGI) to discover important patterns, trends, and potential opportunities.

[0700] 4. Emotional Recognition

[0701] The server uses an emotion engine to recognize and analyze the user's emotions, for example, identifying the user's emotional state (e.g., joy, sadness, anger) from the user's voice data and text data.

[0702] 5. Generating strategic advice

[0703] The server generates specific strategic advice based on the analysis results, and provides advice that also takes into account the results of the user's sentiment analysis.

[0704] 6. Presentation to the User

[0705] The server provides a means to visually present the generated advice to the user, who can then access this information using their terminal and view the results in an easy to understand format.

[0706] Specific examples

[0707] For example, suppose a startup company is planning to launch a new product on the market. This company needs to understand the competitive situation and develop an effective marketing strategy. To do this, they use this system.

[0708] 1. The user uploads the market analysis report (PDF) and video of the consumer interview to the system.

[0709] 2. The server receives these data, extracts the text data, and converts the audio from the video to text.

[0710] 3. The server uses generative AI and AGI to analyze the data and analyse the competitive landscape and consumer opinion.

[0711] 4. The server uses an emotion engine to recognize and analyze consumer emotions from the interview audio data, for example, identifying which products consumers have positive emotions about.

[0712] 5. Based on the analysis, the server generates strategic advice that takes into account the user's emotional state, including how to enter new market segments and effective marketing messages to target customers.

[0713] 6. Users access the dashboard using their devices, review the proposed market strategies and analysis results, and make decisions based on the results.

[0714] In this way, the present invention provides a system that processes a variety of data formats in an integrated manner and provides useful strategic insights that take user sentiment into account. Users can easily upload data and receive analysis results in an intuitively understandable format, supporting quick and appropriate decision-making.

[0715] The processing flow will be explained below.

[0716] Step 1:

[0717] A user uploads data such as text, audio, images, and video to the system.

[0718] Example of specific operation: A user uploads a market analysis report (PDF) and a consumer interview video.

[0719] Step 2:

[0720] The server receives the uploaded data and stores it in a specific directory.

[0721] Example of specific operation: The server stores PDF files and video files in a designated folder.

[0722] Step 3:

[0723] The server initiates pre-processing based on the type of data.

[0724] Examples of how it works: Extracting text from PDF files and separating audio from video files.

[0725] Step 4:

[0726] The server uses speech recognition technology to convert the voice data into text.

[0727] Example of specific operation: Use a speech recognition engine to convert the audio of an interview video into text.

[0728] Step 5:

[0729] The server uses image recognition technology to extract important features from image and video data.

[0730] Example of specific operation: Extract key frames from a video file and analyze them using an image recognition model.

[0731] Step 6:

[0732] The server analyzes the pre-processed data using generative AI and artificial general intelligence.

[0733] Specific operation example: Apply natural language processing (NLP) technology to text data to perform keyword extraction and topic modeling.

[0734] Step 7:

[0735] The server uses an emotion engine to recognize and analyze the user's emotions.

[0736] Specific operation example: Identify the user's emotional state from voice data and text data, and determine joy, sadness, anger, etc.

[0737] Step 8:

[0738] The server synthesizes the analysis results and derives relevant insights.

[0739] Example of how it works: Integrate data from market analysis reports with text data from interviews to identify common trends and patterns.

[0740] Step 9:

[0741] The server generates strategic advice based on the analysis results.

[0742] Example of specific operation: Based on market trends, suggest ways to enter new market segments and create marketing messages that take into account the user's emotional state.

[0743] Step 10:

[0744] The server visualizes the generated advice and presents it to the user.

[0745] Example of how it works: Displays strategic advice and analytical results in graphs and charts in a dashboard format.

[0746] Step 11:

[0747] The terminal displays a dashboard to the user, allowing the user to review the results.

[0748] Example of how it works: A user opens a web browser and views the dashboard with suggested market segments and measures.

[0749] Example 2

[0750] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0751] Conventional data analysis systems lack the ability to process multiple data formats (text, audio, images, and videos) in a unified manner, making it particularly difficult to generate strategic advice that takes into account user sentiment analysis. Furthermore, they lacked a means to provide analysis results to users in an easy-to-understand format, making it difficult for users to make appropriate decisions.

[0752] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0753] In this invention, the server includes means for receiving text, audio, image, and video data, means for preprocessing the received data according to the type of data, means for analyzing the preprocessed data using a generation AI and an intelligent system, means for recognizing and analyzing a user's emotions, means for generating strategic advice based on the analysis results and emotion recognition results, and means for visually presenting the generated advice to the user. This enables unified analysis of multiple data formats and the provision of strategic advice that takes the user's emotions into consideration. Furthermore, visual presentation of the analysis results allows the user to easily understand and make appropriate decisions.

[0754] "Text data" is information expressed in the form of characters and sentences.

[0755] "Audio data" is information that records human voices and sounds in digital format.

[0756] "Image data" is visual information such as photographs and illustrations recorded in digital format.

[0757] "Video data" means a digital recording that combines moving visual and audio information.

[0758] "Means for receiving" refers to a technique or device for capturing data sent from a user within the system.

[0759] A "preprocessing means" is a technique or device for shaping or transforming collected data so that it is easier to analyze.

[0760] "Generative AI" refers to algorithms or systems that use artificial intelligence techniques to generate new information from data.

[0761] An "intelligent system" is an artificial intelligence system that has general knowledge and uses it to solve problems and analyze data.

[0762] "Emotion recognition and analysis means" refers to a technique or device for identifying and analyzing a user's emotional state from data provided by the user.

[0763] The "means for generating strategic advice" is a technology or device for generating guidelines or suggestions that are useful to the user based on the analysis results and emotion recognition results.

[0764] "Means for visually presenting to the user" refers to a technique or device for providing the analysis results and advice to the user in a visual format such as a graph or chart.

[0765] This invention relates to a system that generates and presents strategic advice by analyzing text, voice, image, and video data and combining it with an emotion engine that recognizes the user's emotions. This system is composed of components such as a server, a terminal, and a user, each of which fulfills a specific role to realize the overall function of the invention.

[0766] Data collection

[0767] Users upload data such as text, audio, images, and videos to the system. An interface for uploading this data is provided on the user's terminal. For example, a file can be selected by dragging and dropping using the file upload function of a web browser.

[0768] Data Preprocessing

[0769] The server receives the uploaded data and stores it in a specific directory. The received data is classified by type and the following preprocessing is performed depending on the data format:

[0770] Text data: A natural language processing (NLP) engine (e.g., SpaCy) is used to check grammar and filter out unwanted words.

[0771] Voice Data: We use a speech recognition engine (e.g., Google Speech-to-Text) to convert your voice into text.

[0772] Image data: An image recognition system (e.g., OpenCV) is used to identify objects in the image and generate metadata.

[0773] Video data: The video is broken down into frames, and each frame is analyzed using image recognition technology.

[0774] Data analysis

[0775] The server uses generative AI models (e.g., GPT-4) and intelligent systems to analyze the pre-processed data, allowing it to discover important patterns, trends, and potential opportunities, such as extracting competitor trends and identifying market niches based on market analysis reports.

[0776] Emotion recognition

[0777] The server uses an emotion recognition engine (e.g., IBM Watson Natural Language Understanding) to recognize and analyze emotions from the data provided by the user. It can identify a range of emotional states (such as joy, sadness, and anger) from text and audio data. For example, it can analyze the emotional responses from audio data of consumer interview videos to determine which products have the most positive opinions.

[0778] Strategic Advice Generation

[0779] The server generates strategic advice for users based on the analysis results and emotion recognition results. Specific suggestions are made using a generative AI model (e.g., GPT-4). For example, it can generate a suggestion such as, "Targeting a new product specifically for women in their 20s is likely to increase market share."

[0780] Presenting to the user

[0781] The server provides a dashboard to visually present the generated advice to the user. Users can access this dashboard from their own devices and view the analysis results in an intuitive, easy-to-understand format. The dashboard contains visual elements such as graphs, tables, and heat maps, and users can click on results to access more detailed information.

[0782] Specific examples

[0783] For example, suppose a startup company is planning to launch a new product. This company needs to understand the competitive situation and develop an effective marketing strategy. To do this, they use this system.

[0784] 1. The user uploads the market analysis report (PDF) and video of the consumer interview to the system.

[0785] 2. The server receives these data, extracts the text data, and converts the audio from the video to text.

[0786] 3. The server uses generative AI and intelligence systems to analyze the data and analyse competitors' situations and consumer opinions.

[0787] 4. The server uses an emotion recognition engine to recognize and analyze the consumer's emotions from the interview audio data, for example, to identify which products the consumer has positive feelings about.

[0788] 5. Based on the analysis, the server generates strategic advice that takes into account the user's emotional state, including how to enter new market segments and effective marketing messages to target customers.

[0789] 6. Users access the dashboard using their devices, review the proposed market strategies and analysis results, and make decisions based on the results.

[0790] Prompt Sentence Examples

[0791] "Based on this PDF report and video interviews, you can develop a market strategy and analyze consumer sentiment."

[0792] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0793] Step 1: Collect data

[0794] Input: User-uploaded text, audio, image, and video data

[0795] Specific operation: A user logs in to the system from their own terminal and uses the file upload function to upload data such as text, audio, images, and videos. Files are selected by drag-and-drop operation.

[0796] Output: Data stored on the server

[0797] Step 2: Save the data to the appropriate directory

[0798] Input: User uploaded data

[0799] Specific operation: The server receives the data uploaded by the user and saves it in a specific directory in real time. At this time, the data is temporarily stored in a buffer and then transferred to permanent storage after verification.

[0800] Output: Data stored in a specific directory on the server

[0801] Step 3: Preprocessing the data

[0802] Input: Raw data stored on the server

[0803] Specific operation: The server classifies the received data by type and performs appropriate preprocessing for each type. For example:

[0804] Text data: Use a natural language processing (NLP) engine (e.g., SpaCy) to check grammar and filter unnecessary words.

[0805] Voice Data: Converts speech to text using a speech recognition engine (e.g., Google Speech-to-Text).

[0806] Image data: An image recognition system (e.g., OpenCV) is used to identify objects in the image and generate metadata.

[0807] Video data: The video is broken down into frames, and each frame is analyzed using image recognition technology.

[0808] Output: Preprocessed data

[0809] Step 4: Data analysis

[0810] Input: Preprocessed data

[0811] What it does: The server analyzes the pre-processed data using generative AI models (e.g., GPT-4) and intelligent systems to discover important patterns, trends, and potential opportunities in the data.

[0812] Output: Analysis results (patterns, trends, potential opportunities)

[0813] Step 5: Recognize emotions

[0814] Input: Preprocessed text and audio data

[0815] Specific operation: The server uses an emotion recognition engine (e.g., IBM Watson Natural Language Understanding) to recognize and analyze emotions from the data provided by the user, and identifies a range of emotional states from the voice and text data.

[0816] Output: Emotion recognition results (happiness, sadness, anger, etc.)

[0817] Step 6: Generate strategic advice

[0818] Input: Analysis results and emotion recognition results

[0819] Specific operation: The server uses the generative AI model to generate specific strategic advice based on the analysis and emotion recognition results. For example, it generates a suggestion such as, "Targeting a new product specifically for women in their 20s is likely to increase market share."

[0820] Output: Strategic advice

[0821] Step 7: Present to the user

[0822] Input: Generated strategic advice

[0823] Specific operation: The server provides a dashboard to visually present the generated advice to the user. The user accesses this dashboard using their own device and understands the analysis results through visual elements such as graphs and charts.

[0824] Output: Visual analysis results and advice on a user-accessible dashboard

[0825] (Application example 2)

[0826] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0827] In typical brick-and-mortar stores, it is difficult for staff to accurately grasp customer emotions in real time and provide optimal product recommendations and services accordingly. Furthermore, staff need training and experience to deal with a wide variety of customers individually, and a lack of this training can lead to a decline in customer satisfaction. For this reason, there is a demand for support tools that enable staff to deal with customers quickly and accurately.

[0828] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving text, audio, image, and video data, means for preprocessing the received data, means for analyzing the preprocessed data using generative AI and artificial general intelligence, means for generating strategic advice based on the analysis results, means for presenting the generated advice to the user, means for recognizing emotions from the user's voice and video, and means for presenting advice generated in real time taking into account the emotion recognition results. This enables staff at physical stores to grasp customer emotions in real time and quickly provide optimal product suggestions and services based on that.

[0829] definition statement

[0830] "Text data" is data made up of character information.

[0831] "Audio data" refers to information collected as sound waves recorded in digital format.

[0832] "Image data" is data that stores visual information as a still image.

[0833] "Moving image data" is data that expresses dynamic visual information by combining successive still images along a time axis.

[0834] "Preprocessing" is the initial data processing to convert raw data into an analyzable format.

[0835] "Generative AI" is an artificial intelligence technology that generates new data and information based on input data.

[0836] "Artificial general intelligence" is artificial intelligence with a wide range of capabilities that can handle a wide variety of tasks.

[0837] An "emotion engine" is a technology that recognizes and analyzes a user's emotional state from voice, facial expressions, etc.

[0838] "Strategic advice" refers to specific and effective suggestions and instructions based on the results of analysis.

[0839] "Real-time presentation" means that analysis and advice are performed instantly, and information is provided to the user immediately.

[0840] MODE FOR CARRYING OUT THE INVENTION

[0841] The following describes an embodiment of the present invention.

[0842] The system includes means for receiving text, audio, image, and video data, means for preprocessing the received data, means for analyzing the preprocessed data using generative AI and artificial general intelligence, means for generating strategic advice based on the analysis results, means for presenting the generated advice to a user, means for recognizing emotions from the user's voice and video, and means for presenting the generated advice in real time, taking into account the emotion recognition results.

[0843] Data collection

[0844] Users use their smartphones to collect video and audio data of customers in physical stores via a camera and microphone, and these video and audio data are converted into text and image data.

[0845] Data Preprocessing

[0846] The server preprocesses the received video and audio data. Video data is converted into still images using an image recognition algorithm, and audio data is converted into text using speech recognition technology. Libraries and APIs such as OpenCV, Google Speech Recognition, and Emotion Recognition are used for preprocessing.

[0847] Data analysis

[0848] The pre-processed data is then analyzed using generative AI and artificial general intelligence to identify the products customers are interested in and their emotional state (interest, delight, doubt, etc.). The analysis is performed using OpenAI's API.

[0849] Emotion recognition

[0850] The user's emotional state is recognized from audio and video data. Emotion Recognition technology identifies the customer's emotional state and complements strategic advice based on this information.

[0851] Strategic Advice Generation

[0852] The server generates optimal strategic advice for the customer based on the analysis results and emotion recognition results, including product suggestions and service delivery methods that take into account the customer's emotions.

[0853] Presenting to the user

[0854] Strategic advice is presented to users in real time via a smartphone application, allowing wait staff to take immediate and appropriate action.

[0855] Specific examples

[0856] For example, if a customer asks for details about a particular product, the system might:

[0857] Customer question: "What are the features of this camera?"

[0858] Emotions recognized from customer facial expressions: Interest

[0859] Strategic advice: "This camera takes high-resolution photos and is waterproof. Plus, other customers often buy it."

[0860] Prompt Sentence Examples

[0861] Generate advice based on the user's statements and feelings below.

[0862] Say: "What is special about this camera?"

[0863] Emotion: "Interest"

[0864] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0865] Program processing steps

[0866] Step 1:

[0867] Users use the camera and microphone on their smartphones to collect video and audio data of customers in physical stores. The collected data is input as raw video frames and audio waveform data.

[0868] Step 2:

[0869] The terminal transmits the collected data to a server, which receives the video frames and audio waveform data, stores the video data as an image file, and stores the audio data as an audio file.

[0870] Step 3:

[0871] The server pre-processes the received data: video data is converted into still images using the OpenCV library, and audio data is converted into text using the Google Speech Recognition API. This pre-processing converts video data into an image format and audio data into a text format.

[0872] Step 4:

[0873] The server analyzes the preprocessed data using generative AI and artificial general intelligence (AGI). For the analysis, OpenAI's API is used to extract visual features from image data and customer interests and concerns from text data. This results in data feature quantities.

[0874] Step 5:

[0875] The server uses Emotion Recognition technology to recognize emotions from the user's video and audio data. It uses a facial expression recognition algorithm (e.g., OpenFace) from the video data and a voice emotion recognition algorithm from the audio data. This recognition yields the user's emotional state (interest, joy, etc.).

[0876] Step 6:

[0877] The server generates strategic advice based on the analysis results and emotion recognition results. It uses OpenAI's API to integrate this data and generate prompts. Based on these prompts, the server automatically generates advice that includes optimal product recommendations and service delivery methods for the customer.

[0878] Step 7:

[0879] The server then sends the generated strategic advice to the terminal, which then displays the received advice in real time on the smartphone screen so that the customer service staff can immediately check it.

[0880] Step 8:

[0881] The user (customer service staff) can then provide appropriate product recommendations and services to customers based on the strategic advice provided, which will enable quick and accurate customer service and is expected to improve customer satisfaction.

[0882] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0883] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0884] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0885] [Third embodiment]

[0886] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0887] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0888] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0889] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0890] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0891] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0892] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0893] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0894] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0895] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0896] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0897] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0898] This invention relates to a system that analyzes text, audio, image, and video data to generate and present strategic advice. This system is composed of components such as a server, a terminal, and a user, and each component fulfills its role to realize the overall function of the invention.

[0899] Specific system configuration and operation

[0900] 1. Data Collection

[0901] Users upload data in various formats to the system, such as market analysis reports and interview videos.

[0902] The server receives the uploaded data and stores it in a specific directory.

[0903] 2. Data Preprocessing

[0904] The server processes the received data using different pre-processing methods: text data is formatted using natural language processing (NLP), audio data is converted to text using speech recognition technology, and image and video data is analyzed using image recognition technology.

[0905] For example, it extracts text from PDFs of market analysis reports uploaded by users and converts audio from interview videos into text.

[0906] 3. Data Analysis

[0907] The server then analyzes the pre-processed data using generative AI and artificial general intelligence (AGI) to discover important patterns, trends, and potential opportunities.

[0908] For example, market trends can be identified from extracted text data, and consumer opinions can be analyzed based on the content of interviews.

[0909] 4. Generating strategic advice

[0910] The server generates specific strategic advice based on the analysis results, including strategy suggestions based on market trends and the construction of marketing messages.

[0911] For example, the analysis results can be used to propose new market segments and generate effective messages for target customers.

[0912] 5. Presentation to the User

[0913] The server provides a means to visually present the generated advice to the user, who can then access this information using their terminal and view the results in an easy to understand format.

[0914] For example, users can view suggested market strategies and analysis results on the dashboard.

[0915] Specific examples

[0916] Let's say a startup company is planning to launch a new product into the market. This company needs to understand the competitive situation and develop an effective marketing strategy. To do this, they use this system.

[0917] 1. Users upload market analysis reports (PDF) and consumer interview videos.

[0918] 2. The server receives this data and performs preprocessing by extracting text from the PDF and converting audio from the video to text.

[0919] 3. The server uses generative AI and AGI to analyze the data and analyse the competitive landscape and consumer opinion.

[0920] 4. The server uses the analysis results to develop new market segments and generate effective marketing messages.

[0921] 5. Users access the dashboard from their devices, review the proposed strategies and analysis results, and make decisions based on the results.

[0922] As described above, the present invention provides a system that processes a variety of data formats in an integrated manner and provides useful strategic insights. Users can easily upload data and obtain analytical results in an intuitive format, supporting quick and appropriate decision-making.

[0923] The processing flow will be explained below.

[0924] Step 1:

[0925] A user uploads data such as text, audio, images, and video to the system.

[0926] Example of specific operation: A user uploads a market analysis report (PDF) and an interview video.

[0927] Step 2:

[0928] The server receives the uploaded data and stores it in a specific directory.

[0929] Example of specific operation: The server stores PDF files and video files in a designated folder.

[0930] Step 3:

[0931] The server initiates pre-processing based on the type of data.

[0932] Examples of how it works: Extracting text from PDF files and separating audio from video files.

[0933] Step 4:

[0934] The server uses speech recognition technology to convert the voice data into text.

[0935] Example of specific operation: Use a speech recognition engine to convert the audio of an interview video into text.

[0936] Step 5:

[0937] The server uses image recognition technology to extract important features from image and video data.

[0938] Example of specific operation: Extract key frames from a video file and analyze them using an image recognition model.

[0939] Step 6:

[0940] The server analyzes the pre-processed data using generative AI and artificial general intelligence.

[0941] Specific operation example: Apply natural language processing (NLP) technology to text data to perform keyword extraction and topic modeling.

[0942] Step 7:

[0943] The server synthesizes the analysis results and derives relevant insights.

[0944] Example of how it works: Integrate data from market analysis reports with text data from interviews to identify common trends and patterns.

[0945] Step 8:

[0946] The server generates strategic advice based on the analysis results.

[0947] Example of specific operation: Propose ways to enter new market segments based on market trends.

[0948] Step 9:

[0949] The server visualizes the generated advice and presents it to the user.

[0950] Example of how it works: Displays strategic advice and analytical results in graphs and charts in a dashboard format.

[0951] Step 10:

[0952] The terminal displays a dashboard to the user, allowing the user to review the results.

[0953] Example of how it works: A user opens a web browser and views the dashboard with suggested market segments and measures.

[0954] Example 1

[0955] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0956] There is a need for a method to process diverse data formats (text, audio, images, video) in a unified manner and obtain information that will enable users to make strategic decisions quickly and accurately, but current technology lacks an efficient and effective way to achieve this. Therefore, there is a need for a system that can analyze diverse data formats and provide effective strategic advice.

[0957] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0958] In this invention, the server includes means for receiving data, means for preprocessing the received data based on different formats, means for analyzing the preprocessed data using generative artificial intelligence and general artificial intelligence, means for generating strategic advice based on the analysis results, and means for visually presenting the generated advice to the user. This makes it possible to extract useful information from data in various formats and provide strategic advice in a form that the user can intuitively understand.

[0959] The "means for receiving data" is a function for importing various data formats such as text, audio, images, and video provided by the user into the server.

[0960] The "means for preprocessing based on different formats" is a function for applying appropriate natural language processing technology, speech recognition technology, and image recognition technology depending on the format of the received data, and converting the data into an analyzable state.

[0961] "Generative AI" refers to algorithms and their implementations that learn large amounts of data and perform highly accurate predictions and generation for a variety of tasks.

[0962] "General artificial intelligence" is a general-purpose artificial intelligence technology that does not depend on specific tasks and can autonomously solve a variety of problems.

[0963] "Means for analysis" refers to processing functions that extract patterns and identify trends from preprocessed data to derive useful information contained in the data.

[0964] The "means for generating strategic advice" is a function for specifically proposing feasible strategies and policies to users based on the analysis results.

[0965] "Visual presentation means" refers to a function for displaying the generated strategic advice in a form that can be intuitively understood by the user, including in the form of graphs or dashboards.

[0966] This invention relates to a system that analyzes text, audio, image, and video data to generate and present strategic advice. This system is composed of components such as a server, a terminal, and a user, and each component fulfills its role to realize the overall function of the invention.

[0967] Data collection

[0968] Users can upload various types of data, such as market analysis reports and interview videos, to the system by using the file selection function on the browser to upload the data files.

[0969] The server receives these data and saves them in the specified directory. Uploaded files are sent to the server via HTTP POST request and saved in a specific folder on the server (e.g., / uploaded_data).

[0970] Data Preprocessing

[0971] The server performs preprocessing on the received data. Text data is formatted using natural language processing (NLP) technology, voice data is converted to text using voice recognition technology, and image and video data is analyzed using image recognition technology. For example, the following technologies are used:

[0972] Extract text from PDFs: Uses Python's pdfplumber library.

[0973] Converts voice data to text: Uses Google's Speech-to-Text API.

[0974] Image and video analysis: Uses OpenCV and TensorFlow models.

[0975] Data analysis

[0976] The server uses generative artificial intelligence (generative AI) and artificial general intelligence (AGI) to analyze the pre-processed data. This analysis uncovers important patterns, trends, and potential opportunities. Specifically, it uses the following techniques:

[0977] BERT model: Used for keyword extraction and sentiment analysis of preprocessed text data.

[0978] GPT-4 model: Used to analyze the context of the interview content.

[0979] Strategic Advice Generation

[0980] The server generates specific strategic advice based on the analysis results, including strategy proposals and marketing message construction based on market trends. For example, it uses GPT-4 to generate strategy proposals based on detected trends. It can also take a context-based approach by taking into account the user's past success stories.

[0981] Presenting to the user

[0982] The server provides a means to visually present the generated advice to the user, who can then access this information using their device and view the results in an easy-to-understand format, for example by displaying the results in a dashboard using a JavaScript library (e.g., D3.js or Chart.js).

[0983] Specific examples

[0984] Consider a startup company planning to launch a new product in the market. The company needs to understand the competitive landscape and develop an effective marketing strategy.

[0985] 1. The user uploads the market analysis report (PDF) and consumer interview video to the system. This is done using the file selection function on the browser.

[0986] 2. The server receives these data, extracts text from PDFs, and converts audio from videos to text. For example, the pdfplumber library extracts text from PDFs and Google's Speech-to-Text API converts audio to text.

[0987] 3. The server uses generative AI and AGI to analyze the competitive landscape and consumer opinions, using the BERT model for market analysis and the GPT-4 model for detailed analysis of interview content.

[0988] 4. The server uses the analysis results to develop strategies for entering new market segments and generating effective marketing messages. It uses GPT-4 to generate strategic recommendations based on the trends detected.

[0989] 5. Users access the dashboard from their device to view the proposed strategies and analysis results. The dashboard displays data visualized using D3.js and Chart.js.

[0990] Based on this example, here are some example prompts:

[0991] "Analyze market analysis reports and consumer interview videos regarding the launch of new products, and propose effective marketing strategies based on the situation of competitors and consumer opinions."

[0992] Through the above process, useful information can be extracted from a variety of data formats, enabling users to make quick and appropriate decisions.

[0993] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0994] Step 1: Upload your data

[0995] Subject: User

[0996] Input: Market analysis report (PDF), consumer interview video (MP4, etc.)

[0997] Output: Data file uploaded to the server

[0998] Specific operation: The user uses the file selection function on the browser to select the required data file and upload it to the system, where the data is sent to the server and received.

[0999] Step 2: Receiving and storing data

[1000] Subject: Server

[1001] Input: User uploaded data file

[1002] Output: Data files saved in a specific directory

[1003] Specific operation: The server receives the HTTP POST request and saves the uploaded file in a pre-specified folder (e.g., / uploaded_data), allowing data to be managed centrally.

[1004] Step 3: Data preprocessing (text data)

[1005] Subject: Server

[1006] Input: Saved text data file (PDF)

[1007] Output: Preprocessed text data

[1008] What it does: It uses the Python pdfplumber library to extract text from PDF files and saves the extracted text in a text file or database.

[1009] Step 4: Data preprocessing (audio data)

[1010] Subject: Server

[1011] Input: Saved audio data file (MP4)

[1012] Output: Preprocessed text data

[1013] What it does: The audio data is processed using Google's Speech-to-Text API, which converts the audio into text, which is then saved in a text file or database.

[1014] Step 5: Data preprocessing (image and video data)

[1015] Subject: Server

[1016] Input: Saved image or video data files (JPG, MP4, etc.)

[1017] Output: Preprocessed image recognition data

[1018] Specific operation: Analyzes images and videos using OpenCV and TensorFlow models, for example, performs tasks such as object detection and face recognition, and saves the analysis results.

[1019] Step 6: Data analysis (text data)

[1020] Subject: Server

[1021] Input: Preprocessed text data

[1022] Output: Analyzed patterns and trends

[1023] Specific operation: Using the BERT model, keyword extraction and sentiment analysis are performed. The analysis results are stored in a database.

[1024] Step 7: Data analysis (audio data)

[1025] Subject: Server

[1026] Input: Preprocessed text data (converted from audio)

[1027] Output: Parsed opinions and feedback

[1028] How it works: Using the GPT-4 model, the voice data is analyzed in detail to extract consumer opinions and feedback. The analysis results are stored in a database.

[1029] Step 8: Generate strategic advice

[1030] Subject: Server

[1031] Input: Analysis results (text data, audio data, image and video data)

[1032] Output: Strategic advice

[1033] How it works: Generative AI is used to generate strategic advice based on the analysis results. For example, GPT-4 is used to create strategic proposals based on market trends. The generated advice is then stored in a database.

[1034] Step 9: Present to the user

[1035] Subject: Server

[1036] Input: Generated strategic advice

[1037] Output: Visualized data (dashboard)

[1038] Specific behavior: Strategic advice is visualized and displayed on a dashboard using JavaScript libraries (e.g., D3.js, Chart.js). Users can access the dashboard from their devices and check the visualized results.

[1039] (Application example 1)

[1040] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1041] In autonomous vehicles, it is difficult to efficiently analyze large amounts of diverse data, such as operational data, voice instructions, and camera footage, and provide safe and optimal driving routes. In particular, it is necessary to generate and present strategic advice in real time, which will help prevent accidents and ensure efficient route selection.

[1042] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1043] In this invention, the server includes means for receiving text, audio, image, and video data, means for preprocessing the received data, means for analyzing the preprocessed data using generative AI and artificial general intelligence, means for generating strategic advice based on the analysis results, means for presenting the generated advice to a user, means for receiving camera footage, sensor data, and voice instructions from an autonomous vehicle, means for analyzing road conditions using the received data, and means for proposing a safe driving route based on the analysis results, thereby enabling the autonomous vehicle to select a safe and optimal driving route in real time, prevent accidents, and drive efficiently.

[1044] "Text" is a data format that expresses information using characters and symbols.

[1045] "Voice" refers to data relating to human speech or acoustics transmitted through sound waves.

[1046] An "image" is a still image that represents visual information in pixels.

[1047] "Video" refers to dynamic visual data that expresses movement using successive image frames.

[1048] "Preprocessing" refers to the initial processing of received data to convert it into a form suitable for analysis.

[1049] "Generative AI" is an artificial intelligence technology that learns from large datasets and generates new data.

[1050] "Artificial general intelligence" refers to advanced artificial intelligence that has the ability to solve a wide range of problems, not just specific tasks.

[1051] "Strategic advice" refers to recommendations based on data analysis that provide users with optimal courses of action.

[1052] An "autonomous vehicle" is a vehicle that has the ability to operate autonomously using artificial intelligence and various sensors.

[1053] "Camera footage" refers to video data captured by a camera and recorded as visual information.

[1054] "Sensor data" refers to data about physical phenomena measured by various sensors.

[1055] "Voice instructions" refer to commands or instructions given using voice.

[1056] "Road conditions" refers to environmental information that affects the safety and efficiency of driving, such as road congestion and road surface conditions.

[1057] A "travel route" is the optimal route to a destination.

[1058] The present invention is a system for analyzing operational data of autonomous vehicles to support safe and efficient operation. This system analyzes text, audio, image, and video data to generate and present strategic advice. Specific embodiments of this system are described below.

[1059] System configuration and operation

[1060] This system is mainly composed of components such as servers, terminals, and users, each of which fulfills its respective role to realize the overall function.

[1061] Server Roles

[1062] The server is the core of this system and has the following functions:

[1063] Data reception and preprocessing:

[1064] The server receives text, audio, image, and video data and pre-processes this data: for example, camera footage is processed frame by frame using OpenCV, and audio data is converted to text using the SpeechRecognition library.

[1065] Data Analysis:

[1066] The pre-processed data is then analyzed using deep learning models (generative AI and artificial general intelligence), for example, to detect road conditions and vehicle anomalies using deep learning frameworks such as TensorFlow and Keras.

[1067] Strategic advice generation:

[1068] Based on the analysis results, strategic advice is generated, such as suggesting safe driving routes and generating warning messages for drivers.

[1069] Present your proposal:

[1070] The server sends the generated advice to the terminal and presents it visually to the user. Operation information and advice are displayed in real time in a dashboard format.

[1071] Device Role

[1072] A terminal is a device that is directly operated by the user, such as a smartphone or tablet. A terminal has the following functions:

[1073] Data entry and collection:

[1074] It sends camera footage, sensor data, and voice instructions uploaded by the user to the server.

[1075] View the advice offered:

[1076] The terminal visually displays strategic advice sent from the server, allowing users to see driving routes and warning messages in real time.

[1077] User Roles

[1078] The user is the entity that uses the system and performs the following operations:

[1079] Upload data:

[1080] Users upload a variety of data to the system, including market analysis reports and interview videos.

[1081] Review the suggested advice:

[1082] Users can view the strategic advice and driving routes generated through a dashboard on their device and make decisions based on the results.

[1083] Examples of specific examples and prompts

[1084] When an autonomous vehicle is driving in a city and analyzing camera footage to detect obstacles and pedestrians ahead, the server uses the following prompt sentence:

[1085] Camera video analysis prompt: "Analyze the video data obtained from the vehicle's front camera to detect obstacles on the road."

[1086] When the vehicle selects a particular route based on the driver's voice instructions, the server uses the following prompt sentence:

[1087] Voice command analysis prompt: "Please convert the driver's current voice commands into text and analyze their actions."

[1088] This will enable autonomous vehicles to select safe and optimal routes in real time, prevent accidents, and operate efficiently.

[1089] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1090] Step 1:

[1091] The server receives camera footage, sensor data, and audio instructions from the user.

[1092] Input: Camera video files, sensor data files, audio files

[1093] Specific operation: Captures camera images as frame data, receives sensor data as numerical data, and recognizes audio files as audio data.

[1094] Step 2:

[1095] The server pre-processes the received data.

[1096] Input: Camera image frame data, sensor data values, audio data

[1097] Specific operations: Camera image frames are resized and preprocessed using OpenCV, sensor data is normalized to an appropriate range, and audio data is converted to text using the SpeechRecognition library.

[1098] Output: Preprocessed video data, normalized sensor data, and transcribed audio data

[1099] Step 3:

[1100] The server analyzes the pre-processed data using generative AI models and artificial general intelligence.

[1101] Input: Preprocessed video data, normalized sensor data, and transcribed audio data

[1102] Specific operation: Preprocessed video data is input into deep learning models (TensorFlow and Keras) to analyze road conditions, sensor data is analyzed, anomaly detection algorithms are applied, and voice data is analyzed to determine instructions.

[1103] Output: Analyzed road conditions, abnormal sensor data, voice instructions

[1104] Step 4:

[1105] The server generates strategic advice based on the analysis results.

[1106] Input: Analyzed road conditions, abnormal sensor data, voice instructions

[1107] Specific operation: Based on the analysis results, a strategic algorithm is applied to generate safe driving routes and warning messages for drivers.

[1108] Output: Generated route suggestions and warning messages for the driver

[1109] Step 5:

[1110] The server transmits the generated advice to the terminal.

[1111] Input: Generated route suggestions, warning messages for drivers

[1112] Specific operation: The generated advice is sent to the terminal and displayed on the terminal in real time.

[1113] Output: Route suggestions displayed on the terminal, warning messages to the driver

[1114] Step 6:

[1115] The terminal visually presents the generated advice to the user.

[1116] Input: Route suggestions displayed on the device, warning messages to the driver

[1117] Specific operation: Routes and warning messages are presented to the user in real time using a dashboard-style UI.

[1118] Output: User-visible route and warning messages

[1119] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1120] This invention relates to a system that generates and presents strategic advice by analyzing text, voice, image, and video data and combining it with an emotion engine that recognizes the user's emotions. This system is composed of components such as a server, a terminal, and a user, and each component fulfills its role to realize the overall function of the invention.

[1121] Specific system configuration and operation

[1122] 1. Data Collection

[1123] A user uploads data such as text, audio, images, and video to the system.

[1124] The server receives the uploaded data and stores it in a specific directory.

[1125] 2. Data Preprocessing

[1126] The server processes the received data using different pre-processing methods: for example, text data is formatted using natural language processing (NLP), audio data is converted to text using speech recognition technology, and image and video data is analyzed using image recognition technology.

[1127] 3. Data Analysis

[1128] The server then analyzes the pre-processed data using generative AI and artificial general intelligence (AGI) to discover important patterns, trends, and potential opportunities.

[1129] 4. Emotional Recognition

[1130] The server uses an emotion engine to recognize and analyze the user's emotions, for example, identifying the user's emotional state (e.g., joy, sadness, anger) from the user's voice data and text data.

[1131] 5. Generating strategic advice

[1132] The server generates specific strategic advice based on the analysis results, and provides advice that also takes into account the results of the user's sentiment analysis.

[1133] 6. Presentation to the User

[1134] The server provides a means to visually present the generated advice to the user, who can then access this information using their terminal and view the results in an easy to understand format.

[1135] Specific examples

[1136] For example, suppose a startup company is planning to launch a new product on the market. This company needs to understand the competitive situation and develop an effective marketing strategy. To do this, they use this system.

[1137] 1. The user uploads the market analysis report (PDF) and video of the consumer interview to the system.

[1138] 2. The server receives these data, extracts the text data, and converts the audio from the video to text.

[1139] 3. The server uses generative AI and AGI to analyze the data and analyse the competitive landscape and consumer opinion.

[1140] 4. The server uses an emotion engine to recognize and analyze consumer emotions from the interview audio data, for example, identifying which products consumers have positive emotions about.

[1141] 5. Based on the analysis, the server generates strategic advice that takes into account the user's emotional state, including how to enter new market segments and effective marketing messages to target customers.

[1142] 6. Users access the dashboard using their devices, review the proposed market strategies and analysis results, and make decisions based on the results.

[1143] In this way, the present invention provides a system that processes a variety of data formats in an integrated manner and provides useful strategic insights that take user sentiment into account. Users can easily upload data and receive analysis results in an intuitively understandable format, supporting quick and appropriate decision-making.

[1144] The processing flow will be explained below.

[1145] Step 1:

[1146] A user uploads data such as text, audio, images, and video to the system.

[1147] Example of specific operation: A user uploads a market analysis report (PDF) and a consumer interview video.

[1148] Step 2:

[1149] The server receives the uploaded data and stores it in a specific directory.

[1150] Example of specific operation: The server stores PDF files and video files in a designated folder.

[1151] Step 3:

[1152] The server initiates pre-processing based on the type of data.

[1153] Examples of how it works: Extracting text from PDF files and separating audio from video files.

[1154] Step 4:

[1155] The server uses speech recognition technology to convert the voice data into text.

[1156] Example of specific operation: Use a speech recognition engine to convert the audio of an interview video into text.

[1157] Step 5:

[1158] The server uses image recognition technology to extract important features from image and video data.

[1159] Example of specific operation: Extract key frames from a video file and analyze them using an image recognition model.

[1160] Step 6:

[1161] The server analyzes the pre-processed data using generative AI and artificial general intelligence.

[1162] Specific operation example: Apply natural language processing (NLP) technology to text data to perform keyword extraction and topic modeling.

[1163] Step 7:

[1164] The server uses an emotion engine to recognize and analyze the user's emotions.

[1165] Specific operation example: Identify the user's emotional state from voice data and text data, and determine joy, sadness, anger, etc.

[1166] Step 8:

[1167] The server synthesizes the analysis results and derives relevant insights.

[1168] Example of how it works: Integrate data from market analysis reports with text data from interviews to identify common trends and patterns.

[1169] Step 9:

[1170] The server generates strategic advice based on the analysis results.

[1171] Example of specific operation: Based on market trends, suggest ways to enter new market segments and create marketing messages that take into account the user's emotional state.

[1172] Step 10:

[1173] The server visualizes the generated advice and presents it to the user.

[1174] Example of how it works: Displays strategic advice and analytical results in graphs and charts in a dashboard format.

[1175] Step 11:

[1176] The terminal displays a dashboard to the user, allowing the user to review the results.

[1177] Example of how it works: A user opens a web browser and views the dashboard with suggested market segments and measures.

[1178] Example 2

[1179] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1180] Conventional data analysis systems lack the ability to process multiple data formats (text, audio, images, and videos) in a unified manner, making it particularly difficult to generate strategic advice that takes into account user sentiment analysis. Furthermore, they lacked a means to provide analysis results to users in an easy-to-understand format, making it difficult for users to make appropriate decisions.

[1181] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1182] In this invention, the server includes means for receiving text, audio, image, and video data, means for preprocessing the received data according to the type of data, means for analyzing the preprocessed data using a generation AI and an intelligent system, means for recognizing and analyzing a user's emotions, means for generating strategic advice based on the analysis results and emotion recognition results, and means for visually presenting the generated advice to the user. This enables unified analysis of multiple data formats and the provision of strategic advice that takes the user's emotions into consideration. Furthermore, visual presentation of the analysis results allows the user to easily understand and make appropriate decisions.

[1183] "Text data" is information expressed in the form of characters and sentences.

[1184] "Audio data" is information that records human voices and sounds in digital format.

[1185] "Image data" is visual information such as photographs and illustrations recorded in digital format.

[1186] "Video data" means a digital recording that combines moving visual and audio information.

[1187] "Means for receiving" refers to a technique or device for capturing data sent from a user within the system.

[1188] A "preprocessing means" is a technique or device for shaping or transforming collected data so that it is easier to analyze.

[1189] "Generative AI" refers to algorithms or systems that use artificial intelligence techniques to generate new information from data.

[1190] An "intelligent system" is an artificial intelligence system that has general knowledge and uses it to solve problems and analyze data.

[1191] "Emotion recognition and analysis means" refers to a technique or device for identifying and analyzing a user's emotional state from data provided by the user.

[1192] The "means for generating strategic advice" is a technology or device for generating guidelines or suggestions that are useful to the user based on the analysis results and emotion recognition results.

[1193] "Means for visually presenting to the user" refers to a technique or device for providing the analysis results and advice to the user in a visual format such as a graph or chart.

[1194] This invention relates to a system that generates and presents strategic advice by analyzing text, voice, image, and video data and combining it with an emotion engine that recognizes the user's emotions. This system is composed of components such as a server, a terminal, and a user, each of which fulfills a specific role to realize the overall function of the invention.

[1195] Data collection

[1196] Users upload data such as text, audio, images, and videos to the system. An interface for uploading this data is provided on the user's terminal. For example, a file can be selected by dragging and dropping using the file upload function of a web browser.

[1197] Data Preprocessing

[1198] The server receives the uploaded data and stores it in a specific directory. The received data is classified by type and the following preprocessing is performed depending on the data format:

[1199] Text data: A natural language processing (NLP) engine (e.g., SpaCy) is used to check grammar and filter out unwanted words.

[1200] Voice Data: We use a speech recognition engine (e.g., Google Speech-to-Text) to convert your voice into text.

[1201] Image data: An image recognition system (e.g., OpenCV) is used to identify objects in the image and generate metadata.

[1202] Video data: The video is broken down into frames, and each frame is analyzed using image recognition technology.

[1203] Data analysis

[1204] The server uses generative AI models (e.g., GPT-4) and intelligent systems to analyze the pre-processed data, allowing it to discover important patterns, trends, and potential opportunities, such as extracting competitor trends and identifying market niches based on market analysis reports.

[1205] Emotion recognition

[1206] The server uses an emotion recognition engine (e.g., IBM Watson Natural Language Understanding) to recognize and analyze emotions from the data provided by the user. It can identify a range of emotional states (such as joy, sadness, and anger) from text and audio data. For example, it can analyze the emotional responses from audio data of consumer interview videos to determine which products have the most positive opinions.

[1207] Strategic Advice Generation

[1208] The server generates strategic advice for users based on the analysis results and emotion recognition results. Specific suggestions are made using a generative AI model (e.g., GPT-4). For example, it can generate a suggestion such as, "Targeting a new product specifically for women in their 20s is likely to increase market share."

[1209] Presenting to the user

[1210] The server provides a dashboard to visually present the generated advice to the user. Users can access this dashboard from their own devices and view the analysis results in an intuitive, easy-to-understand format. The dashboard contains visual elements such as graphs, tables, and heat maps, and users can click on results to access more detailed information.

[1211] Specific examples

[1212] For example, suppose a startup company is planning to launch a new product. This company needs to understand the competitive situation and develop an effective marketing strategy. To do this, they use this system.

[1213] 1. The user uploads the market analysis report (PDF) and video of the consumer interview to the system.

[1214] 2. The server receives these data, extracts the text data, and converts the audio from the video to text.

[1215] 3. The server uses generative AI and intelligence systems to analyze the data and analyse competitors' situations and consumer opinions.

[1216] 4. The server uses an emotion recognition engine to recognize and analyze the consumer's emotions from the interview audio data, for example, to identify which products the consumer has positive feelings about.

[1217] 5. Based on the analysis, the server generates strategic advice that takes into account the user's emotional state, including how to enter new market segments and effective marketing messages to target customers.

[1218] 6. Users access the dashboard using their devices, review the proposed market strategies and analysis results, and make decisions based on the results.

[1219] Prompt Sentence Examples

[1220] "Based on this PDF report and video interviews, you can develop a market strategy and analyze consumer sentiment."

[1221] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1222] Step 1: Collect data

[1223] Input: User-uploaded text, audio, image, and video data

[1224] Specific operation: A user logs in to the system from their own terminal and uses the file upload function to upload data such as text, audio, images, and videos. Files are selected by drag-and-drop operation.

[1225] Output: Data stored on the server

[1226] Step 2: Save the data to the appropriate directory

[1227] Input: User uploaded data

[1228] Specific operation: The server receives the data uploaded by the user and saves it in a specific directory in real time. At this time, the data is temporarily stored in a buffer and then transferred to permanent storage after verification.

[1229] Output: Data stored in a specific directory on the server

[1230] Step 3: Preprocessing the data

[1231] Input: Raw data stored on the server

[1232] Specific operation: The server classifies the received data by type and performs appropriate preprocessing for each type. For example:

[1233] Text data: Use a natural language processing (NLP) engine (e.g., SpaCy) to check grammar and filter unnecessary words.

[1234] Voice Data: Converts speech to text using a speech recognition engine (e.g., Google Speech-to-Text).

[1235] Image data: An image recognition system (e.g., OpenCV) is used to identify objects in the image and generate metadata.

[1236] Video data: The video is broken down into frames, and each frame is analyzed using image recognition technology.

[1237] Output: Preprocessed data

[1238] Step 4: Data analysis

[1239] Input: Preprocessed data

[1240] What it does: The server analyzes the pre-processed data using generative AI models (e.g., GPT-4) and intelligent systems to discover important patterns, trends, and potential opportunities in the data.

[1241] Output: Analysis results (patterns, trends, potential opportunities)

[1242] Step 5: Recognize emotions

[1243] Input: Preprocessed text and audio data

[1244] Specific operation: The server uses an emotion recognition engine (e.g., IBM Watson Natural Language Understanding) to recognize and analyze emotions from the data provided by the user, and identifies a range of emotional states from the voice and text data.

[1245] Output: Emotion recognition results (happiness, sadness, anger, etc.)

[1246] Step 6: Generate strategic advice

[1247] Input: Analysis results and emotion recognition results

[1248] Specific operation: The server uses the generative AI model to generate specific strategic advice based on the analysis and emotion recognition results. For example, it generates a suggestion such as, "Targeting a new product specifically for women in their 20s is likely to increase market share."

[1249] Output: Strategic advice

[1250] Step 7: Present to the user

[1251] Input: Generated strategic advice

[1252] Specific operation: The server provides a dashboard to visually present the generated advice to the user. The user accesses this dashboard using their own device and understands the analysis results through visual elements such as graphs and charts.

[1253] Output: Visual analysis results and advice on a user-accessible dashboard

[1254] (Application example 2)

[1255] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1256] In typical brick-and-mortar stores, it is difficult for staff to accurately grasp customer emotions in real time and provide optimal product recommendations and services accordingly. Furthermore, staff need training and experience to deal with a wide variety of customers individually, and a lack of this training can lead to a decline in customer satisfaction. For this reason, there is a demand for support tools that enable staff to deal with customers quickly and accurately.

[1257] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving text, audio, image, and video data, means for preprocessing the received data, means for analyzing the preprocessed data using generative AI and artificial general intelligence, means for generating strategic advice based on the analysis results, means for presenting the generated advice to the user, means for recognizing emotions from the user's voice and video, and means for presenting advice generated in real time taking into account the emotion recognition results. This enables staff at physical stores to grasp customer emotions in real time and quickly provide optimal product suggestions and services based on that.

[1258] definition statement

[1259] "Text data" is data made up of character information.

[1260] "Audio data" refers to information collected as sound waves recorded in digital format.

[1261] "Image data" is data that stores visual information as a still image.

[1262] "Moving image data" is data that expresses dynamic visual information by combining successive still images along a time axis.

[1263] "Preprocessing" is the initial data processing to convert raw data into an analyzable format.

[1264] "Generative AI" is an artificial intelligence technology that generates new data and information based on input data.

[1265] "Artificial general intelligence" is artificial intelligence with a wide range of capabilities that can handle a wide variety of tasks.

[1266] An "emotion engine" is a technology that recognizes and analyzes a user's emotional state from voice, facial expressions, etc.

[1267] "Strategic advice" refers to specific and effective suggestions and instructions based on the results of analysis.

[1268] "Real-time presentation" means that analysis and advice are performed instantly, and information is provided to the user immediately.

[1269] MODE FOR CARRYING OUT THE INVENTION

[1270] The following describes an embodiment of the present invention.

[1271] The system includes means for receiving text, audio, image, and video data, means for preprocessing the received data, means for analyzing the preprocessed data using generative AI and artificial general intelligence, means for generating strategic advice based on the analysis results, means for presenting the generated advice to a user, means for recognizing emotions from the user's voice and video, and means for presenting the generated advice in real time, taking into account the emotion recognition results.

[1272] Data collection

[1273] Users use their smartphones to collect video and audio data of customers in physical stores via a camera and microphone, and these video and audio data are converted into text and image data.

[1274] Data Preprocessing

[1275] The server preprocesses the received video and audio data. Video data is converted into still images using an image recognition algorithm, and audio data is converted into text using speech recognition technology. Libraries and APIs such as OpenCV, Google Speech Recognition, and Emotion Recognition are used for preprocessing.

[1276] Data analysis

[1277] The pre-processed data is then analyzed using generative AI and artificial general intelligence to identify the products customers are interested in and their emotional state (interest, delight, doubt, etc.). The analysis is performed using OpenAI's API.

[1278] Emotion recognition

[1279] The user's emotional state is recognized from audio and video data. Emotion Recognition technology identifies the customer's emotional state and complements strategic advice based on this information.

[1280] Strategic Advice Generation

[1281] The server generates optimal strategic advice for the customer based on the analysis results and emotion recognition results, including product suggestions and service delivery methods that take into account the customer's emotions.

[1282] Presenting to the user

[1283] Strategic advice is presented to users in real time via a smartphone application, allowing wait staff to take immediate and appropriate action.

[1284] Specific examples

[1285] For example, if a customer asks for details about a particular product, the system might:

[1286] Customer question: "What are the features of this camera?"

[1287] Emotions recognized from customer facial expressions: Interest

[1288] Strategic advice: "This camera takes high-resolution photos and is waterproof. Plus, other customers often buy it."

[1289] Prompt Sentence Examples

[1290] Generate advice based on the user's statements and feelings below.

[1291] Say: "What is special about this camera?"

[1292] Emotion: "Interest"

[1293] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1294] Program processing steps

[1295] Step 1:

[1296] Users use the camera and microphone on their smartphones to collect video and audio data of customers in physical stores. The collected data is input as raw video frames and audio waveform data.

[1297] Step 2:

[1298] The terminal transmits the collected data to a server, which receives the video frames and audio waveform data, stores the video data as an image file, and stores the audio data as an audio file.

[1299] Step 3:

[1300] The server pre-processes the received data: video data is converted into still images using the OpenCV library, and audio data is converted into text using the Google Speech Recognition API. This pre-processing converts video data into an image format and audio data into a text format.

[1301] Step 4:

[1302] The server analyzes the preprocessed data using generative AI and artificial general intelligence (AGI). For the analysis, OpenAI's API is used to extract visual features from image data and customer interests and concerns from text data. This results in data feature quantities.

[1303] Step 5:

[1304] The server uses Emotion Recognition technology to recognize emotions from the user's video and audio data. It uses a facial expression recognition algorithm (e.g., OpenFace) from the video data and a voice emotion recognition algorithm from the audio data. This recognition yields the user's emotional state (interest, joy, etc.).

[1305] Step 6:

[1306] The server generates strategic advice based on the analysis results and emotion recognition results. It uses OpenAI's API to integrate this data and generate prompts. Based on these prompts, the server automatically generates advice that includes optimal product recommendations and service delivery methods for the customer.

[1307] Step 7:

[1308] The server then sends the generated strategic advice to the terminal, which then displays the received advice in real time on the smartphone screen so that the customer service staff can immediately check it.

[1309] Step 8:

[1310] The user (customer service staff) can then provide appropriate product recommendations and services to customers based on the strategic advice provided, which will enable quick and accurate customer service and is expected to improve customer satisfaction.

[1311] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1312] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1313] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1314] [Fourth embodiment]

[1315] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1316] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1317] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1318] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1319] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1320] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1321] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1322] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1323] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1324] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1325] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1326] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1327] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1328] This invention relates to a system that analyzes text, audio, image, and video data to generate and present strategic advice. This system is composed of components such as a server, a terminal, and a user, and each component fulfills its role to realize the overall function of the invention.

[1329] Specific system configuration and operation

[1330] 1. Data Collection

[1331] Users upload data in various formats to the system, such as market analysis reports and interview videos.

[1332] The server receives the uploaded data and stores it in a specific directory.

[1333] 2. Data Preprocessing

[1334] The server processes the received data using different pre-processing methods: text data is formatted using natural language processing (NLP), audio data is converted to text using speech recognition technology, and image and video data is analyzed using image recognition technology.

[1335] For example, it extracts text from PDFs of market analysis reports uploaded by users and converts audio from interview videos into text.

[1336] 3. Data Analysis

[1337] The server then analyzes the pre-processed data using generative AI and artificial general intelligence (AGI) to discover important patterns, trends, and potential opportunities.

[1338] For example, market trends can be identified from extracted text data, and consumer opinions can be analyzed based on the content of interviews.

[1339] 4. Generating strategic advice

[1340] The server generates specific strategic advice based on the analysis results, including strategy suggestions based on market trends and the construction of marketing messages.

[1341] For example, the analysis results can be used to propose new market segments and generate effective messages for target customers.

[1342] 5. Presentation to the User

[1343] The server provides a means to visually present the generated advice to the user, who can then access this information using their terminal and view the results in an easy to understand format.

[1344] For example, users can view suggested market strategies and analysis results on the dashboard.

[1345] Specific examples

[1346] Let's say a startup company is planning to launch a new product into the market. This company needs to understand the competitive situation and develop an effective marketing strategy. To do this, they use this system.

[1347] 1. Users upload market analysis reports (PDF) and consumer interview videos.

[1348] 2. The server receives this data and performs preprocessing by extracting text from the PDF and converting audio from the video to text.

[1349] 3. The server uses generative AI and AGI to analyze the data and analyse the competitive landscape and consumer opinion.

[1350] 4. The server uses the analysis results to develop new market segments and generate effective marketing messages.

[1351] 5. Users access the dashboard from their devices, review the proposed strategies and analysis results, and make decisions based on the results.

[1352] As described above, the present invention provides a system that processes a variety of data formats in an integrated manner and provides useful strategic insights. Users can easily upload data and obtain analytical results in an intuitive format, supporting quick and appropriate decision-making.

[1353] The processing flow will be explained below.

[1354] Step 1:

[1355] A user uploads data such as text, audio, images, and video to the system.

[1356] Example of specific operation: A user uploads a market analysis report (PDF) and an interview video.

[1357] Step 2:

[1358] The server receives the uploaded data and stores it in a specific directory.

[1359] Example of specific operation: The server stores PDF files and video files in a designated folder.

[1360] Step 3:

[1361] The server initiates pre-processing based on the type of data.

[1362] Examples of how it works: Extracting text from PDF files and separating audio from video files.

[1363] Step 4:

[1364] The server uses speech recognition technology to convert the voice data into text.

[1365] Example of specific operation: Use a speech recognition engine to convert the audio of an interview video into text.

[1366] Step 5:

[1367] The server uses image recognition technology to extract important features from image and video data.

[1368] Example of specific operation: Extract key frames from a video file and analyze them using an image recognition model.

[1369] Step 6:

[1370] The server analyzes the pre-processed data using generative AI and artificial general intelligence.

[1371] Specific operation example: Apply natural language processing (NLP) technology to text data to perform keyword extraction and topic modeling.

[1372] Step 7:

[1373] The server synthesizes the analysis results and derives relevant insights.

[1374] Example of how it works: Integrate data from market analysis reports with text data from interviews to identify common trends and patterns.

[1375] Step 8:

[1376] The server generates strategic advice based on the analysis results.

[1377] Example of specific operation: Propose ways to enter new market segments based on market trends.

[1378] Step 9:

[1379] The server visualizes the generated advice and presents it to the user.

[1380] Example of how it works: Displays strategic advice and analytical results in graphs and charts in a dashboard format.

[1381] Step 10:

[1382] The terminal displays a dashboard to the user, allowing the user to review the results.

[1383] Example of how it works: A user opens a web browser and views the dashboard with suggested market segments and measures.

[1384] Example 1

[1385] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1386] There is a need for a method to process diverse data formats (text, audio, images, video) in a unified manner and obtain information that will enable users to make strategic decisions quickly and accurately, but current technology lacks an efficient and effective way to achieve this. Therefore, there is a need for a system that can analyze diverse data formats and provide effective strategic advice.

[1387] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1388] In this invention, the server includes means for receiving data, means for preprocessing the received data based on different formats, means for analyzing the preprocessed data using generative artificial intelligence and general artificial intelligence, means for generating strategic advice based on the analysis results, and means for visually presenting the generated advice to the user. This makes it possible to extract useful information from data in various formats and provide strategic advice in a form that the user can intuitively understand.

[1389] The "means for receiving data" is a function for importing various data formats such as text, audio, images, and video provided by the user into the server.

[1390] The "means for preprocessing based on different formats" is a function for applying appropriate natural language processing technology, speech recognition technology, and image recognition technology depending on the format of the received data, and converting the data into an analyzable state.

[1391] "Generative AI" refers to algorithms and their implementations that learn large amounts of data and perform highly accurate predictions and generation for a variety of tasks.

[1392] "General artificial intelligence" is a general-purpose artificial intelligence technology that does not depend on specific tasks and can autonomously solve a variety of problems.

[1393] "Means for analysis" refers to processing functions that extract patterns and identify trends from preprocessed data to derive useful information contained in the data.

[1394] The "means for generating strategic advice" is a function for specifically proposing feasible strategies and policies to users based on the analysis results.

[1395] "Visual presentation means" refers to a function for displaying the generated strategic advice in a form that can be intuitively understood by the user, including in the form of graphs or dashboards.

[1396] This invention relates to a system that analyzes text, audio, image, and video data to generate and present strategic advice. This system is composed of components such as a server, a terminal, and a user, and each component fulfills its role to realize the overall function of the invention.

[1397] Data collection

[1398] Users can upload various types of data, such as market analysis reports and interview videos, to the system by using the file selection function on the browser to upload the data files.

[1399] The server receives these data and saves them in the specified directory. Uploaded files are sent to the server via HTTP POST request and saved in a specific folder on the server (e.g., / uploaded_data).

[1400] Data Preprocessing

[1401] The server performs preprocessing on the received data. Text data is formatted using natural language processing (NLP) technology, voice data is converted to text using voice recognition technology, and image and video data is analyzed using image recognition technology. For example, the following technologies are used:

[1402] Extract text from PDFs: Uses Python's pdfplumber library.

[1403] Converts voice data to text: Uses Google's Speech-to-Text API.

[1404] Image and video analysis: Uses OpenCV and TensorFlow models.

[1405] Data analysis

[1406] The server uses generative artificial intelligence (generative AI) and artificial general intelligence (AGI) to analyze the pre-processed data. This analysis uncovers important patterns, trends, and potential opportunities. Specifically, it uses the following techniques:

[1407] BERT model: Used for keyword extraction and sentiment analysis of preprocessed text data.

[1408] GPT-4 model: Used to analyze the context of the interview content.

[1409] Strategic Advice Generation

[1410] The server generates specific strategic advice based on the analysis results, including strategy proposals and marketing message construction based on market trends. For example, it uses GPT-4 to generate strategy proposals based on detected trends. It can also take a context-based approach by taking into account the user's past success stories.

[1411] Presenting to the user

[1412] The server provides a means to visually present the generated advice to the user, who can then access this information using their device and view the results in an easy-to-understand format, for example by displaying the results in a dashboard using a JavaScript library (e.g., D3.js or Chart.js).

[1413] Specific examples

[1414] Consider a startup company planning to launch a new product in the market. The company needs to understand the competitive landscape and develop an effective marketing strategy.

[1415] 1. The user uploads the market analysis report (PDF) and consumer interview video to the system. This is done using the file selection function on the browser.

[1416] 2. The server receives these data, extracts text from PDFs, and converts audio from videos to text. For example, the pdfplumber library extracts text from PDFs and Google's Speech-to-Text API converts audio to text.

[1417] 3. The server uses generative AI and AGI to analyze the competitive landscape and consumer opinions, using the BERT model for market analysis and the GPT-4 model for detailed analysis of interview content.

[1418] 4. The server uses the analysis results to develop strategies for entering new market segments and generating effective marketing messages. It uses GPT-4 to generate strategic recommendations based on the trends detected.

[1419] 5. Users access the dashboard from their device to view the proposed strategies and analysis results. The dashboard displays data visualized using D3.js and Chart.js.

[1420] Based on this example, here are some example prompts:

[1421] "Analyze market analysis reports and consumer interview videos regarding the launch of new products, and propose effective marketing strategies based on the situation of competitors and consumer opinions."

[1422] Through the above process, useful information can be extracted from a variety of data formats, enabling users to make quick and appropriate decisions.

[1423] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1424] Step 1: Upload your data

[1425] Subject: User

[1426] Input: Market analysis report (PDF), consumer interview video (MP4, etc.)

[1427] Output: Data file uploaded to the server

[1428] Specific operation: The user uses the file selection function on the browser to select the required data file and upload it to the system, where the data is sent to the server and received.

[1429] Step 2: Receiving and storing data

[1430] Subject: Server

[1431] Input: User uploaded data file

[1432] Output: Data files saved in a specific directory

[1433] Specific operation: The server receives the HTTP POST request and saves the uploaded file in a pre-specified folder (e.g., / uploaded_data), allowing data to be managed centrally.

[1434] Step 3: Data preprocessing (text data)

[1435] Subject: Server

[1436] Input: Saved text data file (PDF)

[1437] Output: Preprocessed text data

[1438] What it does: It uses the Python pdfplumber library to extract text from PDF files and saves the extracted text in a text file or database.

[1439] Step 4: Data preprocessing (audio data)

[1440] Subject: Server

[1441] Input: Saved audio data file (MP4)

[1442] Output: Preprocessed text data

[1443] What it does: The audio data is processed using Google's Speech-to-Text API, which converts the audio into text, which is then saved in a text file or database.

[1444] Step 5: Data preprocessing (image and video data)

[1445] Subject: Server

[1446] Input: Saved image or video data files (JPG, MP4, etc.)

[1447] Output: Preprocessed image recognition data

[1448] Specific operation: Analyzes images and videos using OpenCV and TensorFlow models, for example, performs tasks such as object detection and face recognition, and saves the analysis results.

[1449] Step 6: Data analysis (text data)

[1450] Subject: Server

[1451] Input: Preprocessed text data

[1452] Output: Analyzed patterns and trends

[1453] Specific operation: Using the BERT model, keyword extraction and sentiment analysis are performed. The analysis results are stored in a database.

[1454] Step 7: Data analysis (audio data)

[1455] Subject: Server

[1456] Input: Preprocessed text data (converted from audio)

[1457] Output: Parsed opinions and feedback

[1458] How it works: Using the GPT-4 model, the voice data is analyzed in detail to extract consumer opinions and feedback. The analysis results are stored in a database.

[1459] Step 8: Generate strategic advice

[1460] Subject: Server

[1461] Input: Analysis results (text data, audio data, image and video data)

[1462] Output: Strategic advice

[1463] How it works: Generative AI is used to generate strategic advice based on the analysis results. For example, GPT-4 is used to create strategic proposals based on market trends. The generated advice is then stored in a database.

[1464] Step 9: Present to the user

[1465] Subject: Server

[1466] Input: Generated strategic advice

[1467] Output: Visualized data (dashboard)

[1468] Specific behavior: Strategic advice is visualized and displayed on a dashboard using JavaScript libraries (e.g., D3.js, Chart.js). Users can access the dashboard from their devices and check the visualized results.

[1469] (Application example 1)

[1470] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1471] In autonomous vehicles, it is difficult to efficiently analyze large amounts of diverse data, such as operational data, voice instructions, and camera footage, and provide safe and optimal driving routes. In particular, it is necessary to generate and present strategic advice in real time, which will help prevent accidents and ensure efficient route selection.

[1472] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1473] In this invention, the server includes means for receiving text, audio, image, and video data, means for preprocessing the received data, means for analyzing the preprocessed data using generative AI and artificial general intelligence, means for generating strategic advice based on the analysis results, means for presenting the generated advice to a user, means for receiving camera footage, sensor data, and voice instructions from an autonomous vehicle, means for analyzing road conditions using the received data, and means for proposing a safe driving route based on the analysis results, thereby enabling the autonomous vehicle to select a safe and optimal driving route in real time, prevent accidents, and drive efficiently.

[1474] "Text" is a data format that expresses information using characters and symbols.

[1475] "Voice" refers to data relating to human speech or acoustics transmitted through sound waves.

[1476] An "image" is a still image that represents visual information in pixels.

[1477] "Video" refers to dynamic visual data that expresses movement using successive image frames.

[1478] "Preprocessing" refers to the initial processing of received data to convert it into a form suitable for analysis.

[1479] "Generative AI" is an artificial intelligence technology that learns from large datasets and generates new data.

[1480] "Artificial general intelligence" refers to advanced artificial intelligence that has the ability to solve a wide range of problems, not just specific tasks.

[1481] "Strategic advice" refers to recommendations based on data analysis that provide users with optimal courses of action.

[1482] An "autonomous vehicle" is a vehicle that has the ability to operate autonomously using artificial intelligence and various sensors.

[1483] "Camera footage" refers to video data captured by a camera and recorded as visual information.

[1484] "Sensor data" refers to data about physical phenomena measured by various sensors.

[1485] "Voice instructions" refer to commands or instructions given using voice.

[1486] "Road conditions" refers to environmental information that affects the safety and efficiency of driving, such as road congestion and road surface conditions.

[1487] A "travel route" is the optimal route to a destination.

[1488] The present invention is a system for analyzing operational data of autonomous vehicles to support safe and efficient operation. This system analyzes text, audio, image, and video data to generate and present strategic advice. Specific embodiments of this system are described below.

[1489] System configuration and operation

[1490] This system is mainly composed of components such as servers, terminals, and users, each of which fulfills its respective role to realize the overall function.

[1491] Server Roles

[1492] The server is the core of this system and has the following functions:

[1493] Data reception and preprocessing:

[1494] The server receives text, audio, image, and video data and pre-processes this data: for example, camera footage is processed frame by frame using OpenCV, and audio data is converted to text using the SpeechRecognition library.

[1495] Data Analysis:

[1496] The pre-processed data is then analyzed using deep learning models (generative AI and artificial general intelligence), for example, to detect road conditions and vehicle anomalies using deep learning frameworks such as TensorFlow and Keras.

[1497] Strategic advice generation:

[1498] Based on the analysis results, strategic advice is generated, such as suggesting safe driving routes and generating warning messages for drivers.

[1499] Present your proposal:

[1500] The server sends the generated advice to the terminal and presents it visually to the user. Operation information and advice are displayed in real time in a dashboard format.

[1501] Device Role

[1502] A terminal is a device that is directly operated by the user, such as a smartphone or tablet. A terminal has the following functions:

[1503] Data entry and collection:

[1504] It sends camera footage, sensor data, and voice instructions uploaded by the user to the server.

[1505] View the advice offered:

[1506] The terminal visually displays strategic advice sent from the server, allowing users to see driving routes and warning messages in real time.

[1507] User Roles

[1508] The user is the entity that uses the system and performs the following operations:

[1509] Upload data:

[1510] Users upload a variety of data to the system, including market analysis reports and interview videos.

[1511] Review the suggested advice:

[1512] Users can view the strategic advice and driving routes generated through a dashboard on their device and make decisions based on the results.

[1513] Examples of specific examples and prompts

[1514] When an autonomous vehicle is driving in a city and analyzing camera footage to detect obstacles and pedestrians ahead, the server uses the following prompt sentence:

[1515] Camera video analysis prompt: "Analyze the video data obtained from the vehicle's front camera to detect obstacles on the road."

[1516] When the vehicle selects a particular route based on the driver's voice instructions, the server uses the following prompt sentence:

[1517] Voice command analysis prompt: "Please convert the driver's current voice commands into text and analyze their actions."

[1518] This will enable autonomous vehicles to select safe and optimal routes in real time, prevent accidents, and operate efficiently.

[1519] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1520] Step 1:

[1521] The server receives camera footage, sensor data, and audio instructions from the user.

[1522] Input: Camera video files, sensor data files, audio files

[1523] Specific operation: Captures camera images as frame data, receives sensor data as numerical data, and recognizes audio files as audio data.

[1524] Step 2:

[1525] The server pre-processes the received data.

[1526] Input: Camera image frame data, sensor data values, audio data

[1527] Specific operations: Camera image frames are resized and preprocessed using OpenCV, sensor data is normalized to an appropriate range, and audio data is converted to text using the SpeechRecognition library.

[1528] Output: Preprocessed video data, normalized sensor data, and transcribed audio data

[1529] Step 3:

[1530] The server analyzes the pre-processed data using generative AI models and artificial general intelligence.

[1531] Input: Preprocessed video data, normalized sensor data, and transcribed audio data

[1532] Specific operation: Preprocessed video data is input into deep learning models (TensorFlow and Keras) to analyze road conditions, sensor data is analyzed, anomaly detection algorithms are applied, and voice data is analyzed to determine instructions.

[1533] Output: Analyzed road conditions, abnormal sensor data, voice instructions

[1534] Step 4:

[1535] The server generates strategic advice based on the analysis results.

[1536] Input: Analyzed road conditions, abnormal sensor data, voice instructions

[1537] Specific operation: Based on the analysis results, a strategic algorithm is applied to generate safe driving routes and warning messages for drivers.

[1538] Output: Generated route suggestions and warning messages for the driver

[1539] Step 5:

[1540] The server transmits the generated advice to the terminal.

[1541] Input: Generated route suggestions, warning messages for drivers

[1542] Specific operation: The generated advice is sent to the terminal and displayed on the terminal in real time.

[1543] Output: Route suggestions displayed on the terminal, warning messages to the driver

[1544] Step 6:

[1545] The terminal visually presents the generated advice to the user.

[1546] Input: Route suggestions displayed on the device, warning messages to the driver

[1547] Specific operation: Routes and warning messages are presented to the user in real time using a dashboard-style UI.

[1548] Output: User-visible route and warning messages

[1549] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1550] This invention relates to a system that generates and presents strategic advice by analyzing text, voice, image, and video data and combining it with an emotion engine that recognizes the user's emotions. This system is composed of components such as a server, a terminal, and a user, and each component fulfills its role to realize the overall function of the invention.

[1551] Specific system configuration and operation

[1552] 1. Data Collection

[1553] A user uploads data such as text, audio, images, and video to the system.

[1554] The server receives the uploaded data and stores it in a specific directory.

[1555] 2. Data Preprocessing

[1556] The server processes the received data using different pre-processing methods: for example, text data is formatted using natural language processing (NLP), audio data is converted to text using speech recognition technology, and image and video data is analyzed using image recognition technology.

[1557] 3. Data Analysis

[1558] The server then analyzes the pre-processed data using generative AI and artificial general intelligence (AGI) to discover important patterns, trends, and potential opportunities.

[1559] 4. Emotional Recognition

[1560] The server uses an emotion engine to recognize and analyze the user's emotions, for example, identifying the user's emotional state (e.g., joy, sadness, anger) from the user's voice data and text data.

[1561] 5. Generating strategic advice

[1562] The server generates specific strategic advice based on the analysis results, and provides advice that also takes into account the results of the user's sentiment analysis.

[1563] 6. Presentation to the User

[1564] The server provides a means to visually present the generated advice to the user, who can then access this information using their terminal and view the results in an easy to understand format.

[1565] Specific examples

[1566] For example, suppose a startup company is planning to launch a new product on the market. This company needs to understand the competitive situation and develop an effective marketing strategy. To do this, they use this system.

[1567] 1. The user uploads the market analysis report (PDF) and video of the consumer interview to the system.

[1568] 2. The server receives these data, extracts the text data, and converts the audio from the video to text.

[1569] 3. The server uses generative AI and AGI to analyze the data and analyse the competitive landscape and consumer opinion.

[1570] 4. The server uses an emotion engine to recognize and analyze consumer emotions from the interview audio data, for example, identifying which products consumers have positive emotions about.

[1571] 5. Based on the analysis, the server generates strategic advice that takes into account the user's emotional state, including how to enter new market segments and effective marketing messages to target customers.

[1572] 6. Users access the dashboard using their devices, review the proposed market strategies and analysis results, and make decisions based on the results.

[1573] In this way, the present invention provides a system that processes a variety of data formats in an integrated manner and provides useful strategic insights that take user sentiment into account. Users can easily upload data and receive analysis results in an intuitively understandable format, supporting quick and appropriate decision-making.

[1574] The processing flow will be explained below.

[1575] Step 1:

[1576] A user uploads data such as text, audio, images, and video to the system.

[1577] Example of specific operation: A user uploads a market analysis report (PDF) and a consumer interview video.

[1578] Step 2:

[1579] The server receives the uploaded data and stores it in a specific directory.

[1580] Example of specific operation: The server stores PDF files and video files in a designated folder.

[1581] Step 3:

[1582] The server initiates pre-processing based on the type of data.

[1583] Examples of how it works: Extracting text from PDF files and separating audio from video files.

[1584] Step 4:

[1585] The server uses speech recognition technology to convert the voice data into text.

[1586] Example of specific operation: Use a speech recognition engine to convert the audio of an interview video into text.

[1587] Step 5:

[1588] The server uses image recognition technology to extract important features from image and video data.

[1589] Example of specific operation: Extract key frames from a video file and analyze them using an image recognition model.

[1590] Step 6:

[1591] The server analyzes the pre-processed data using generative AI and artificial general intelligence.

[1592] Specific operation example: Apply natural language processing (NLP) technology to text data to perform keyword extraction and topic modeling.

[1593] Step 7:

[1594] The server uses an emotion engine to recognize and analyze the user's emotions.

[1595] Specific operation example: Identify the user's emotional state from voice data and text data, and determine joy, sadness, anger, etc.

[1596] Step 8:

[1597] The server synthesizes the analysis results and derives relevant insights.

[1598] Example of how it works: Integrate data from market analysis reports with text data from interviews to identify common trends and patterns.

[1599] Step 9:

[1600] The server generates strategic advice based on the analysis results.

[1601] Example of specific operation: Based on market trends, suggest ways to enter new market segments and create marketing messages that take into account the user's emotional state.

[1602] Step 10:

[1603] The server visualizes the generated advice and presents it to the user.

[1604] Example of how it works: Displays strategic advice and analytical results in graphs and charts in a dashboard format.

[1605] Step 11:

[1606] The terminal displays a dashboard to the user, allowing the user to review the results.

[1607] Example of how it works: A user opens a web browser and views the dashboard with suggested market segments and measures.

[1608] Example 2

[1609] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1610] Conventional data analysis systems lack the ability to process multiple data formats (text, audio, images, and videos) in a unified manner, making it particularly difficult to generate strategic advice that takes into account user sentiment analysis. Furthermore, they lacked a means to provide analysis results to users in an easy-to-understand format, making it difficult for users to make appropriate decisions.

[1611] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1612] In this invention, the server includes means for receiving text, audio, image, and video data, means for preprocessing the received data according to the type of data, means for analyzing the preprocessed data using a generation AI and an intelligent system, means for recognizing and analyzing a user's emotions, means for generating strategic advice based on the analysis results and emotion recognition results, and means for visually presenting the generated advice to the user. This enables unified analysis of multiple data formats and the provision of strategic advice that takes the user's emotions into consideration. Furthermore, visual presentation of the analysis results allows the user to easily understand and make appropriate decisions.

[1613] "Text data" is information expressed in the form of characters and sentences.

[1614] "Audio data" is information that records human voices and sounds in digital format.

[1615] "Image data" is visual information such as photographs and illustrations recorded in digital format.

[1616] "Video data" means a digital recording that combines moving visual and audio information.

[1617] "Means for receiving" refers to a technique or device for capturing data sent from a user within the system.

[1618] A "preprocessing means" is a technique or device for shaping or transforming collected data so that it is easier to analyze.

[1619] "Generative AI" refers to algorithms or systems that use artificial intelligence techniques to generate new information from data.

[1620] An "intelligent system" is an artificial intelligence system that has general knowledge and uses it to solve problems and analyze data.

[1621] "Emotion recognition and analysis means" refers to a technique or device for identifying and analyzing a user's emotional state from data provided by the user.

[1622] The "means for generating strategic advice" is a technology or device for generating guidelines or suggestions that are useful to the user based on the analysis results and emotion recognition results.

[1623] "Means for visually presenting to the user" refers to a technique or device for providing the analysis results and advice to the user in a visual format such as a graph or chart.

[1624] This invention relates to a system that generates and presents strategic advice by analyzing text, voice, image, and video data and combining it with an emotion engine that recognizes the user's emotions. This system is composed of components such as a server, a terminal, and a user, each of which fulfills a specific role to realize the overall function of the invention.

[1625] Data collection

[1626] Users upload data such as text, audio, images, and videos to the system. An interface for uploading this data is provided on the user's terminal. For example, a file can be selected by dragging and dropping using the file upload function of a web browser.

[1627] Data Preprocessing

[1628] The server receives the uploaded data and stores it in a specific directory. The received data is classified by type and the following preprocessing is performed depending on the data format:

[1629] Text data: A natural language processing (NLP) engine (e.g., SpaCy) is used to check grammar and filter out unwanted words.

[1630] Voice Data: We use a speech recognition engine (e.g., Google Speech-to-Text) to convert your voice into text.

[1631] Image data: An image recognition system (e.g., OpenCV) is used to identify objects in the image and generate metadata.

[1632] Video data: The video is broken down into frames, and each frame is analyzed using image recognition technology.

[1633] Data analysis

[1634] The server uses generative AI models (e.g., GPT-4) and intelligent systems to analyze the pre-processed data, allowing it to discover important patterns, trends, and potential opportunities, such as extracting competitor trends and identifying market niches based on market analysis reports.

[1635] Emotion recognition

[1636] The server uses an emotion recognition engine (e.g., IBM Watson Natural Language Understanding) to recognize and analyze emotions from the data provided by the user. It can identify a range of emotional states (such as joy, sadness, and anger) from text and audio data. For example, it can analyze the emotional responses from audio data of consumer interview videos to determine which products have the most positive opinions.

[1637] Strategic Advice Generation

[1638] The server generates strategic advice for users based on the analysis results and emotion recognition results. Specific suggestions are made using a generative AI model (e.g., GPT-4). For example, it can generate a suggestion such as, "Targeting a new product specifically for women in their 20s is likely to increase market share."

[1639] Presenting to the user

[1640] The server provides a dashboard to visually present the generated advice to the user. Users can access this dashboard from their own devices and view the analysis results in an intuitive, easy-to-understand format. The dashboard contains visual elements such as graphs, tables, and heat maps, and users can click on results to access more detailed information.

[1641] Specific examples

[1642] For example, suppose a startup company is planning to launch a new product. This company needs to understand the competitive situation and develop an effective marketing strategy. To do this, they use this system.

[1643] 1. The user uploads the market analysis report (PDF) and video of the consumer interview to the system.

[1644] 2. The server receives these data, extracts the text data, and converts the audio from the video to text.

[1645] 3. The server uses generative AI and intelligence systems to analyze the data and analyse competitors' situations and consumer opinions.

[1646] 4. The server uses an emotion recognition engine to recognize and analyze the consumer's emotions from the interview audio data, for example, to identify which products the consumer has positive feelings about.

[1647] 5. Based on the analysis, the server generates strategic advice that takes into account the user's emotional state, including how to enter new market segments and effective marketing messages to target customers.

[1648] 6. Users access the dashboard using their devices, review the proposed market strategies and analysis results, and make decisions based on the results.

[1649] Prompt Sentence Examples

[1650] "Based on this PDF report and video interviews, you can develop a market strategy and analyze consumer sentiment."

[1651] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1652] Step 1: Collect data

[1653] Input: User-uploaded text, audio, image, and video data

[1654] Specific operation: A user logs in to the system from their own terminal and uses the file upload function to upload data such as text, audio, images, and videos. Files are selected by drag-and-drop operation.

[1655] Output: Data stored on the server

[1656] Step 2: Save the data to the appropriate directory

[1657] Input: User uploaded data

[1658] Specific operation: The server receives the data uploaded by the user and saves it in a specific directory in real time. At this time, the data is temporarily stored in a buffer and then transferred to permanent storage after verification.

[1659] Output: Data stored in a specific directory on the server

[1660] Step 3: Preprocessing the data

[1661] Input: Raw data stored on the server

[1662] Specific operation: The server classifies the received data by type and performs appropriate preprocessing for each type. For example:

[1663] Text data: Use a natural language processing (NLP) engine (e.g., SpaCy) to check grammar and filter unnecessary words.

[1664] Voice Data: Converts speech to text using a speech recognition engine (e.g., Google Speech-to-Text).

[1665] Image data: An image recognition system (e.g., OpenCV) is used to identify objects in the image and generate metadata.

[1666] Video data: The video is broken down into frames, and each frame is analyzed using image recognition technology.

[1667] Output: Preprocessed data

[1668] Step 4: Data analysis

[1669] Input: Preprocessed data

[1670] What it does: The server analyzes the pre-processed data using generative AI models (e.g., GPT-4) and intelligent systems to discover important patterns, trends, and potential opportunities in the data.

[1671] Output: Analysis results (patterns, trends, potential opportunities)

[1672] Step 5: Recognize emotions

[1673] Input: Preprocessed text and audio data

[1674] Specific operation: The server uses an emotion recognition engine (e.g., IBM Watson Natural Language Understanding) to recognize and analyze emotions from the data provided by the user, and identifies a range of emotional states from the voice and text data.

[1675] Output: Emotion recognition results (happiness, sadness, anger, etc.)

[1676] Step 6: Generate strategic advice

[1677] Input: Analysis results and emotion recognition results

[1678] Specific operation: The server uses the generative AI model to generate specific strategic advice based on the analysis and emotion recognition results. For example, it generates a suggestion such as, "Targeting a new product specifically for women in their 20s is likely to increase market share."

[1679] Output: Strategic advice

[1680] Step 7: Present to the user

[1681] Input: Generated strategic advice

[1682] Specific operation: The server provides a dashboard to visually present the generated advice to the user. The user accesses this dashboard using their own device and understands the analysis results through visual elements such as graphs and charts.

[1683] Output: Visual analysis results and advice on a user-accessible dashboard

[1684] (Application example 2)

[1685] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1686] In typical brick-and-mortar stores, it is difficult for staff to accurately grasp customer emotions in real time and provide optimal product recommendations and services accordingly. Furthermore, staff need training and experience to deal with a wide variety of customers individually, and a lack of this training can lead to a decline in customer satisfaction. For this reason, there is a demand for support tools that enable staff to deal with customers quickly and accurately.

[1687] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving text, audio, image, and video data, means for preprocessing the received data, means for analyzing the preprocessed data using generative AI and artificial general intelligence, means for generating strategic advice based on the analysis results, means for presenting the generated advice to the user, means for recognizing emotions from the user's voice and video, and means for presenting advice generated in real time taking into account the emotion recognition results. This enables staff at physical stores to grasp customer emotions in real time and quickly provide optimal product suggestions and services based on that.

[1688] definition statement

[1689] "Text data" is data made up of character information.

[1690] "Audio data" refers to information collected as sound waves recorded in digital format.

[1691] "Image data" is data that stores visual information as a still image.

[1692] "Moving image data" is data that expresses dynamic visual information by combining successive still images along a time axis.

[1693] "Preprocessing" is the initial data processing to convert raw data into an analyzable format.

[1694] "Generative AI" is an artificial intelligence technology that generates new data and information based on input data.

[1695] "Artificial general intelligence" is artificial intelligence with a wide range of capabilities that can handle a wide variety of tasks.

[1696] An "emotion engine" is a technology that recognizes and analyzes a user's emotional state from voice, facial expressions, etc.

[1697] "Strategic advice" refers to specific and effective suggestions and instructions based on the results of analysis.

[1698] "Real-time presentation" means that analysis and advice are performed instantly, and information is provided to the user immediately.

[1699] MODE FOR CARRYING OUT THE INVENTION

[1700] The following describes an embodiment of the present invention.

[1701] The system includes means for receiving text, audio, image, and video data, means for preprocessing the received data, means for analyzing the preprocessed data using generative AI and artificial general intelligence, means for generating strategic advice based on the analysis results, means for presenting the generated advice to a user, means for recognizing emotions from the user's voice and video, and means for presenting the generated advice in real time, taking into account the emotion recognition results.

[1702] Data collection

[1703] Users use their smartphones to collect video and audio data of customers in physical stores via a camera and microphone, and these video and audio data are converted into text and image data.

[1704] Data Preprocessing

[1705] The server preprocesses the received video and audio data. Video data is converted into still images using an image recognition algorithm, and audio data is converted into text using speech recognition technology. Libraries and APIs such as OpenCV, Google Speech Recognition, and Emotion Recognition are used for preprocessing.

[1706] Data analysis

[1707] The pre-processed data is then analyzed using generative AI and artificial general intelligence to identify the products customers are interested in and their emotional state (interest, delight, doubt, etc.). The analysis is performed using OpenAI's API.

[1708] Emotion recognition

[1709] The user's emotional state is recognized from audio and video data. Emotion Recognition technology identifies the customer's emotional state and complements strategic advice based on this information.

[1710] Strategic Advice Generation

[1711] The server generates optimal strategic advice for the customer based on the analysis results and emotion recognition results, including product suggestions and service delivery methods that take into account the customer's emotions.

[1712] Presenting to the user

[1713] Strategic advice is presented to users in real time via a smartphone application, allowing wait staff to take immediate and appropriate action.

[1714] Specific examples

[1715] For example, if a customer asks for details about a particular product, the system might:

[1716] Customer question: "What are the features of this camera?"

[1717] Emotions recognized from customer facial expressions: Interest

[1718] Strategic advice: "This camera takes high-resolution photos and is waterproof. Plus, other customers often buy it."

[1719] Prompt Sentence Examples

[1720] Generate advice based on the user's statements and feelings below.

[1721] Say: "What is special about this camera?"

[1722] Emotion: "Interest"

[1723] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1724] Program processing steps

[1725] Step 1:

[1726] Users use the camera and microphone on their smartphones to collect video and audio data of customers in physical stores. The collected data is input as raw video frames and audio waveform data.

[1727] Step 2:

[1728] The terminal transmits the collected data to a server, which receives the video frames and audio waveform data, stores the video data as an image file, and stores the audio data as an audio file.

[1729] Step 3:

[1730] The server pre-processes the received data: video data is converted into still images using the OpenCV library, and audio data is converted into text using the Google Speech Recognition API. This pre-processing converts video data into an image format and audio data into a text format.

[1731] Step 4:

[1732] The server analyzes the preprocessed data using generative AI and artificial general intelligence (AGI). For the analysis, OpenAI's API is used to extract visual features from image data and customer interests and concerns from text data. This results in data feature quantities.

[1733] Step 5:

[1734] The server uses Emotion Recognition technology to recognize emotions from the user's video and audio data. It uses a facial expression recognition algorithm (e.g., OpenFace) from the video data and a voice emotion recognition algorithm from the audio data. This recognition yields the user's emotional state (interest, joy, etc.).

[1735] Step 6:

[1736] The server generates strategic advice based on the analysis results and emotion recognition results. It uses OpenAI's API to integrate this data and generate prompts. Based on these prompts, the server automatically generates advice that includes optimal product recommendations and service delivery methods for the customer.

[1737] Step 7:

[1738] The server then sends the generated strategic advice to the terminal, which then displays the received advice in real time on the smartphone screen so that the customer service staff can immediately check it.

[1739] Step 8:

[1740] The user (customer service staff) can then provide appropriate product recommendations and services to customers based on the strategic advice provided, which will enable quick and accurate customer service and is expected to improve customer satisfaction.

[1741] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1742] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1743] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1744] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1745] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1746] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1747] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1748] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1749] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1750] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1751] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1752] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1753] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1754] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1755] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1756] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1757] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1758] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1759] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1760] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1761] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1762] The following is further disclosed regarding the above embodiment.

[1763] (Claim 1)

[1764] means for receiving text, audio, image, and video data;

[1765] means for preprocessing the received data;

[1766] A means for analyzing the pre-processed data using generative AI and artificial general intelligence;

[1767] means for generating strategic advice based on the analysis results;

[1768] means for presenting the generated advice to a user;

[1769] A system including:

[1770] (Claim 2)

[1771] 10. The system of claim 1, wherein different preprocessing is performed depending on the type of data.

[1772] (Claim 3)

[1773] 2. The system according to claim 1, wherein the analysis results are visualized and presented to the user.

[1774] "Example 1"

[1775] (Claim 1)

[1776] means for receiving data;

[1777] means for preprocessing the received data based on different formats;

[1778] A means for analyzing the preprocessed data using artificial intelligence and general artificial intelligence;

[1779] means for generating strategic advice based on the analysis results;

[1780] a means for visually presenting the generated advice to a user;

[1781] A system including:

[1782] (Claim 2)

[1783] 2. The system according to claim 1, wherein preprocessing is performed to extract text from text data using natural language processing technology and convert voice data into character information using voice recognition technology.

[1784] (Claim 3)

[1785] 10. The system of claim 1, further comprising means for displaying the analysis results on a dashboard.

[1786] "Application Example 1"

[1787] (Claim 1)

[1788] means for receiving text, audio, image, and video data;

[1789] means for preprocessing the received data;

[1790] A means for analyzing the pre-processed data using generative AI and artificial general intelligence;

[1791] means for generating strategic advice based on the analysis results;

[1792] means for presenting the generated advice to a user;

[1793] means for receiving camera footage, sensor data, and voice instructions from the autonomous vehicle;

[1794] means for analyzing road conditions using the received data;

[1795] A means of proposing safe driving routes based on the analysis results;

[1796] A system including:

[1797] (Claim 2)

[1798] 10. The system of claim 1, wherein different preprocessing is performed depending on the type of data.

[1799] (Claim 3)

[1800] 2. The system according to claim 1, wherein the analysis results are visualized and presented to the user.

[1801] "Example 2: Combining Emotion Engines"

[1802] (Claim 1)

[1803] means for receiving text, audio, image, and video data;

[1804] means for preprocessing the received data according to the type of data;

[1805] means for analyzing the pre-processed data using generative AI and intelligent systems;

[1806] means for recognizing and analyzing user emotions;

[1807] means for generating strategic advice based on the analysis results and emotion recognition results;

[1808] means for visually presenting the generated advice to the user;

[1809] A system including:

[1810] (Claim 2)

[1811] 10. The system of claim 1, wherein different preprocessing is performed depending on the type of data.

[1812] (Claim 3)

[1813] 2. The system according to claim 1, wherein strategic advice based on the analysis results and emotion recognition results is visualized and presented to the user.

[1814] "Application example 2 when combining emotion engines"

[1815] Claims

[1816] (Claim 1)

[1817] means for receiving text, audio, image, and video data;

[1818] means for preprocessing the received data;

[1819] A means for analyzing the pre-processed data using generative AI and artificial general intelligence;

[1820] means for generating strategic advice based on the analysis results;

[1821] means for presenting the generated advice to a user;

[1822] means for recognizing emotions from the user's voice and video;

[1823] a means for presenting advice generated in real time taking into account the emotion recognition results;

[1824] A system including:

[1825] (Claim 2)

[1826] 10. The system of claim 1, wherein different preprocessing is performed depending on the type of data.

[1827] (Claim 3)

[1828] 2. The system according to claim 1, wherein the analysis results are visualized and presented to the user. [Explanation of symbols]

[1829] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving text, audio, image, and video data; means for preprocessing the received data; A means for analyzing the pre-processed data using generative AI and artificial general intelligence; means for generating strategic advice based on the analysis results; means for presenting the generated advice to a user; A system including:

2. The system of claim 1 , wherein different pre-processing is performed depending on the type of data.

3. The system according to claim 1, wherein the analysis results are visualized and presented to the user.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A