system
A battlefield system integrates AI technologies for real-time data processing and decision-making, addressing the challenge of rapid and accurate tactical analysis and response.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-17
- Publication Date
- 2026-04-30
AI Technical Summary
In battlefield scenarios, the vast amount of diverse information generated makes it difficult for humans to quickly and accurately analyze and make strategic decisions, leading to delayed judgments and reduced tactical effectiveness.
A system that integrates information collection, analysis, decision-making, and action execution components, utilizing AI technologies to process video, audio, and text data in real-time, enabling rapid and accurate tactical decision-making.
The system accelerates information processing and decision-making on the battlefield, securing a strategic advantage by providing real-time, accurate, and flexible tactical responses.
Smart Images

Figure 2026071622000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] On the battlefield, diverse and large amounts of information are constantly generated, and it is very difficult for humans to quickly and accurately judge it. Therefore, there are problems that strategic decision-making is delayed due to information delay and misjudgment, and the effectiveness of tactics is reduced. In particular, there is a need for the ability to integrally analyze data obtained from multiple information sources and make decisions quickly.
Means for Solving the Problems
[0005] This invention provides a system that collects, analyzes, and integrates multiple data formats such as video, audio, and text in real time, and determines and executes the optimal action according to the situation. By including information collection means, analysis means, decision-making means, action execution means, and control means that link them together, this system can process information on the battlefield quickly and accurately and secure a tactical advantage.
[0006] "Information gathering means" refers to components equipped with the ability to acquire multiple data formats, such as video, audio, and text, in real time on the battlefield.
[0007] An "analysis tool" is a component that has the function of understanding the situation by analyzing and integrating collected data.
[0008] A "decision-making tool" is a component that has the function of determining the optimal action based on analyzed data.
[0009] "Means of action execution" refers to components equipped with the function of concretely carrying out a decided action.
[0010] "Control means" refers to a component equipped with functions for coordinating information gathering means, analysis means, decision-making means, and action execution means. [Brief explanation of the drawing]
[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5]This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0013] First, let's explain the terminology used in the following explanation.
[0014] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0015] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0016] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0017] In the following embodiments, the labeled communication I / F (Interface) is an interface that includes a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0019] [First Embodiment]
[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0032] The system of the present invention aims to process information quickly and accurately on the battlefield and execute optimal actions. This system comprises information gathering means, analysis means, decision-making means, action execution means, and control means for coordinating and coordinating each of these means. Embodiments for carrying out the present invention are described below.
[0033] Information gathering methods acquire data from various sensors, cameras, and microphones deployed on the battlefield. This allows the server to collect diverse information, including video, audio, and text data, in real time.
[0034] The analysis method processes the collected data using specialized AI technology. The server uses image recognition algorithms to identify the enemy's location and movement from the video data, and speech recognition technology to convert the enemy's communications into text data. Natural language processing (NLP) is applied to the text data to analyze the enemy's intentions and situation.
[0035] The decision-making mechanism uses generative AI to assess the situation based on the analysis results and select the optimal action. The server compares past data with the current situation and makes tactical predictions to accurately predict the enemy's next move.
[0036] The execution mechanism takes specific actions based on instructions from the server. The terminal provides instructions to friendly units in real time and visualizes the information needed to take necessary actions on the terminal. This allows users (friendly units) to execute instructions quickly and efficiently.
[0037] As a concrete example, suppose a server analyzes video data collected from a drone to understand the movement patterns of enemy vehicles in real time. The server then predicts the enemy's escape routes and transmits this information to friendly terminals. Based on the received data, the terminals can deploy friendly forces to the optimal positions, making it possible to more efficiently thwart enemy actions.
[0038] As described above, the present invention provides a system that can accelerate information processing and decision-making on the battlefield and secure a strategic advantage.
[0039] The following describes the processing flow.
[0040] Step 1:
[0041] The server collects video, audio, and text data in real time from sensors and drones deployed on the battlefield. This includes streaming video from fixed cameras and recordings of ambient and communication sounds. The server centrally acquires this data and prepares it for processing.
[0042] Step 2:
[0043] The server preprocesses the collected data. Video data is divided frame by frame, and image filters are applied to remove unnecessary noise. Audio data is converted into a clear audio signal using noise cancellation technology and then into a format suitable for speech recognition. Text data is formatted into an appropriate format and prepared for analysis.
[0044] Step 3:
[0045] The server analyzes data using various AI models. Specifically, it uses image recognition algorithms to automatically detect enemy vehicles and people from video footage and map their locations. It uses speech recognition models to transcribe enemy communications into text and analyze their content. The text data is then processed using natural language processing to understand the intent behind commands and reports.
[0046] Step 4:
[0047] The server integrates multiple analysis results to model the overall situation. It utilizes a Geographic Information System (GIS) to map the positions and movement patterns of both friendly and enemy forces. It also conducts simulations to predict enemy behavior patterns, taking into account temporal changes.
[0048] Step 5:
[0049] The server uses generative AI to determine the optimal action plan based on comprehensive data. It simulates multiple tactical options and selects the most effective course of action. It evaluates risks and benefits and completes the decision-making process.
[0050] Step 6:
[0051] The terminal transmits information to friendly commanders and soldiers in real time based on tactical instructions received from the server. It visualizes the instructions by updating map information and the positions of friendly and enemy forces on the display device. It also provides specific action instructions via voice or message as needed to encourage a rapid response.
[0052] (Example 1)
[0053] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0054] On-site information processing requires handling a wide variety of data formats simultaneously, integrating and analyzing them in real time, and making rapid decisions. However, there are difficulties in establishing appropriate infrastructure for processing diverse data and in rapidly implementing action plans based on analysis results. Furthermore, systems that improve the accuracy of decision-making while communicating instructions in a way that users can immediately execute are also insufficient.
[0055] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0056] In this invention, the server includes means for acquiring multiple pieces of information, including video data, audio data, and text data, in real time via a detection device at the site; preprocessing means for preprocessing the acquired information, reducing noise, and shaping it for analysis; and analysis means for analyzing the preprocessed data, identifying the enemy's movements and location using image recognition and speech recognition technologies, and analyzing their intentions using natural language processing technologies. This enables rapid and accurate processing of diverse data, allowing for immediate action decisions and the transmission of those instructions.
[0057] "Video data" refers to visual information captured by cameras or other visual detection devices.
[0058] "Audio data" refers to acoustic information recorded by microphones or other sound detection devices.
[0059] "Text data" refers to information that has been converted into a string of characters through speech recognition or other methods.
[0060] A "detection device" refers to equipment installed on-site to acquire various types of data, such as video, audio, temperature, and motion.
[0061] "Preprocessing means" refers to the process of removing noise from acquired data and converting it into a format suitable for analysis.
[0062] "Analysis methods" refer to techniques that use pre-processed data and employ machine learning algorithms and AI technologies to examine information in detail.
[0063] "Decision-making tools" refer to the process of making decisions to predict future events by utilizing analyzed information.
[0064] "Action execution means" refers to a method of generating specific instructions based on a decided action plan and transmitting them to a terminal in an executable format.
[0065] "Management means" refers to a system for efficiently controlling the overall process by linking the processes of information acquisition, preprocessing, analysis, decision-making, and execution.
[0066] This invention relates to a system that efficiently processes diverse information in the field, enabling rapid decision-making and its implementation. This system supports strategic actions by combining means with multiple functions.
[0067] The server acquires video, audio, and text data in real time using detection devices installed on-site. The hardware used includes cameras, microphones, and various sensors. Software-wise, an integrated data acquisition application works in conjunction with these devices to efficiently collect data.
[0068] The server performs data cleaning on the acquired data using preprocessing mechanisms. Specifically, it removes background noise from audio data and cuts video data to the required frames. This process utilizes advanced algorithms and AI-assisted technologies.
[0069] As an analysis method, the server uses a generative AI model and employs image recognition and speech recognition technologies to identify the enemy's location and movements. Furthermore, natural language processing is performed on text data to analyze the enemy's intentions and tactics. These analyses require dedicated AI toolkits and libraries.
[0070] In decision-making, the server predicts the enemy's next move by comparing a vast dataset of past data with current information. This prediction process utilizes generative AI models to formulate the optimal strategy for each situation.
[0071] In the execution mechanism, the terminal receives instructions from the server and visually provides the user (friendly force) with an action plan. The terminal is equipped with an intuitive interface that displays enemy location information and predicted movements on a map.
[0072] A concrete example is the process where a server analyzes video data collected from a drone to determine the current location and predicted movement path of enemy vehicles. Based on this data, the server transmits information to terminals, which then provide allied forces with immediate visualization of how to deploy. This supports optimal decision-making based on the situation.
[0073] The following is an example of a prompt message to input into the generative AI model.
[0074] "Input data: Video data, audio data, text data. Objective: Identify enemy location and movement, analyze communication content, and predict next actions."
[0075] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0076] Step 1:
[0077] The server acquires video, audio, and text data from detection devices in real time. Input includes raw data transmitted from cameras and microphones installed on-site. The server uses an interface to acquire and store this data in a database. Output is the prepared raw dataset.
[0078] Step 2:
[0079] The server preprocesses the acquired data. The input is the raw data obtained in step 1. Specifically, it performs noise reduction on the audio data and extracts the necessary frames from the video data. This generates a dataset suitable for analysis. The output is denoised audio data and formatted video data.
[0080] Step 3:
[0081] The server inputs the pre-processed data into the generating AI model for analysis. The input is the data processed in step 2. The server applies an image recognition algorithm to identify the enemy's movement and location in the video, and performs speech recognition on the audio data to convert the information into text. It also uses natural language processing to understand the intent. The output provides analysis results regarding the enemy's location, movement, and intent.
[0082] Step 4:
[0083] The server uses the analysis results to execute the decision-making process. The input is the analysis results obtained in step 3. The server compares this with past database data and uses a generated AI model to predict the enemy's next move. The output generates an optimal action plan or prediction.
[0084] Step 5:
[0085] The server communicates the execution plan to the terminal. The input is the action plan formulated in step 4. The server sends information about the action plan and predicted enemy movements to the terminal. The output is the instruction information displayed on the terminal.
[0086] Step 6:
[0087] The terminal transmits information received from the server to the user. The input is the instruction information received in step 5. The terminal visualizes this information on a map and provides an interface to prompt the user to take quick action. As output, information for the user to take action is visualized.
[0088] (Application Example 1)
[0089] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0090] Security monitoring in public and commercial facilities currently relies on traditional methods such as security cameras and visual checks by security guards, which can lead to delays in response. Therefore, there is a need for means to quickly detect potential dangers and ensure safety.
[0091] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0092] In this invention, the server includes information gathering means for collecting multiple pieces of information, including video, audio, and text, in real time; analysis means for analyzing and integrating the collected information; and safety monitoring means for analyzing information related to safety and identifying hazards. This enables real-time monitoring of safety conditions and rapid response.
[0093] "Information gathering means" refers to a device or method for collecting multiple pieces of information, such as images, sounds, and text, in real time.
[0094] "Analysis means" refers to a process or technique for processing and integrating collected information to understand a situation.
[0095] A "decision-making tool" is a system that serves as the foundation for determining the optimal course of action based on analysis results.
[0096] An "action execution mechanism" is a mechanism for actually carrying out a decided action and transmitting instructions.
[0097] A "control system" is a system that integrates and integrates various means of information gathering, analysis, decision-making, and action execution to function as a single unit.
[0098] "Safety monitoring measures" refer to methods and techniques for monitoring safety within a facility and identifying potential hazards.
[0099] This system is designed for safety monitoring in public and commercial facilities, and functions through the collaboration of a server, terminals, and users. The server first acquires data from various video and acoustic sensors. This is supported by an information gathering system that has the function of aggregating video, audio, and text data in real time.
[0100] Next, the server processes the collected information using analysis tools. In this step, the video data is analyzed using image processing libraries such as OpenCV, and object detection is performed using the Hugging Face Transformers library. Based on the collected information, AI technology is used to identify potential risks.
[0101] Based on the analysis results, the server's decision-making system evaluates and determines the optimal security response. As a result, if a threat is identified, it can immediately send a warning or issue instructions to security staff through the action execution system. This enables a rapid response and the implementation of appropriate security measures.
[0102] For example, a system can continuously monitor camera feeds within a shopping mall, detect suspicious behavior or dangers in real time, and alert the security team. Such a system would also be useful in crowded conditions during holidays or events.
[0103] The generative AI model can be input using prompts like the following:
[0104] "Analyze real-time camera feeds from within the shopping mall to identify suspicious individuals and congestion levels."
[0105] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0106] Step 1:
[0107] The server acquires real-time data from video and acoustic sensors within the facility. This input data is collected by the server's information gathering mechanism, and image data is sent to the server as frames, while acoustic data is sent as waveforms. This allows for the centralized aggregation of diverse information, preparing it for subsequent analysis.
[0108] Step 2:
[0109] The server processes the collected data using analytical tools. For video data, it analyzes each frame using image processing libraries such as OpenCV, and identifies suspicious individuals or abnormal behavior using an object detection model. For audio data, it identifies abnormal sound patterns using a speech recognition algorithm. The output is the analysis result, including matches with suspicious behavior and abnormal sounds.
[0110] Step 3:
[0111] The server selects the optimal countermeasure using a decision-making mechanism based on the analysis results. In this step, a generative AI model is used to determine the alert level based on existing data and the current situation, and to select the necessary response. The input is the analysis results, and the output is the selection of the countermeasure.
[0112] Step 4:
[0113] The terminal transmits the selected response measures as specific instructions to security staff through the action execution mechanism. These instructions are displayed on the terminal both audibly and visually, allowing security staff to quickly implement on-site responses. The input is the response measures from the decision-making mechanism, and the output is the specific instructions.
[0114] Step 5:
[0115] The security staff, as users, carry out actual security-enhancing actions on-site based on instructions from the terminal. Based on the specific instructions from the terminal, the security staff can take flexible measures according to the situation on site. Here, the input is instructions received from the terminal, and the output is the execution action for ensuring security.
[0116] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0117] In addition to the conventional functions of collecting and analyzing diverse data on the battlefield, the system of the present invention also has the function of recognizing the user's emotions. This system incorporates an emotion engine in addition to information gathering means, analysis means, decision-making means, action execution means, and control means, thereby enabling the optimization of tactics that take into account the user's emotional state.
[0118] Information gathering methods include not only data from various sensors, cameras, and microphones on the battlefield, but also emotional data collected through user biometric signals and voice surveys. This allows the server to centrally acquire a wide range of information, forming a foundation for analysis.
[0119] The analysis method processes the collected data using AI technology. The server analyzes video data using image recognition, communication sounds using speech recognition, and text using natural language processing. In addition, it recognizes the user's emotional state by analyzing the intonation of their voice and biometric data using an emotion engine.
[0120] The decision-making process involves using a generated AI to assess the situation based on integrated analysis results and determine the optimal course of action. In particular, it detects the stress and distress that users encounter on the battlefield based on emotional data and proposes options and support methods to mitigate them.
[0121] The action execution mechanism sends instructions to the terminal to carry out tactical actions selected by the server. The terminal provides the user with specific instructions in real time and visualizes information that is adjusted to minimize user stress.
[0122] For example, if a server analyzes a user's heart rate and voice patterns using an emotion engine and determines that the user is in a high-stress state, the decision-making system has the ability to suggest low-risk options. Depending on the device, the user can receive this information in an easily accessible format, allowing them to choose the optimal course of action while reducing psychological burden.
[0123] In this way, the present invention can provide a system that not only accelerates information processing and decision-making on the battlefield but also enables flexible responses that take into account the psychological state of the human being.
[0124] The following describes the processing flow.
[0125] Step 1:
[0126] The server acquires video, audio, and text data from battlefield sensors, cameras, and microphones, while simultaneously collecting biometric signals such as heart rate and voice intonation from users' mobile devices and wearables. This also allows for the acquisition of user emotional data.
[0127] Step 2:
[0128] The server preprocesses the collected data. Video data is de-noised and converted into an easily analyzable format. Audio data is converted to text using speech recognition technology. For emotional data, an emotion engine is used to numerically represent the user's current emotional state (stress, joy, anxiety, etc.).
[0129] Step 3:
[0130] The server analyzes data using various AI models. Image recognition models identify the locations of enemies and obstacles, important keywords are extracted from audio data, and intentions and commands are analyzed from text data using natural language processing. The emotion engine evaluates how the user's emotions affect the overall situation.
[0131] Step 4:
[0132] The server integrates the analyzed data and builds a comprehensive situational model. Using GIS, it maps the positions of friendly and enemy forces along with the users' emotional states to grasp the overall tactical situation. Based on this situational model, the generative AI performs tactical predictions and simulations.
[0133] Step 5:
[0134] The server uses generative AI to develop the optimal action plan. In doing so, it takes the user's emotional state into consideration, selecting tactics to reduce risk and methods to provide emotional support to the user in high-stress situations. The plan is then prepared as an action plan.
[0135] Step 6:
[0136] The terminal receives action plans sent from the server and provides instructions to the user in real time. These instructions include maps that visually represent the situation and suggestions for mental health support as mitigation measures. With this support, the user can make appropriate decisions and take action.
[0137] (Example 2)
[0138] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0139] Existing tactical support systems have limitations in real-time situational awareness and decision-making, and have not adequately optimized to consider the psychological and physiological state of users on the battlefield. Furthermore, instructions given to users can be burdensome, hindering rapid and reliable decision-making. Therefore, there is a need for more flexible and advanced tactical support that takes psychological and physiological states into account.
[0140] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0141] In this invention, the server includes information gathering means for collecting multiple pieces of information, including video, audio, text, and biosignals, in real time; analysis means for analyzing the collected information and generating integrated analysis results, including the user's emotional state; and decision-making means for recognizing the situation based on the integrated analysis results and determining the optimal action using a generated AI model. This enables the selection of rapid and appropriate tactical actions while taking the user's emotional state into consideration.
[0142] "Information gathering means" refers to means for collecting multiple types of information in real time, including video, audio, text, and biosignals, using measuring instruments deployed in a tactical area.
[0143] "Analysis means" refers to the means of analyzing collected information and generating integrated analysis results, including the user's emotional state.
[0144] A "decision-making tool" is a means of recognizing a situation based on integrated analysis results and determining the optimal action using a generative AI model.
[0145] "Action execution means" refers to the means for carrying out a determined action and providing the user with predetermined instructions.
[0146] "Control means" refers to means for coordinating information gathering means, analysis means, decision-making means, and action execution means.
[0147] A "generative AI model" is an artificial intelligence model used to determine the optimal course of action by utilizing integrated analysis results.
[0148] A "prompt statement" is a sentence of text containing instructions or information that is input to a generative AI model.
[0149] The embodiments for carrying out the present invention are shown below.
[0150] This system supports information gathering, analysis, decision-making, and action execution in tactical areas. The server uses information gathering tools to collect diverse real-time data from sensors, cameras, and microphones. This data includes images, audio, text, and biometric signals, as well as information about the tactical area's situation and the user's emotional state.
[0151] The server processes the collected data using analytical tools. This analysis utilizes image recognition software, speech recognition software, and a natural language processing engine. Specifically, it recognizes image data as visual information and converts speech data into text. It also uses an emotion engine to analyze the user's emotions from biometric data and voice input.
[0152] Furthermore, it incorporates a decision-making mechanism using a generative AI model, which makes situational judgments and determines the optimal action based on integrated analysis results. This allows for the formation of safe and effective tactical actions that take into account the user's psychological state. As a concrete example of a prompt, the generative AI model is input with instructions such as, "Recommend the optimal action when the user is in a high-stress state."
[0153] Upon receiving instructions from the server, the terminal uses its execution mechanisms to provide intuitive and easy-to-understand instructions to the user. The terminal communicates information to the user through visual displays and voice, supporting real-time decision-making.
[0154] In this way, the system of the present invention provides advanced information processing and situational response capabilities in the tactical zone. Furthermore, it can reduce the psychological burden on users and enable efficient decision-making. This system realizes cutting-edge tactical support utilizing information technology and artificial intelligence.
[0155] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0156] Step 1:
[0157] The server acquires data from sensors, cameras, and microphones in the tactical area using information gathering tools. Inputs include a variety of data such as images, audio, and biosignals. This data is collected in real time and temporarily stored.
[0158] Step 2:
[0159] The server preprocesses the acquired data into a format that can be analyzed. Specifically, it adjusts image resolution and removes noise, filters audio data, and tokenizes text data. The preprocessed data is then used as input for the next analysis step.
[0160] Step 3:
[0161] The server analyzes the pre-processed data using analytical tools. In this step, visual information is analyzed using image recognition software, and audio data is converted to text using speech recognition software. Furthermore, an emotion engine is used to determine the user's emotional state from biometric data. The analysis results are output as integrated information.
[0162] Step 4:
[0163] The server uses a generative AI model to make decisions based on the integrated analysis results. At this stage, a prompt message such as "Recommend the best course of action when the user is in a high-stress state" is input, and the generative AI proposes the optimal tactical action. As a result, optimized action instructions are output.
[0164] Step 5:
[0165] The server sends action instructions to the terminal. The terminal receives this information and issues instructions to the user using the means of execution. Specifically, this includes providing visual information on a visual display and audio guidance. Based on this information, the user can take appropriate action.
[0166] (Application Example 2)
[0167] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0168] In traditional brick-and-mortar stores, accurately understanding customers' emotional states and providing timely service accordingly has been difficult. While the importance of emotion-based responses in improving customer satisfaction and promoting purchases is recognized, there is a lack of concrete methods for putting this into practice.
[0169] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0170] In this invention, the server includes information acquisition means for collecting the customer's emotional state in real time through biosignal data and voice data, analysis means for analyzing the collected emotional data and determining the customer's psychological state, and decision-making means for determining the optimal response policy based on the determined emotional state. This makes it possible to provide optimal service tailored to the customer's emotional state in the store.
[0171] "Information acquisition means" refers to means for collecting customers' emotional states in real time through biosignal data and voice data.
[0172] "Analysis methods" refer to means of analyzing collected emotional data to determine the customer's psychological state.
[0173] A "decision-making tool" is a means of determining the optimal course of action based on the identified emotional state.
[0174] "Means of action" refer to the means of implementing the decided response policy and sending appropriate instructions to service providers.
[0175] "Control means" refers to means that coordinate information acquisition means, analysis means, decision-making means, and action execution means to improve services in the store environment.
[0176] This invention realizes a system that integrates emotion recognition and decision-making processes to improve customer service in physical stores. The server collects biosignal data and voice data acquired from sensors and voice recognition devices in real time and functions as an information acquisition means. This allows data on the emotional state of customers to be obtained.
[0177] The server analyzes the collected data and uses analytical methods to determine the customer's psychological state. This analysis employs artificial intelligence technology, specifically machine learning libraries such as TENSORFLOW® and PyTorch. Furthermore, Google® Cloud Speech-to-Text may be used for speech recognition. This allows for accurate detection of the customer's emotional state.
[0178] Based on the analysis results, a decision-making mechanism is activated to determine the optimal customer service strategy. This decision-making process utilizes a generative AI model to propose appropriate service methods tailored to the customer's emotional state.
[0179] For example, if a customer appears anxious, the server might send an instruction to the service provider such as, "Please explain the current promotion again to alleviate the customer's anxiety."
[0180] Another example of a prompt message is, "If the customer is feeling anxious, please suggest the best course of action to take." This allows store staff to provide service that responds to the customer's emotions in real time, thereby improving customer satisfaction.
[0181] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0182] Step 1:
[0183] The server uses sensors and voice recognition devices to collect customer biometric and voice data in real time. The sensors detect data such as pulse rate and facial expressions, while the voice recognition device records the customer's voice. This biometric and voice data is used as input, and the output is raw data ready for analysis.
[0184] Step 2:
[0185] The server processes the collected data using analytical tools. Specifically, it uses machine learning models such as TensorFlow and PyTorch to perform advanced data analysis to determine the customer's emotional state. The input data is the raw data obtained in step 1, and the output is the analysis result indicating the customer's emotional state. This analysis result estimates the customer's psychological state and is used in the next decision-making step.
[0186] Step 3:
[0187] Based on the analysis results, the server uses decision-making tools to determine the optimal course of action. A generative AI model is used to generate prompt messages that match the customer's emotional state. The input is the analysis results from step 2, and the output is a concrete action plan. This action plan contributes to improving customer service and increasing satisfaction.
[0188] Step 4:
[0189] The server transmits the determined action plan to the terminal via the action execution mechanism. The terminal displays instructions to the service provider in the form of a prompt message, such as "Please explain the current promotion again to alleviate the customer's concerns." The input data is the action plan obtained in step 3, and the output is specific instructions for the service provider. This enables the user to respond in real time on-site.
[0190] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0191] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0192] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0193] [Second Embodiment]
[0194] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0195] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0196] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0197] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0198] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0199] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0200] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0201] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0202] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0203] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0204] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0205] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0206] The system of the present invention aims to process information quickly and accurately on the battlefield and execute optimal actions. This system comprises information gathering means, analysis means, decision-making means, action execution means, and control means for coordinating and coordinating each of these means. Embodiments for carrying out the present invention are described below.
[0207] Information gathering methods acquire data from various sensors, cameras, and microphones deployed on the battlefield. This allows the server to collect diverse information, including video, audio, and text data, in real time.
[0208] The analysis method processes the collected data using specialized AI technology. The server uses image recognition algorithms to identify the enemy's location and movement from the video data, and speech recognition technology to convert the enemy's communications into text data. Natural language processing (NLP) is applied to the text data to analyze the enemy's intentions and situation.
[0209] The decision-making mechanism uses generative AI to assess the situation based on the analysis results and select the optimal action. The server compares past data with the current situation and makes tactical predictions to accurately predict the enemy's next move.
[0210] The execution mechanism takes specific actions based on instructions from the server. The terminal provides instructions to friendly units in real time and visualizes the information needed to take necessary actions on the terminal. This allows users (friendly units) to execute instructions quickly and efficiently.
[0211] As a concrete example, suppose a server analyzes video data collected from a drone to understand the movement patterns of enemy vehicles in real time. The server then predicts the enemy's escape routes and transmits this information to friendly terminals. Based on the received data, the terminals can deploy friendly forces to the optimal positions, making it possible to more efficiently thwart enemy actions.
[0212] As described above, the present invention provides a system that can accelerate information processing and decision-making on the battlefield and secure a strategic advantage.
[0213] The following describes the processing flow.
[0214] Step 1:
[0215] The server collects video, audio, and text data in real time from sensors and drones deployed on the battlefield. This includes streaming video from fixed cameras and recordings of ambient and communication sounds. The server centrally acquires this data and prepares it for processing.
[0216] Step 2:
[0217] The server preprocesses the collected data. Video data is divided frame by frame, and image filters are applied to remove unnecessary noise. Audio data is converted into a clear audio signal using noise cancellation technology and then into a format suitable for speech recognition. Text data is formatted into an appropriate format and prepared for analysis.
[0218] Step 3:
[0219] The server analyzes data using various AI models. Specifically, it uses image recognition algorithms to automatically detect enemy vehicles and people from video footage and map their locations. It uses speech recognition models to transcribe enemy communications into text and analyze their content. The text data is then processed using natural language processing to understand the intent behind commands and reports.
[0220] Step 4:
[0221] The server integrates multiple analysis results to model the overall situation. It utilizes a Geographic Information System (GIS) to map the positions and movement patterns of both friendly and enemy forces. It also conducts simulations to predict enemy behavior patterns, taking into account temporal changes.
[0222] Step 5:
[0223] The server uses generative AI to determine the optimal action plan based on comprehensive data. It simulates multiple tactical options and selects the most effective course of action. It evaluates risks and benefits and completes the decision-making process.
[0224] Step 6:
[0225] The terminal transmits information to friendly commanders and soldiers in real time based on tactical instructions received from the server. It visualizes the instructions by updating map information and the positions of friendly and enemy forces on the display device. It also provides specific action instructions via voice or message as needed to encourage a rapid response.
[0226] (Example 1)
[0227] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0228] On-site information processing requires handling a wide variety of data formats simultaneously, integrating and analyzing them in real time, and making rapid decisions. However, there are difficulties in establishing appropriate infrastructure for processing diverse data and in rapidly implementing action plans based on analysis results. Furthermore, systems that improve the accuracy of decision-making while communicating instructions in a way that users can immediately execute are also insufficient.
[0229] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0230] In this invention, the server includes means for acquiring multiple pieces of information, including video data, audio data, and text data, in real time via a detection device at the site; preprocessing means for preprocessing the acquired information, reducing noise, and shaping it for analysis; and analysis means for analyzing the preprocessed data, identifying the enemy's movements and location using image recognition and speech recognition technologies, and analyzing their intentions using natural language processing technologies. This enables rapid and accurate processing of diverse data, allowing for immediate action decisions and the transmission of those instructions.
[0231] "Video data" refers to visual information captured by cameras or other visual detection devices.
[0232] "Audio data" refers to acoustic information recorded by microphones or other sound detection devices.
[0233] "Text data" refers to information that has been converted into a string of characters through speech recognition or other methods.
[0234] A "detection device" refers to equipment installed on-site to acquire various types of data, such as video, audio, temperature, and motion.
[0235] "Preprocessing means" refers to the process of removing noise from acquired data and converting it into a format suitable for analysis.
[0236] "Analysis methods" refer to techniques that use pre-processed data and employ machine learning algorithms and AI technologies to examine information in detail.
[0237] "Decision-making tools" refer to the process of making decisions to predict future events by utilizing analyzed information.
[0238] "Action execution means" refers to a method of generating specific instructions based on a decided action plan and transmitting them to a terminal in an executable format.
[0239] "Management means" refers to a system for efficiently controlling the overall process by linking the processes of information acquisition, preprocessing, analysis, decision-making, and execution.
[0240] This invention relates to a system that efficiently processes diverse information in the field, enabling rapid decision-making and its implementation. This system supports strategic actions by combining means with multiple functions.
[0241] The server acquires video, audio, and text data in real time using detection devices installed on-site. The hardware used includes cameras, microphones, and various sensors. Software-wise, an integrated data acquisition application works in conjunction with these devices to efficiently collect data.
[0242] The server performs data cleaning on the acquired data using preprocessing mechanisms. Specifically, it removes background noise from audio data and cuts video data to the required frames. This process utilizes advanced algorithms and AI-assisted technologies.
[0243] As an analysis method, the server uses a generative AI model and employs image recognition and speech recognition technologies to identify the enemy's location and movements. Furthermore, natural language processing is performed on text data to analyze the enemy's intentions and tactics. These analyses require dedicated AI toolkits and libraries.
[0244] In decision-making, the server predicts the enemy's next move by comparing a vast dataset of past data with current information. This prediction process utilizes generative AI models to formulate the optimal strategy for each situation.
[0245] In the execution mechanism, the terminal receives instructions from the server and visually provides the user (friendly force) with an action plan. The terminal is equipped with an intuitive interface that displays enemy location information and predicted movements on a map.
[0246] A concrete example is the process where a server analyzes video data collected from a drone to determine the current location and predicted movement path of enemy vehicles. Based on this data, the server transmits information to terminals, which then provide allied forces with immediate visualization of how to deploy. This supports optimal decision-making based on the situation.
[0247] The following is an example of a prompt message to input into the generative AI model.
[0248] "Input data: Video data, audio data, text data. Objective: Identify enemy location and movement, analyze communication content, and predict next actions."
[0249] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0250] Step 1:
[0251] The server acquires video, audio, and text data from detection devices in real time. Input includes raw data transmitted from cameras and microphones installed on-site. The server uses an interface to acquire and store this data in a database. Output is the prepared raw dataset.
[0252] Step 2:
[0253] The server preprocesses the acquired data. The input is the raw data obtained in step 1. Specifically, it performs noise reduction on the audio data and extracts the necessary frames from the video data. This generates a dataset suitable for analysis. The output is denoised audio data and formatted video data.
[0254] Step 3:
[0255] The server inputs the pre-processed data into the generating AI model for analysis. The input is the data processed in step 2. The server applies an image recognition algorithm to identify the enemy's movement and location in the video, and performs speech recognition on the audio data to convert the information into text. It also uses natural language processing to understand the intent. The output provides analysis results regarding the enemy's location, movement, and intent.
[0256] Step 4:
[0257] The server uses the analysis results to execute the decision-making process. The input is the analysis results obtained in step 3. The server compares this with past database data and uses a generated AI model to predict the enemy's next move. The output generates an optimal action plan or prediction.
[0258] Step 5:
[0259] The server communicates the execution plan to the terminal. The input is the action plan formulated in step 4. The server sends information about the action plan and predicted enemy movements to the terminal. The output is the instruction information displayed on the terminal.
[0260] Step 6:
[0261] The terminal transmits information received from the server to the user. The input is the instruction information received in step 5. The terminal visualizes this information on a map and provides an interface to prompt the user to take quick action. As output, information for the user to take action is visualized.
[0262] (Application Example 1)
[0263] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0264] Security monitoring in public and commercial facilities currently relies on traditional methods such as security cameras and visual checks by security guards, which can lead to delays in response. Therefore, there is a need for means to quickly detect potential dangers and ensure safety.
[0265] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0266] In this invention, the server includes information gathering means for collecting multiple pieces of information, including video, audio, and text, in real time; analysis means for analyzing and integrating the collected information; and safety monitoring means for analyzing information related to safety and identifying hazards. This enables real-time monitoring of safety conditions and rapid response.
[0267] "Information gathering means" refers to a device or method for collecting multiple pieces of information, such as images, sounds, and text, in real time.
[0268] "Analysis means" refers to a process or technique for processing and integrating collected information to understand a situation.
[0269] A "decision-making tool" is a system that serves as the foundation for determining the optimal course of action based on analysis results.
[0270] An "action execution mechanism" is a mechanism for actually carrying out a decided action and transmitting instructions.
[0271] A "control system" is a system that integrates and integrates various means of information gathering, analysis, decision-making, and action execution to function as a single unit.
[0272] "Safety monitoring measures" refer to methods and techniques for monitoring safety within a facility and identifying potential hazards.
[0273] This system is designed for safety monitoring in public and commercial facilities, and functions through the collaboration of a server, terminals, and users. The server first acquires data from various video and acoustic sensors. This is supported by an information gathering system that has the function of aggregating video, audio, and text data in real time.
[0274] Next, the server processes the collected information using analysis tools. In this step, the video data is analyzed using image processing libraries such as OpenCV, and object detection is performed using the Hugging Face Transformers library. Based on the collected information, AI technology is used to identify potential risks.
[0275] Based on the analysis results, the server's decision-making system evaluates and determines the optimal security response. As a result, if a threat is identified, it can immediately send a warning or issue instructions to security staff through the action execution system. This enables a rapid response and the implementation of appropriate security measures.
[0276] For example, a system can continuously monitor camera feeds within a shopping mall, detect suspicious behavior or dangers in real time, and alert the security team. Such a system would also be useful in crowded conditions during holidays or events.
[0277] The generative AI model can be input using prompts like the following:
[0278] "Analyze real-time camera feeds from within the shopping mall to identify suspicious individuals and congestion levels."
[0279] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0280] Step 1:
[0281] The server acquires real-time data from video and acoustic sensors within the facility. This input data is collected by the server's information gathering mechanism, and image data is sent to the server as frames, while acoustic data is sent as waveforms. This allows for the centralized aggregation of diverse information, preparing it for subsequent analysis.
[0282] Step 2:
[0283] The server processes the collected data using analysis means. For video data, it analyzes each frame using an image processing library such as OpenCV and identifies suspicious persons or abnormal behaviors using an object detection model. For acoustic data, it identifies abnormal sound patterns using a speech recognition algorithm. This output is the analysis result including the match with suspicious behaviors and abnormal sounds.
[0284] Step 3:
[0285] The server selects an optimal countermeasure by the decision-making means based on the analysis result. In this step, using a generative AI model, based on the existing data and the current situation, it determines the alert level and selects the necessary countermeasures. This input is the analysis result, and the output is the selection of the countermeasure.
[0286] Step 4:
[0287] The terminal transmits the selected countermeasure as a specific instruction to the security staff through the action execution means. This instruction is displayed on the terminal both audibly and visually, and the security staff can quickly execute the on-site response according to this. The input is the countermeasure from the decision-making means, and the output is the specific instruction.
[0288] Step 5:
[0289] The security staff, who is the user, implements the actual security actions on-site based on the instructions from the terminal. Based on the specific instructions from the terminal, the security staff can take flexible measures according to the on-site situation. Here, it receives the instruction from the terminal as the input, and the output is the execution action for ensuring security.
[0290] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform specific processing using the user's emotion.
[0291] In addition to the conventional functions of collecting and analyzing diverse data on the battlefield, the system of the present invention also has the function of recognizing the user's emotions. This system incorporates an emotion engine in addition to information gathering means, analysis means, decision-making means, action execution means, and control means, thereby enabling the optimization of tactics that take into account the user's emotional state.
[0292] Information gathering methods include not only data from various sensors, cameras, and microphones on the battlefield, but also emotional data collected through user biometric signals and voice surveys. This allows the server to centrally acquire a wide range of information, forming a foundation for analysis.
[0293] The analysis method processes the collected data using AI technology. The server analyzes video data using image recognition, communication sounds using speech recognition, and text using natural language processing. In addition, it recognizes the user's emotional state by analyzing the intonation of their voice and biometric data using an emotion engine.
[0294] The decision-making process involves using a generated AI to assess the situation based on integrated analysis results and determine the optimal course of action. In particular, it detects the stress and distress that users encounter on the battlefield based on emotional data and proposes options and support methods to mitigate them.
[0295] The action execution mechanism sends instructions to the terminal to carry out tactical actions selected by the server. The terminal provides the user with specific instructions in real time and visualizes information that is adjusted to minimize user stress.
[0296] For example, if a server analyzes a user's heart rate and voice patterns using an emotion engine and determines that the user is in a high-stress state, the decision-making system has the ability to suggest low-risk options. Depending on the device, the user can receive this information in an easily accessible format, allowing them to choose the optimal course of action while reducing psychological burden.
[0297] In this way, the present invention can provide a system that not only accelerates information processing and decision-making on the battlefield but also enables flexible responses that take into account the psychological state of the human being.
[0298] The following describes the processing flow.
[0299] Step 1:
[0300] The server acquires video, audio, and text data from battlefield sensors, cameras, and microphones, while simultaneously collecting biometric signals such as heart rate and voice intonation from users' mobile devices and wearables. This also allows for the acquisition of user emotional data.
[0301] Step 2:
[0302] The server preprocesses the collected data. Video data is de-noised and converted into an easily analyzable format. Audio data is converted to text using speech recognition technology. For emotional data, an emotion engine is used to numerically represent the user's current emotional state (stress, joy, anxiety, etc.).
[0303] Step 3:
[0304] The server analyzes data using various AI models. Image recognition models identify the locations of enemies and obstacles, important keywords are extracted from audio data, and intentions and commands are analyzed from text data using natural language processing. The emotion engine evaluates how the user's emotions affect the overall situation.
[0305] Step 4:
[0306] The server integrates the analyzed data and builds a comprehensive situational model. Using GIS, it maps the positions of friendly and enemy forces along with the users' emotional states to grasp the overall tactical situation. Based on this situational model, the generative AI performs tactical predictions and simulations.
[0307] Step 5:
[0308] The server formulates an optimal action plan using generative AI. In this process, it takes into account the user's emotional state and selects tactics to reduce risks or methods to provide mental support to the user in high-stress situations. The plan is prepared as an action plan.
[0309] Step 6:
[0310] The terminal receives the action plan transmitted from the server and provides instructions to the user in real time. The instructions also include a map visually showing the situation and proposals for mental health support as relaxation measures. With this support, the user can make appropriate judgments and take actions.
[0311] (Example 2)
[0312] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0313] Existing tactical support systems have limitations in real-time situation awareness and decision-making, and optimization considering the psychological and physiological states of users on the battlefield has not been fully achieved. In addition, the instructions for users may be burdensome, which is a factor hindering quick and reliable decision-making. Therefore, more flexible and advanced tactical support considering the psychological and physiological states is required.
[0314] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0315] In this invention, the server includes information gathering means for collecting multiple pieces of information, including video, audio, text, and biosignals, in real time; analysis means for analyzing the collected information and generating integrated analysis results, including the user's emotional state; and decision-making means for recognizing the situation based on the integrated analysis results and determining the optimal action using a generated AI model. This enables the selection of rapid and appropriate tactical actions while taking the user's emotional state into consideration.
[0316] "Information gathering means" refers to means for collecting multiple types of information in real time, including video, audio, text, and biosignals, using measuring instruments deployed in a tactical area.
[0317] "Analysis means" refers to the means of analyzing collected information and generating integrated analysis results, including the user's emotional state.
[0318] A "decision-making tool" is a means of recognizing a situation based on integrated analysis results and determining the optimal action using a generative AI model.
[0319] "Action execution means" refers to the means for carrying out a determined action and providing the user with predetermined instructions.
[0320] "Control means" refers to means for coordinating information gathering means, analysis means, decision-making means, and action execution means.
[0321] A "generative AI model" is an artificial intelligence model used to determine the optimal course of action by utilizing integrated analysis results.
[0322] A "prompt statement" is a sentence of text containing instructions or information that is input to a generative AI model.
[0323] The embodiments for carrying out the present invention are shown below.
[0324] This system supports information gathering, analysis, decision-making, and action execution in tactical areas. The server uses information gathering tools to collect diverse real-time data from sensors, cameras, and microphones. This data includes images, audio, text, and biometric signals, as well as information about the tactical area's situation and the user's emotional state.
[0325] The server processes the collected data using analytical tools. This analysis utilizes image recognition software, speech recognition software, and a natural language processing engine. Specifically, it recognizes image data as visual information and converts speech data into text. It also uses an emotion engine to analyze the user's emotions from biometric data and voice input.
[0326] Furthermore, it incorporates a decision-making mechanism using a generative AI model, which makes situational judgments and determines the optimal action based on integrated analysis results. This allows for the formation of safe and effective tactical actions that take into account the user's psychological state. As a concrete example of a prompt, the generative AI model is input with instructions such as, "Recommend the optimal action when the user is in a high-stress state."
[0327] Upon receiving instructions from the server, the terminal uses its execution mechanisms to provide intuitive and easy-to-understand instructions to the user. The terminal communicates information to the user through visual displays and voice, supporting real-time decision-making.
[0328] In this way, the system of the present invention provides advanced information processing and situational response capabilities in the tactical zone. Furthermore, it can reduce the psychological burden on users and enable efficient decision-making. This system realizes cutting-edge tactical support utilizing information technology and artificial intelligence.
[0329] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0330] Step 1:
[0331] The server acquires data from sensors, cameras, and microphones in the tactical area using information gathering tools. Inputs include a variety of data such as images, audio, and biosignals. This data is collected in real time and temporarily stored.
[0332] Step 2:
[0333] The server preprocesses the acquired data into a format that can be analyzed. Specifically, it adjusts image resolution and removes noise, filters audio data, and tokenizes text data. The preprocessed data is then used as input for the next analysis step.
[0334] Step 3:
[0335] The server analyzes the pre-processed data using analytical tools. In this step, visual information is analyzed using image recognition software, and audio data is converted to text using speech recognition software. Furthermore, an emotion engine is used to determine the user's emotional state from biometric data. The analysis results are output as integrated information.
[0336] Step 4:
[0337] The server uses a generative AI model to make decisions based on the integrated analysis results. At this stage, a prompt message such as "Recommend the best course of action when the user is in a high-stress state" is input, and the generative AI proposes the optimal tactical action. As a result, optimized action instructions are output.
[0338] Step 5:
[0339] The server sends action instructions to the terminal. The terminal receives this information and issues instructions to the user using the means of execution. Specifically, this includes providing visual information on a visual display and audio guidance. Based on this information, the user can take appropriate action.
[0340] (Application Example 2)
[0341] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0342] In traditional brick-and-mortar stores, accurately understanding customers' emotional states and providing timely service accordingly has been difficult. While the importance of emotion-based responses in improving customer satisfaction and promoting purchases is recognized, there is a lack of concrete methods for putting this into practice.
[0343] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0344] In this invention, the server includes information acquisition means for collecting the customer's emotional state in real time through biosignal data and voice data, analysis means for analyzing the collected emotional data and determining the customer's psychological state, and decision-making means for determining the optimal response policy based on the determined emotional state. This makes it possible to provide optimal service tailored to the customer's emotional state in the store.
[0345] "Information acquisition means" refers to means for collecting customers' emotional states in real time through biosignal data and voice data.
[0346] "Analysis methods" refer to means of analyzing collected emotional data to determine the customer's psychological state.
[0347] A "decision-making tool" is a means of determining the optimal course of action based on the identified emotional state.
[0348] "Means of action" refer to the means of implementing the decided response policy and sending appropriate instructions to service providers.
[0349] "Control means" refers to means that coordinate information acquisition means, analysis means, decision-making means, and action execution means to improve services in the store environment.
[0350] This invention realizes a system that integrates emotion recognition and decision-making processes to improve customer service in physical stores. The server collects biosignal data and voice data acquired from sensors and voice recognition devices in real time and functions as an information acquisition means. This allows data on the emotional state of customers to be obtained.
[0351] The server analyzes the collected data and uses analytical methods to determine the customer's psychological state. This analysis employs artificial intelligence technology, specifically machine learning libraries such as TensorFlow and PyTorch. Furthermore, Google Cloud Speech-to-Text may be used for speech recognition. This allows for accurate detection of the customer's emotional state.
[0352] Based on the analysis results, a decision-making mechanism is activated to determine the optimal customer service strategy. This decision-making process utilizes a generative AI model to propose appropriate service methods tailored to the customer's emotional state.
[0353] For example, if a customer appears anxious, the server might send an instruction to the service provider such as, "Please explain the current promotion again to alleviate the customer's anxiety."
[0354] Another example of a prompt message is, "If the customer is feeling anxious, please suggest the best course of action to take." This allows store staff to provide service that responds to the customer's emotions in real time, thereby improving customer satisfaction.
[0355] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0356] Step 1:
[0357] The server uses sensors and voice recognition devices to collect customer biometric and voice data in real time. The sensors detect data such as pulse rate and facial expressions, while the voice recognition device records the customer's voice. This biometric and voice data is used as input, and the output is raw data ready for analysis.
[0358] Step 2:
[0359] The server processes the collected data using analytical tools. Specifically, it uses machine learning models such as TensorFlow and PyTorch to perform advanced data analysis to determine the customer's emotional state. The input data is the raw data obtained in step 1, and the output is the analysis result indicating the customer's emotional state. This analysis result estimates the customer's psychological state and is used in the next decision-making step.
[0360] Step 3:
[0361] Based on the analysis results, the server uses decision-making tools to determine the optimal course of action. A generative AI model is used to generate prompt messages that match the customer's emotional state. The input is the analysis results from step 2, and the output is a concrete action plan. This action plan contributes to improving customer service and increasing satisfaction.
[0362] Step 4:
[0363] The server transmits the determined action plan to the terminal via the action execution mechanism. The terminal displays instructions to the service provider in the form of a prompt message, such as "Please explain the current promotion again to alleviate the customer's concerns." The input data is the action plan obtained in step 3, and the output is specific instructions for the service provider. This enables the user to respond in real time on-site.
[0364] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0365] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0366] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0367] [Third Embodiment]
[0368] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0369] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0370] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0371] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0372] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0373] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0374] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0375] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0376] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0377] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0378] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0379] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0380] The system of the present invention aims to process information quickly and accurately on the battlefield and execute optimal actions. This system comprises information gathering means, analysis means, decision-making means, action execution means, and control means for coordinating and coordinating each of these means. Embodiments for carrying out the present invention are described below.
[0381] Information gathering methods acquire data from various sensors, cameras, and microphones deployed on the battlefield. This allows the server to collect diverse information, including video, audio, and text data, in real time.
[0382] The analysis method processes the collected data using specialized AI technology. The server uses image recognition algorithms to identify the enemy's location and movement from the video data, and speech recognition technology to convert the enemy's communications into text data. Natural language processing (NLP) is applied to the text data to analyze the enemy's intentions and situation.
[0383] The decision-making mechanism uses generative AI to assess the situation based on the analysis results and select the optimal action. The server compares past data with the current situation and makes tactical predictions to accurately predict the enemy's next move.
[0384] The execution mechanism takes specific actions based on instructions from the server. The terminal provides instructions to friendly units in real time and visualizes the information needed to take necessary actions on the terminal. This allows users (friendly units) to execute instructions quickly and efficiently.
[0385] As a concrete example, suppose a server analyzes video data collected from a drone to understand the movement patterns of enemy vehicles in real time. The server then predicts the enemy's escape routes and transmits this information to friendly terminals. Based on the received data, the terminals can deploy friendly forces to the optimal positions, making it possible to more efficiently thwart enemy actions.
[0386] As described above, the present invention provides a system that can accelerate information processing and decision-making on the battlefield and secure a strategic advantage.
[0387] The following describes the processing flow.
[0388] Step 1:
[0389] The server collects video, audio, and text data in real time from sensors and drones deployed on the battlefield. This includes streaming video from fixed cameras and recordings of ambient and communication sounds. The server centrally acquires this data and prepares it for processing.
[0390] Step 2:
[0391] The server preprocesses the collected data. Video data is divided frame by frame, and image filters are applied to remove unnecessary noise. Audio data is converted into a clear audio signal using noise cancellation technology and then into a format suitable for speech recognition. Text data is formatted into an appropriate format and prepared for analysis.
[0392] Step 3:
[0393] The server analyzes data using various AI models. Specifically, it uses image recognition algorithms to automatically detect enemy vehicles and people from video footage and map their locations. It uses speech recognition models to transcribe enemy communications into text and analyze their content. The text data is then processed using natural language processing to understand the intent behind commands and reports.
[0394] Step 4:
[0395] The server integrates multiple analysis results to model the overall situation. It utilizes a Geographic Information System (GIS) to map the positions and movement patterns of both friendly and enemy forces. It also conducts simulations to predict enemy behavior patterns, taking into account temporal changes.
[0396] Step 5:
[0397] The server uses generative AI to determine the optimal action plan based on comprehensive data. It simulates multiple tactical options and selects the most effective course of action. It evaluates risks and benefits and completes the decision-making process.
[0398] Step 6:
[0399] The terminal transmits information to friendly commanders and soldiers in real time based on tactical instructions received from the server. It visualizes the instructions by updating map information and the positions of friendly and enemy forces on the display device. It also provides specific action instructions via voice or message as needed to encourage a rapid response.
[0400] (Example 1)
[0401] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0402] On-site information processing requires handling a wide variety of data formats simultaneously, integrating and analyzing them in real time, and making rapid decisions. However, there are difficulties in establishing appropriate infrastructure for processing diverse data and in rapidly implementing action plans based on analysis results. Furthermore, systems that improve the accuracy of decision-making while communicating instructions in a way that users can immediately execute are also insufficient.
[0403] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0404] In this invention, the server includes means for acquiring multiple pieces of information, including video data, audio data, and text data, in real time via a detection device at the site; preprocessing means for preprocessing the acquired information, reducing noise, and shaping it for analysis; and analysis means for analyzing the preprocessed data, identifying the enemy's movements and location using image recognition and speech recognition technologies, and analyzing their intentions using natural language processing technologies. This enables rapid and accurate processing of diverse data, allowing for immediate action decisions and the transmission of those instructions.
[0405] "Video data" refers to visual information captured by cameras or other visual detection devices.
[0406] "Audio data" refers to acoustic information recorded by microphones or other sound detection devices.
[0407] "Text data" refers to information that has been converted into a string of characters through speech recognition or other methods.
[0408] A "detection device" refers to equipment installed on-site to acquire various types of data, such as video, audio, temperature, and motion.
[0409] "Preprocessing means" refers to the process of removing noise from acquired data and converting it into a format suitable for analysis.
[0410] "Analysis methods" refer to techniques that use pre-processed data and employ machine learning algorithms and AI technologies to examine information in detail.
[0411] "Decision-making tools" refer to the process of making decisions to predict future events by utilizing analyzed information.
[0412] "Action execution means" refers to a method of generating specific instructions based on a decided action plan and transmitting them to a terminal in an executable format.
[0413] "Management means" refers to a system for efficiently controlling the overall process by linking the processes of information acquisition, preprocessing, analysis, decision-making, and execution.
[0414] This invention relates to a system that efficiently processes diverse information in the field, enabling rapid decision-making and its implementation. This system supports strategic actions by combining means with multiple functions.
[0415] The server acquires video, audio, and text data in real time using detection devices installed on-site. The hardware used includes cameras, microphones, and various sensors. Software-wise, an integrated data acquisition application works in conjunction with these devices to efficiently collect data.
[0416] The server performs data cleaning on the acquired data using preprocessing mechanisms. Specifically, it removes background noise from audio data and cuts video data to the required frames. This process utilizes advanced algorithms and AI-assisted technologies.
[0417] As an analysis method, the server uses a generative AI model and employs image recognition and speech recognition technologies to identify the enemy's location and movements. Furthermore, natural language processing is performed on text data to analyze the enemy's intentions and tactics. These analyses require dedicated AI toolkits and libraries.
[0418] In decision-making, the server predicts the enemy's next move by comparing a vast dataset of past data with current information. This prediction process utilizes generative AI models to formulate the optimal strategy for each situation.
[0419] In the execution mechanism, the terminal receives instructions from the server and visually provides the user (friendly force) with an action plan. The terminal is equipped with an intuitive interface that displays enemy location information and predicted movements on a map.
[0420] A concrete example is the process where a server analyzes video data collected from a drone to determine the current location and predicted movement path of enemy vehicles. Based on this data, the server transmits information to terminals, which then provide allied forces with immediate visualization of how to deploy. This supports optimal decision-making based on the situation.
[0421] The following is an example of a prompt message to input into the generative AI model.
[0422] "Input data: Video data, audio data, text data. Objective: Identify enemy location and movement, analyze communication content, and predict next actions."
[0423] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0424] Step 1:
[0425] The server acquires video, audio, and text data from detection devices in real time. Input includes raw data transmitted from cameras and microphones installed on-site. The server uses an interface to acquire and store this data in a database. Output is the prepared raw dataset.
[0426] Step 2:
[0427] The server preprocesses the acquired data. The input is the raw data obtained in step 1. Specifically, it performs noise reduction on the audio data and extracts the necessary frames from the video data. This generates a dataset suitable for analysis. The output is denoised audio data and formatted video data.
[0428] Step 3:
[0429] The server inputs the pre-processed data into the generating AI model for analysis. The input is the data processed in step 2. The server applies an image recognition algorithm to identify the enemy's movement and location in the video, and performs speech recognition on the audio data to convert the information into text. It also uses natural language processing to understand the intent. The output provides analysis results regarding the enemy's location, movement, and intent.
[0430] Step 4:
[0431] The server uses the analysis results to execute the decision-making process. The input is the analysis results obtained in step 3. The server compares this with past database data and uses a generated AI model to predict the enemy's next move. The output generates an optimal action plan or prediction.
[0432] Step 5:
[0433] The server communicates the execution plan to the terminal. The input is the action plan formulated in step 4. The server sends information about the action plan and predicted enemy movements to the terminal. The output is the instruction information displayed on the terminal.
[0434] Step 6:
[0435] The terminal transmits information received from the server to the user. The input is the instruction information received in step 5. The terminal visualizes this information on a map and provides an interface to prompt the user to take quick action. As output, information for the user to take action is visualized.
[0436] (Application Example 1)
[0437] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0438] Security monitoring in public and commercial facilities currently relies on traditional methods such as security cameras and visual checks by security guards, which can lead to delays in response. Therefore, there is a need for means to quickly detect potential dangers and ensure safety.
[0439] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0440] In this invention, the server includes information gathering means for collecting multiple pieces of information, including video, audio, and text, in real time; analysis means for analyzing and integrating the collected information; and safety monitoring means for analyzing information related to safety and identifying hazards. This enables real-time monitoring of safety conditions and rapid response.
[0441] "Information gathering means" refers to a device or method for collecting multiple pieces of information, such as images, sounds, and text, in real time.
[0442] "Analysis means" refers to a process or technique for processing and integrating collected information to understand a situation.
[0443] A "decision-making tool" is a system that serves as the foundation for determining the optimal course of action based on analysis results.
[0444] An "action execution mechanism" is a mechanism for actually carrying out a decided action and transmitting instructions.
[0445] A "control system" is a system that integrates and integrates various means of information gathering, analysis, decision-making, and action execution to function as a single unit.
[0446] "Safety monitoring measures" refer to methods and techniques for monitoring safety within a facility and identifying potential hazards.
[0447] This system is designed for safety monitoring in public and commercial facilities, and functions through the collaboration of a server, terminals, and users. The server first acquires data from various video and acoustic sensors. This is supported by an information gathering system that has the function of aggregating video, audio, and text data in real time.
[0448] Next, the server processes the collected information using analysis tools. In this step, the video data is analyzed using image processing libraries such as OpenCV, and object detection is performed using the Hugging Face Transformers library. Based on the collected information, AI technology is used to identify potential risks.
[0449] Based on the analysis results, the server's decision-making system evaluates and determines the optimal security response. As a result, if a threat is identified, it can immediately send a warning or issue instructions to security staff through the action execution system. This enables a rapid response and the implementation of appropriate security measures.
[0450] For example, a system can continuously monitor camera feeds within a shopping mall, detect suspicious behavior or dangers in real time, and alert the security team. Such a system would also be useful in crowded conditions during holidays or events.
[0451] The generative AI model can be input using prompts like the following:
[0452] "Analyze real-time camera feeds from within the shopping mall to identify suspicious individuals and congestion levels."
[0453] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0454] Step 1:
[0455] The server acquires real-time data from video and acoustic sensors within the facility. This input data is collected by the server's information gathering mechanism, and image data is sent to the server as frames, while acoustic data is sent as waveforms. This allows for the centralized aggregation of diverse information, preparing it for subsequent analysis.
[0456] Step 2:
[0457] The server processes the collected data using analytical tools. For video data, it analyzes each frame using image processing libraries such as OpenCV, and identifies suspicious individuals or abnormal behavior using an object detection model. For audio data, it identifies abnormal sound patterns using a speech recognition algorithm. The output is the analysis result, including matches with suspicious behavior and abnormal sounds.
[0458] Step 3:
[0459] The server selects the optimal countermeasure using a decision-making mechanism based on the analysis results. In this step, a generative AI model is used to determine the alert level based on existing data and the current situation, and to select the necessary response. The input is the analysis results, and the output is the selection of the countermeasure.
[0460] Step 4:
[0461] The terminal transmits the selected response measures as specific instructions to security staff through the action execution mechanism. These instructions are displayed on the terminal both audibly and visually, allowing security staff to quickly implement on-site responses. The input is the response measures from the decision-making mechanism, and the output is the specific instructions.
[0462] Step 5:
[0463] The security staff, as users, carry out actual security-enhancing actions on-site based on instructions from the terminal. Based on the specific instructions from the terminal, the security staff can take flexible measures according to the situation on site. Here, the input is instructions received from the terminal, and the output is the execution action for ensuring security.
[0464] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0465] In addition to the conventional functions of collecting and analyzing diverse data on the battlefield, the system of the present invention also has the function of recognizing the user's emotions. This system incorporates an emotion engine in addition to information gathering means, analysis means, decision-making means, action execution means, and control means, thereby enabling the optimization of tactics that take into account the user's emotional state.
[0466] Information gathering methods include not only data from various sensors, cameras, and microphones on the battlefield, but also emotional data collected through user biometric signals and voice surveys. This allows the server to centrally acquire a wide range of information, forming a foundation for analysis.
[0467] The analysis method processes the collected data using AI technology. The server analyzes video data using image recognition, communication sounds using speech recognition, and text using natural language processing. In addition, it recognizes the user's emotional state by analyzing the intonation of their voice and biometric data using an emotion engine.
[0468] The decision-making process involves using a generated AI to assess the situation based on integrated analysis results and determine the optimal course of action. In particular, it detects the stress and distress that users encounter on the battlefield based on emotional data and proposes options and support methods to mitigate them.
[0469] The action execution mechanism sends instructions to the terminal to carry out tactical actions selected by the server. The terminal provides the user with specific instructions in real time and visualizes information that is adjusted to minimize user stress.
[0470] For example, if a server analyzes a user's heart rate and voice patterns using an emotion engine and determines that the user is in a high-stress state, the decision-making system has the ability to suggest low-risk options. Depending on the device, the user can receive this information in an easily accessible format, allowing them to choose the optimal course of action while reducing psychological burden.
[0471] In this way, the present invention can provide a system that not only accelerates information processing and decision-making on the battlefield but also enables flexible responses that take into account the psychological state of the human being.
[0472] The following describes the processing flow.
[0473] Step 1:
[0474] The server acquires video, audio, and text data from battlefield sensors, cameras, and microphones, while simultaneously collecting biometric signals such as heart rate and voice intonation from users' mobile devices and wearables. This also allows for the acquisition of user emotional data.
[0475] Step 2:
[0476] The server preprocesses the collected data. Video data is de-noised and converted into an easily analyzable format. Audio data is converted to text using speech recognition technology. For emotional data, an emotion engine is used to numerically represent the user's current emotional state (stress, joy, anxiety, etc.).
[0477] Step 3:
[0478] The server analyzes data using various AI models. Image recognition models identify the locations of enemies and obstacles, important keywords are extracted from audio data, and intentions and commands are analyzed from text data using natural language processing. The emotion engine evaluates how the user's emotions affect the overall situation.
[0479] Step 4:
[0480] The server integrates the analyzed data and builds a comprehensive situational model. Using GIS, it maps the positions of friendly and enemy forces along with the users' emotional states to grasp the overall tactical situation. Based on this situational model, the generative AI performs tactical predictions and simulations.
[0481] Step 5:
[0482] The server uses generative AI to develop the optimal action plan. In doing so, it takes the user's emotional state into consideration, selecting tactics to reduce risk and methods to provide emotional support to the user in high-stress situations. The plan is then prepared as an action plan.
[0483] Step 6:
[0484] The terminal receives action plans sent from the server and provides instructions to the user in real time. These instructions include maps that visually represent the situation and suggestions for mental health support as mitigation measures. With this support, the user can make appropriate decisions and take action.
[0485] (Example 2)
[0486] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0487] Existing tactical support systems have limitations in real-time situational awareness and decision-making, and have not adequately optimized to consider the psychological and physiological state of users on the battlefield. Furthermore, instructions given to users can be burdensome, hindering rapid and reliable decision-making. Therefore, there is a need for more flexible and advanced tactical support that takes psychological and physiological states into account.
[0488] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0489] In this invention, the server includes information gathering means for collecting multiple pieces of information, including video, audio, text, and biosignals, in real time; analysis means for analyzing the collected information and generating integrated analysis results, including the user's emotional state; and decision-making means for recognizing the situation based on the integrated analysis results and determining the optimal action using a generated AI model. This enables the selection of rapid and appropriate tactical actions while taking the user's emotional state into consideration.
[0490] "Information gathering means" refers to means for collecting multiple types of information in real time, including video, audio, text, and biosignals, using measuring instruments deployed in a tactical area.
[0491] "Analysis means" refers to the means of analyzing collected information and generating integrated analysis results, including the user's emotional state.
[0492] A "decision-making tool" is a means of recognizing a situation based on integrated analysis results and determining the optimal action using a generative AI model.
[0493] "Action execution means" refers to the means for carrying out a determined action and providing the user with predetermined instructions.
[0494] "Control means" refers to means for coordinating information gathering means, analysis means, decision-making means, and action execution means.
[0495] A "generative AI model" is an artificial intelligence model used to determine the optimal course of action by utilizing integrated analysis results.
[0496] A "prompt statement" is a sentence of text containing instructions or information that is input to a generative AI model.
[0497] The embodiments for carrying out the present invention are shown below.
[0498] This system supports information gathering, analysis, decision-making, and action execution in tactical areas. The server uses information gathering tools to collect diverse real-time data from sensors, cameras, and microphones. This data includes images, audio, text, and biometric signals, as well as information about the tactical area's situation and the user's emotional state.
[0499] The server processes the collected data using analytical tools. This analysis utilizes image recognition software, speech recognition software, and a natural language processing engine. Specifically, it recognizes image data as visual information and converts speech data into text. It also uses an emotion engine to analyze the user's emotions from biometric data and voice input.
[0500] Furthermore, it incorporates a decision-making mechanism using a generative AI model, which makes situational judgments and determines the optimal action based on integrated analysis results. This allows for the formation of safe and effective tactical actions that take into account the user's psychological state. As a concrete example of a prompt, the generative AI model is input with instructions such as, "Recommend the optimal action when the user is in a high-stress state."
[0501] Upon receiving instructions from the server, the terminal uses its execution mechanisms to provide intuitive and easy-to-understand instructions to the user. The terminal communicates information to the user through visual displays and voice, supporting real-time decision-making.
[0502] In this way, the system of the present invention provides advanced information processing and situational response capabilities in the tactical zone. Furthermore, it can reduce the psychological burden on users and enable efficient decision-making. This system realizes cutting-edge tactical support utilizing information technology and artificial intelligence.
[0503] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0504] Step 1:
[0505] The server acquires data from sensors, cameras, and microphones in the tactical area using information gathering tools. Inputs include a variety of data such as images, audio, and biosignals. This data is collected in real time and temporarily stored.
[0506] Step 2:
[0507] The server preprocesses the acquired data into a format that can be analyzed. Specifically, it adjusts image resolution and removes noise, filters audio data, and tokenizes text data. The preprocessed data is then used as input for the next analysis step.
[0508] Step 3:
[0509] The server analyzes the pre-processed data using analytical tools. In this step, visual information is analyzed using image recognition software, and audio data is converted to text using speech recognition software. Furthermore, an emotion engine is used to determine the user's emotional state from biometric data. The analysis results are output as integrated information.
[0510] Step 4:
[0511] The server uses a generative AI model to make decisions based on the integrated analysis results. At this stage, a prompt message such as "Recommend the best course of action when the user is in a high-stress state" is input, and the generative AI proposes the optimal tactical action. As a result, optimized action instructions are output.
[0512] Step 5:
[0513] The server sends action instructions to the terminal. The terminal receives this information and issues instructions to the user using the means of execution. Specifically, this includes providing visual information on a visual display and audio guidance. Based on this information, the user can take appropriate action.
[0514] (Application Example 2)
[0515] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0516] In traditional brick-and-mortar stores, accurately understanding customers' emotional states and providing timely service accordingly has been difficult. While the importance of emotion-based responses in improving customer satisfaction and promoting purchases is recognized, there is a lack of concrete methods for putting this into practice.
[0517] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0518] In this invention, the server includes information acquisition means for collecting the customer's emotional state in real time through biosignal data and voice data, analysis means for analyzing the collected emotional data and determining the customer's psychological state, and decision-making means for determining the optimal response policy based on the determined emotional state. This makes it possible to provide optimal service tailored to the customer's emotional state in the store.
[0519] "Information acquisition means" refers to means for collecting customers' emotional states in real time through biosignal data and voice data.
[0520] "Analysis methods" refer to means of analyzing collected emotional data to determine the customer's psychological state.
[0521] A "decision-making tool" is a means of determining the optimal course of action based on the identified emotional state.
[0522] "Means of action" refer to the means of implementing the decided response policy and sending appropriate instructions to service providers.
[0523] "Control means" refers to means that coordinate information acquisition means, analysis means, decision-making means, and action execution means to improve services in the store environment.
[0524] This invention realizes a system that integrates emotion recognition and decision-making processes to improve customer service in physical stores. The server collects biosignal data and voice data acquired from sensors and voice recognition devices in real time and functions as an information acquisition means. This allows data on the emotional state of customers to be obtained.
[0525] The server analyzes the collected data and uses analytical methods to determine the customer's psychological state. This analysis employs artificial intelligence technology, specifically machine learning libraries such as TensorFlow and PyTorch. Furthermore, Google Cloud Speech-to-Text may be used for speech recognition. This allows for accurate detection of the customer's emotional state.
[0526] Based on the analysis results, a decision-making mechanism is activated to determine the optimal customer service strategy. This decision-making process utilizes a generative AI model to propose appropriate service methods tailored to the customer's emotional state.
[0527] For example, if a customer appears anxious, the server might send an instruction to the service provider such as, "Please explain the current promotion again to alleviate the customer's anxiety."
[0528] Another example of a prompt message is, "If the customer is feeling anxious, please suggest the best course of action to take." This allows store staff to provide service that responds to the customer's emotions in real time, thereby improving customer satisfaction.
[0529] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0530] Step 1:
[0531] The server uses sensors and voice recognition devices to collect customer biometric and voice data in real time. The sensors detect data such as pulse rate and facial expressions, while the voice recognition device records the customer's voice. This biometric and voice data is used as input, and the output is raw data ready for analysis.
[0532] Step 2:
[0533] The server processes the collected data using analytical tools. Specifically, it uses machine learning models such as TensorFlow and PyTorch to perform advanced data analysis to determine the customer's emotional state. The input data is the raw data obtained in step 1, and the output is the analysis result indicating the customer's emotional state. This analysis result estimates the customer's psychological state and is used in the next decision-making step.
[0534] Step 3:
[0535] Based on the analysis results, the server uses decision-making tools to determine the optimal course of action. A generative AI model is used to generate prompt messages that match the customer's emotional state. The input is the analysis results from step 2, and the output is a concrete action plan. This action plan contributes to improving customer service and increasing satisfaction.
[0536] Step 4:
[0537] The server transmits the determined action plan to the terminal via the action execution mechanism. The terminal displays instructions to the service provider in the form of a prompt message, such as "Please explain the current promotion again to alleviate the customer's concerns." The input data is the action plan obtained in step 3, and the output is specific instructions for the service provider. This enables the user to respond in real time on-site.
[0538] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0539] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0540] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0541] [Fourth Embodiment]
[0542] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0543] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0544] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0545] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0546] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0547] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0548] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0549] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0550] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0551] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0552] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0553] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0554] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0555] The system of the present invention aims to process information quickly and accurately on the battlefield and execute optimal actions. This system comprises information gathering means, analysis means, decision-making means, action execution means, and control means for coordinating and coordinating each of these means. Embodiments for carrying out the present invention are described below.
[0556] Information gathering methods acquire data from various sensors, cameras, and microphones deployed on the battlefield. This allows the server to collect diverse information, including video, audio, and text data, in real time.
[0557] The analysis method processes the collected data using specialized AI technology. The server uses image recognition algorithms to identify the enemy's location and movement from the video data, and speech recognition technology to convert the enemy's communications into text data. Natural language processing (NLP) is applied to the text data to analyze the enemy's intentions and situation.
[0558] The decision-making mechanism uses generative AI to assess the situation based on the analysis results and select the optimal action. The server compares past data with the current situation and makes tactical predictions to accurately predict the enemy's next move.
[0559] The execution mechanism takes specific actions based on instructions from the server. The terminal provides instructions to friendly units in real time and visualizes the information needed to take necessary actions on the terminal. This allows users (friendly units) to execute instructions quickly and efficiently.
[0560] As a concrete example, suppose a server analyzes video data collected from a drone to understand the movement patterns of enemy vehicles in real time. The server then predicts the enemy's escape routes and transmits this information to friendly terminals. Based on the received data, the terminals can deploy friendly forces to the optimal positions, making it possible to more efficiently thwart enemy actions.
[0561] As described above, the present invention provides a system that can accelerate information processing and decision-making on the battlefield and secure a strategic advantage.
[0562] The following describes the processing flow.
[0563] Step 1:
[0564] The server collects video, audio, and text data in real time from sensors and drones deployed on the battlefield. This includes streaming video from fixed cameras and recordings of ambient and communication sounds. The server centrally acquires this data and prepares it for processing.
[0565] Step 2:
[0566] The server preprocesses the collected data. Video data is divided frame by frame, and image filters are applied to remove unnecessary noise. Audio data is converted into a clear audio signal using noise cancellation technology and then into a format suitable for speech recognition. Text data is formatted into an appropriate format and prepared for analysis.
[0567] Step 3:
[0568] The server analyzes data using various AI models. Specifically, it uses image recognition algorithms to automatically detect enemy vehicles and people from video footage and map their locations. It uses speech recognition models to transcribe enemy communications into text and analyze their content. The text data is then processed using natural language processing to understand the intent behind commands and reports.
[0569] Step 4:
[0570] The server integrates multiple analysis results to model the overall situation. It utilizes a Geographic Information System (GIS) to map the positions and movement patterns of both friendly and enemy forces. It also conducts simulations to predict enemy behavior patterns, taking into account temporal changes.
[0571] Step 5:
[0572] The server uses generative AI to determine the optimal action plan based on comprehensive data. It simulates multiple tactical options and selects the most effective course of action. It evaluates risks and benefits and completes the decision-making process.
[0573] Step 6:
[0574] The terminal transmits information to friendly commanders and soldiers in real time based on tactical instructions received from the server. It visualizes the instructions by updating map information and the positions of friendly and enemy forces on the display device. It also provides specific action instructions via voice or message as needed to encourage a rapid response.
[0575] (Example 1)
[0576] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0577] On-site information processing requires handling a wide variety of data formats simultaneously, integrating and analyzing them in real time, and making rapid decisions. However, there are difficulties in establishing appropriate infrastructure for processing diverse data and in rapidly implementing action plans based on analysis results. Furthermore, systems that improve the accuracy of decision-making while communicating instructions in a way that users can immediately execute are also insufficient.
[0578] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0579] In this invention, the server includes means for acquiring multiple pieces of information, including video data, audio data, and text data, in real time via a detection device at the site; preprocessing means for preprocessing the acquired information, reducing noise, and shaping it for analysis; and analysis means for analyzing the preprocessed data, identifying the enemy's movements and location using image recognition and speech recognition technologies, and analyzing their intentions using natural language processing technologies. This enables rapid and accurate processing of diverse data, allowing for immediate action decisions and the transmission of those instructions.
[0580] "Video data" refers to visual information captured by cameras or other visual detection devices.
[0581] "Audio data" refers to acoustic information recorded by microphones or other sound detection devices.
[0582] "Text data" refers to information that has been converted into a string of characters through speech recognition or other methods.
[0583] A "detection device" refers to equipment installed on-site to acquire various types of data, such as video, audio, temperature, and motion.
[0584] "Preprocessing means" refers to the process of removing noise from acquired data and converting it into a format suitable for analysis.
[0585] "Analysis methods" refer to techniques that use pre-processed data and employ machine learning algorithms and AI technologies to examine information in detail.
[0586] "Decision-making tools" refer to the process of making decisions to predict future events by utilizing analyzed information.
[0587] "Action execution means" refers to a method of generating specific instructions based on a decided action plan and transmitting them to a terminal in an executable format.
[0588] "Management means" refers to a system for efficiently controlling the overall process by linking the processes of information acquisition, preprocessing, analysis, decision-making, and execution.
[0589] This invention relates to a system that efficiently processes diverse information in the field, enabling rapid decision-making and its implementation. This system supports strategic actions by combining means with multiple functions.
[0590] The server acquires video, audio, and text data in real time using detection devices installed on-site. The hardware used includes cameras, microphones, and various sensors. Software-wise, an integrated data acquisition application works in conjunction with these devices to efficiently collect data.
[0591] The server performs data cleaning on the acquired data using preprocessing mechanisms. Specifically, it removes background noise from audio data and cuts video data to the required frames. This process utilizes advanced algorithms and AI-assisted technologies.
[0592] As an analysis method, the server uses a generative AI model and employs image recognition and speech recognition technologies to identify the enemy's location and movements. Furthermore, natural language processing is performed on text data to analyze the enemy's intentions and tactics. These analyses require dedicated AI toolkits and libraries.
[0593] In decision-making, the server predicts the enemy's next move by comparing a vast dataset of past data with current information. This prediction process utilizes generative AI models to formulate the optimal strategy for each situation.
[0594] In the execution mechanism, the terminal receives instructions from the server and visually provides the user (friendly force) with an action plan. The terminal is equipped with an intuitive interface that displays enemy location information and predicted movements on a map.
[0595] A concrete example is the process where a server analyzes video data collected from a drone to determine the current location and predicted movement path of enemy vehicles. Based on this data, the server transmits information to terminals, which then provide allied forces with immediate visualization of how to deploy. This supports optimal decision-making based on the situation.
[0596] The following is an example of a prompt message to input into the generative AI model.
[0597] "Input data: Video data, audio data, text data. Objective: Identify enemy location and movement, analyze communication content, and predict next actions."
[0598] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0599] Step 1:
[0600] The server acquires video, audio, and text data from detection devices in real time. Input includes raw data transmitted from cameras and microphones installed on-site. The server uses an interface to acquire and store this data in a database. Output is the prepared raw dataset.
[0601] Step 2:
[0602] The server preprocesses the acquired data. The input is the raw data obtained in step 1. Specifically, it performs noise reduction on the audio data and extracts the necessary frames from the video data. This generates a dataset suitable for analysis. The output is denoised audio data and formatted video data.
[0603] Step 3:
[0604] The server inputs the pre-processed data into the generating AI model for analysis. The input is the data processed in step 2. The server applies an image recognition algorithm to identify the enemy's movement and location in the video, and performs speech recognition on the audio data to convert the information into text. It also uses natural language processing to understand the intent. The output provides analysis results regarding the enemy's location, movement, and intent.
[0605] Step 4:
[0606] The server uses the analysis results to execute the decision-making process. The input is the analysis results obtained in step 3. The server compares this with past database data and uses a generated AI model to predict the enemy's next move. The output generates an optimal action plan or prediction.
[0607] Step 5:
[0608] The server communicates the execution plan to the terminal. The input is the action plan formulated in step 4. The server sends information about the action plan and predicted enemy movements to the terminal. The output is the instruction information displayed on the terminal.
[0609] Step 6:
[0610] The terminal transmits information received from the server to the user. The input is the instruction information received in step 5. The terminal visualizes this information on a map and provides an interface to prompt the user to take quick action. As output, information for the user to take action is visualized.
[0611] (Application Example 1)
[0612] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0613] Security monitoring in public and commercial facilities currently relies on traditional methods such as security cameras and visual checks by security guards, which can lead to delays in response. Therefore, there is a need for means to quickly detect potential dangers and ensure safety.
[0614] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0615] In this invention, the server includes information gathering means for collecting multiple pieces of information, including video, audio, and text, in real time; analysis means for analyzing and integrating the collected information; and safety monitoring means for analyzing information related to safety and identifying hazards. This enables real-time monitoring of safety conditions and rapid response.
[0616] "Information gathering means" refers to a device or method for collecting multiple pieces of information, such as images, sounds, and text, in real time.
[0617] "Analysis means" refers to a process or technique for processing and integrating collected information to understand a situation.
[0618] A "decision-making tool" is a system that serves as the foundation for determining the optimal course of action based on analysis results.
[0619] An "action execution mechanism" is a mechanism for actually carrying out a decided action and transmitting instructions.
[0620] A "control system" is a system that integrates and integrates various means of information gathering, analysis, decision-making, and action execution to function as a single unit.
[0621] "Safety monitoring measures" refer to methods and techniques for monitoring safety within a facility and identifying potential hazards.
[0622] This system is designed for safety monitoring in public and commercial facilities, and functions through the collaboration of a server, terminals, and users. The server first acquires data from various video and acoustic sensors. This is supported by an information gathering system that has the function of aggregating video, audio, and text data in real time.
[0623] Next, the server processes the collected information using analysis tools. In this step, the video data is analyzed using image processing libraries such as OpenCV, and object detection is performed using the Hugging Face Transformers library. Based on the collected information, AI technology is used to identify potential risks.
[0624] Based on the analysis results, the server's decision-making system evaluates and determines the optimal security response. As a result, if a threat is identified, it can immediately send a warning or issue instructions to security staff through the action execution system. This enables a rapid response and the implementation of appropriate security measures.
[0625] For example, a system can continuously monitor camera feeds within a shopping mall, detect suspicious behavior or dangers in real time, and alert the security team. Such a system would also be useful in crowded conditions during holidays or events.
[0626] The generative AI model can be input using prompts like the following:
[0627] "Analyze real-time camera feeds from within the shopping mall to identify suspicious individuals and congestion levels."
[0628] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0629] Step 1:
[0630] The server acquires real-time data from video and acoustic sensors within the facility. This input data is collected by the server's information gathering mechanism, and image data is sent to the server as frames, while acoustic data is sent as waveforms. This allows for the centralized aggregation of diverse information, preparing it for subsequent analysis.
[0631] Step 2:
[0632] The server processes the collected data using analytical tools. For video data, it analyzes each frame using image processing libraries such as OpenCV, and identifies suspicious individuals or abnormal behavior using an object detection model. For audio data, it identifies abnormal sound patterns using a speech recognition algorithm. The output is the analysis result, including matches with suspicious behavior and abnormal sounds.
[0633] Step 3:
[0634] The server selects the optimal countermeasure using a decision-making mechanism based on the analysis results. In this step, a generative AI model is used to determine the alert level based on existing data and the current situation, and to select the necessary response. The input is the analysis results, and the output is the selection of the countermeasure.
[0635] Step 4:
[0636] The terminal transmits the selected response measures as specific instructions to security staff through the action execution mechanism. These instructions are displayed on the terminal both audibly and visually, allowing security staff to quickly implement on-site responses. The input is the response measures from the decision-making mechanism, and the output is the specific instructions.
[0637] Step 5:
[0638] The security staff, as users, carry out actual security-enhancing actions on-site based on instructions from the terminal. Based on the specific instructions from the terminal, the security staff can take flexible measures according to the situation on site. Here, the input is instructions received from the terminal, and the output is the execution action for ensuring security.
[0639] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0640] In addition to the conventional functions of collecting and analyzing diverse data on the battlefield, the system of the present invention also has the function of recognizing the user's emotions. This system incorporates an emotion engine in addition to information gathering means, analysis means, decision-making means, action execution means, and control means, thereby enabling the optimization of tactics that take into account the user's emotional state.
[0641] Information gathering methods include not only data from various sensors, cameras, and microphones on the battlefield, but also emotional data collected through user biometric signals and voice surveys. This allows the server to centrally acquire a wide range of information, forming a foundation for analysis.
[0642] The analysis method processes the collected data using AI technology. The server analyzes video data using image recognition, communication sounds using speech recognition, and text using natural language processing. In addition, it recognizes the user's emotional state by analyzing the intonation of their voice and biometric data using an emotion engine.
[0643] The decision-making process involves using a generated AI to assess the situation based on integrated analysis results and determine the optimal course of action. In particular, it detects the stress and distress that users encounter on the battlefield based on emotional data and proposes options and support methods to mitigate them.
[0644] The action execution mechanism sends instructions to the terminal to carry out tactical actions selected by the server. The terminal provides the user with specific instructions in real time and visualizes information that is adjusted to minimize user stress.
[0645] For example, if a server analyzes a user's heart rate and voice patterns using an emotion engine and determines that the user is in a high-stress state, the decision-making system has the ability to suggest low-risk options. Depending on the device, the user can receive this information in an easily accessible format, allowing them to choose the optimal course of action while reducing psychological burden.
[0646] In this way, the present invention can provide a system that not only accelerates information processing and decision-making on the battlefield but also enables flexible responses that take into account the psychological state of the human being.
[0647] The following describes the processing flow.
[0648] Step 1:
[0649] The server acquires video, audio, and text data from battlefield sensors, cameras, and microphones, while simultaneously collecting biometric signals such as heart rate and voice intonation from users' mobile devices and wearables. This also allows for the acquisition of user emotional data.
[0650] Step 2:
[0651] The server preprocesses the collected data. Video data is de-noised and converted into an easily analyzable format. Audio data is converted to text using speech recognition technology. For emotional data, an emotion engine is used to numerically represent the user's current emotional state (stress, joy, anxiety, etc.).
[0652] Step 3:
[0653] The server analyzes data using various AI models. Image recognition models identify the locations of enemies and obstacles, important keywords are extracted from audio data, and intentions and commands are analyzed from text data using natural language processing. The emotion engine evaluates how the user's emotions affect the overall situation.
[0654] Step 4:
[0655] The server integrates the analyzed data and builds a comprehensive situational model. Using GIS, it maps the positions of friendly and enemy forces along with the users' emotional states to grasp the overall tactical situation. Based on this situational model, the generative AI performs tactical predictions and simulations.
[0656] Step 5:
[0657] The server uses generative AI to develop the optimal action plan. In doing so, it takes the user's emotional state into consideration, selecting tactics to reduce risk and methods to provide emotional support to the user in high-stress situations. The plan is then prepared as an action plan.
[0658] Step 6:
[0659] The terminal receives action plans sent from the server and provides instructions to the user in real time. These instructions include maps that visually represent the situation and suggestions for mental health support as mitigation measures. With this support, the user can make appropriate decisions and take action.
[0660] (Example 2)
[0661] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0662] Existing tactical support systems have limitations in real-time situational awareness and decision-making, and have not adequately optimized to consider the psychological and physiological state of users on the battlefield. Furthermore, instructions given to users can be burdensome, hindering rapid and reliable decision-making. Therefore, there is a need for more flexible and advanced tactical support that takes psychological and physiological states into account.
[0663] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0664] In this invention, the server includes information gathering means for collecting multiple pieces of information, including video, audio, text, and biosignals, in real time; analysis means for analyzing the collected information and generating integrated analysis results, including the user's emotional state; and decision-making means for recognizing the situation based on the integrated analysis results and determining the optimal action using a generated AI model. This enables the selection of rapid and appropriate tactical actions while taking the user's emotional state into consideration.
[0665] "Information gathering means" refers to means for collecting multiple types of information in real time, including video, audio, text, and biosignals, using measuring instruments deployed in a tactical area.
[0666] "Analysis means" refers to the means of analyzing collected information and generating integrated analysis results, including the user's emotional state.
[0667] A "decision-making tool" is a means of recognizing a situation based on integrated analysis results and determining the optimal action using a generative AI model.
[0668] "Action execution means" refers to the means for carrying out a determined action and providing the user with predetermined instructions.
[0669] "Control means" refers to means for coordinating information gathering means, analysis means, decision-making means, and action execution means.
[0670] A "generative AI model" is an artificial intelligence model used to determine the optimal course of action by utilizing integrated analysis results.
[0671] A "prompt statement" is a sentence of text containing instructions or information that is input to a generative AI model.
[0672] The embodiments for carrying out the present invention are shown below.
[0673] This system supports information gathering, analysis, decision-making, and action execution in tactical areas. The server uses information gathering tools to collect diverse real-time data from sensors, cameras, and microphones. This data includes images, audio, text, and biometric signals, as well as information about the tactical area's situation and the user's emotional state.
[0674] The server processes the collected data using analytical tools. This analysis utilizes image recognition software, speech recognition software, and a natural language processing engine. Specifically, it recognizes image data as visual information and converts speech data into text. It also uses an emotion engine to analyze the user's emotions from biometric data and voice input.
[0675] Furthermore, it incorporates a decision-making mechanism using a generative AI model, which makes situational judgments and determines the optimal action based on integrated analysis results. This allows for the formation of safe and effective tactical actions that take into account the user's psychological state. As a concrete example of a prompt, the generative AI model is input with instructions such as, "Recommend the optimal action when the user is in a high-stress state."
[0676] Upon receiving instructions from the server, the terminal uses its execution mechanisms to provide intuitive and easy-to-understand instructions to the user. The terminal communicates information to the user through visual displays and voice, supporting real-time decision-making.
[0677] In this way, the system of the present invention provides advanced information processing and situational response capabilities in the tactical zone. Furthermore, it can reduce the psychological burden on users and enable efficient decision-making. This system realizes cutting-edge tactical support utilizing information technology and artificial intelligence.
[0678] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0679] Step 1:
[0680] The server acquires data from sensors, cameras, and microphones in the tactical area using information gathering tools. Inputs include a variety of data such as images, audio, and biosignals. This data is collected in real time and temporarily stored.
[0681] Step 2:
[0682] The server preprocesses the acquired data into a format that can be analyzed. Specifically, it adjusts image resolution and removes noise, filters audio data, and tokenizes text data. The preprocessed data is then used as input for the next analysis step.
[0683] Step 3:
[0684] The server analyzes the pre-processed data using analytical tools. In this step, visual information is analyzed using image recognition software, and audio data is converted to text using speech recognition software. Furthermore, an emotion engine is used to determine the user's emotional state from biometric data. The analysis results are output as integrated information.
[0685] Step 4:
[0686] The server uses a generative AI model to make decisions based on the integrated analysis results. At this stage, a prompt message such as "Recommend the best course of action when the user is in a high-stress state" is input, and the generative AI proposes the optimal tactical action. As a result, optimized action instructions are output.
[0687] Step 5:
[0688] The server sends action instructions to the terminal. The terminal receives this information and issues instructions to the user using the means of execution. Specifically, this includes providing visual information on a visual display and audio guidance. Based on this information, the user can take appropriate action.
[0689] (Application Example 2)
[0690] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0691] In traditional brick-and-mortar stores, accurately understanding customers' emotional states and providing timely service accordingly has been difficult. While the importance of emotion-based responses in improving customer satisfaction and promoting purchases is recognized, there is a lack of concrete methods for putting this into practice.
[0692] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0693] In this invention, the server includes information acquisition means for collecting the customer's emotional state in real time through biosignal data and voice data, analysis means for analyzing the collected emotional data and determining the customer's psychological state, and decision-making means for determining the optimal response policy based on the determined emotional state. This makes it possible to provide optimal service tailored to the customer's emotional state in the store.
[0694] "Information acquisition means" refers to means for collecting customers' emotional states in real time through biosignal data and voice data.
[0695] "Analysis methods" refer to means of analyzing collected emotional data to determine the customer's psychological state.
[0696] A "decision-making tool" is a means of determining the optimal course of action based on the identified emotional state.
[0697] "Means of action" refer to the means of implementing the decided response policy and sending appropriate instructions to service providers.
[0698] "Control means" refers to means that coordinate information acquisition means, analysis means, decision-making means, and action execution means to improve services in the store environment.
[0699] This invention realizes a system that integrates emotion recognition and decision-making processes to improve customer service in physical stores. The server collects biosignal data and voice data acquired from sensors and voice recognition devices in real time and functions as an information acquisition means. This allows data on the emotional state of customers to be obtained.
[0700] The server analyzes the collected data and uses analytical methods to determine the customer's psychological state. This analysis employs artificial intelligence technology, specifically machine learning libraries such as TensorFlow and PyTorch. Furthermore, Google Cloud Speech-to-Text may be used for speech recognition. This allows for accurate detection of the customer's emotional state.
[0701] Based on the analysis results, a decision-making mechanism is activated to determine the optimal customer service strategy. This decision-making process utilizes a generative AI model to propose appropriate service methods tailored to the customer's emotional state.
[0702] For example, if a customer appears anxious, the server might send an instruction to the service provider such as, "Please explain the current promotion again to alleviate the customer's anxiety."
[0703] Another example of a prompt message is, "If the customer is feeling anxious, please suggest the best course of action to take." This allows store staff to provide service that responds to the customer's emotions in real time, thereby improving customer satisfaction.
[0704] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0705] Step 1:
[0706] The server uses sensors and voice recognition devices to collect customer biometric and voice data in real time. The sensors detect data such as pulse rate and facial expressions, while the voice recognition device records the customer's voice. This biometric and voice data is used as input, and the output is raw data ready for analysis.
[0707] Step 2:
[0708] The server processes the collected data using analytical tools. Specifically, it uses machine learning models such as TensorFlow and PyTorch to perform advanced data analysis to determine the customer's emotional state. The input data is the raw data obtained in step 1, and the output is the analysis result indicating the customer's emotional state. This analysis result estimates the customer's psychological state and is used in the next decision-making step.
[0709] Step 3:
[0710] Based on the analysis results, the server uses decision-making tools to determine the optimal course of action. A generative AI model is used to generate prompt messages that match the customer's emotional state. The input is the analysis results from step 2, and the output is a concrete action plan. This action plan contributes to improving customer service and increasing satisfaction.
[0711] Step 4:
[0712] The server transmits the determined action plan to the terminal via the action execution mechanism. The terminal displays instructions to the service provider in the form of a prompt message, such as "Please explain the current promotion again to alleviate the customer's concerns." The input data is the action plan obtained in step 3, and the output is specific instructions for the service provider. This enables the user to respond in real time on-site.
[0713] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0714] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0715] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0716] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0717] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0718] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0719] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0720] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0721] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0722] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0723] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0724] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0725] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0726] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0727] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0728] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0729] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0730] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0731] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0732] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0733] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0734] The following is further disclosed regarding the embodiments described above.
[0735] (Claim 1)
[0736] An information gathering method that collects multiple data formats, including video, audio, and text, in real time,
[0737] Analytical means for analyzing and integrating collected data,
[0738] Based on the analysis results, a decision-making tool is provided to recognize the situation and determine the optimal course of action.
[0739] Action execution means for carrying out a determined action and transmitting a predetermined instruction,
[0740] Control means for coordinating information gathering means, analysis means, decision-making means and action execution means,
[0741] A system that includes this.
[0742] (Claim 2)
[0743] The system according to claim 1, wherein the information gathering means acquires data using sensors deployed on the battlefield.
[0744] (Claim 3)
[0745] The system according to claim 1, wherein the control means generates a tactically optimized action plan based on integrated data.
[0746] "Example 1"
[0747] (Claim 1)
[0748] A means for acquiring multiple pieces of information, including video data, audio data, and text data, in real time via a detection device at the site,
[0749] A preprocessing means for preprocessing acquired information to reduce noise and to format it for analysis,
[0750] An analytical means for analyzing pre-processed data, identifying enemy movements and locations using image recognition and speech recognition technologies, and analyzing intentions using natural language processing technologies,
[0751] Based on the analysis results, a means of decision-making for predicting the next action by comparing it with past information and formulating the action decided using generative AI,
[0752] Based on the formulated action plan, a means of executing actions is provided to encourage optimal actions by sending instructions to terminals and displaying information.
[0753] A management means that integrates information acquisition means, preprocessing means, analysis means, decision-making means, and action execution means, and enables their coordination,
[0754] A system that includes this.
[0755] (Claim 2)
[0756] The system according to claim 1, wherein the information acquisition means collects various types of information using a detection device installed on site.
[0757] (Claim 3)
[0758] The system according to claim 1, wherein the management means creates a strategically optimized action plan using the integrated analysis results.
[0759] "Application Example 1"
[0760] (Claim 1)
[0761] An information gathering method that collects multiple types of information, including video, audio, and text, in real time,
[0762] Analytical means for analyzing and integrating the collected information,
[0763] Based on the analysis results, a decision-making tool is provided to recognize the situation and determine the optimal course of action.
[0764] An action execution means for carrying out a determined action and transmitting instructions,
[0765] Control means for coordinating information gathering means, analysis means, decision-making means and action execution means,
[0766] Safety monitoring means for analyzing information related to safety and identifying hazards,
[0767] A system that includes this.
[0768] (Claim 2)
[0769] The system according to claim 1, wherein the information gathering means acquires information using sensors placed in the facility.
[0770] (Claim 3)
[0771] The system according to claim 1, wherein the control means generates an optimized safety action plan based on integrated information.
[0772] "Example 2 of combining an emotion engine"
[0773] (Claim 1)
[0774] Information gathering means for collecting multiple types of information in real time, including video, audio, text, and biosignals,
[0775] An analytical means for analyzing collected information and generating integrated analytical results, including the user's emotional state,
[0776] Based on the integrated analysis results, a decision-making mechanism is provided to recognize the situation and determine the optimal action using a generative AI model.
[0777] Action execution means for carrying out a determined action and providing the user with predetermined instructions,
[0778] Control means for coordinating information gathering means, analysis means, decision-making means and action execution means,
[0779] A system that includes this.
[0780] (Claim 2)
[0781] The system according to claim 1, wherein the information gathering means acquires information using measuring instruments placed in a tactical area.
[0782] (Claim 3)
[0783] The system according to claim 1, wherein the control means generates a tactically optimized action plan that takes into account emotional states based on integrated information.
[0784] "Application example 2 when combining with an emotional engine"
[0785] (Claim 1)
[0786] Information acquisition means for collecting the emotional state of customers in real time through biosignal data and voice data,
[0787] An analytical means for analyzing collected emotional data and determining the customer's psychological state,
[0788] A decision-making tool for determining the optimal response based on the identified emotional state,
[0789] Actions to implement the decided response policy and send appropriate instructions to service providers,
[0790] A control system that links information acquisition means, analysis means, decision-making means, and action execution means to improve service in the store environment,
[0791] A system that includes this.
[0792] (Claim 2)
[0793] The system according to claim 1, wherein the information acquisition means acquires data using sensors and voice recognition devices installed in the store.
[0794] (Claim 3)
[0795] The system according to claim 1, wherein the control means generates customer responses as an optimized service plan based on integrated emotional data. [Explanation of symbols]
[0796] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. An information gathering method that collects multiple data formats, including video, audio, and text, in real time, Analytical means for analyzing and integrating collected data, Based on the analysis results, a decision-making tool is provided to recognize the situation and determine the optimal course of action. Action execution means for carrying out a determined action and transmitting a predetermined instruction, Control means for coordinating information gathering means, analysis means, decision-making means, and action execution means, A system that includes this.
2. The system according to claim 1, wherein the information gathering means acquires data using sensors deployed on the battlefield.
3. The system according to claim 1, wherein the control means generates a tactically optimized action plan based on integrated data.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A