System
The system addresses the complexity of conventional drone control by enabling voice-operated drone control with AI-generated flight paths, enhancing user accessibility and reliability.
Patent Information
- Application Number
- JP2024133422
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2026-02-20
AI Technical Summary
Conventional drone control systems require specialized control devices and pilot licenses, posing a barrier for average users due to advanced expertise and lack of voice control, resulting in low convenience and reliability.
A system that includes voice input, voice recognition, command generation using artificial intelligence, communication, and drone control means, allowing users to operate drones with voice commands without the need for piloting skills or specialized knowledge, ensuring safe and accurate flight through high-speed communication.
Enables easy and efficient drone operation for general users by converting voice commands into flight paths, ensuring real-time, safe, and accurate drone control.
Smart Images

Figure 2026030439000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional drone control systems require users to have specialized control devices and pilot licenses, and require advanced expertise and skills, making them difficult for the average user to use. Furthermore, the lack of voice control places a significant burden on users, resulting in low convenience. [Means for solving the problem]
[0005] The present invention solves the problem by providing a system that includes a voice input means for capturing user voice input, a voice recognition means for converting voice into text data, a command generation means for generating drone flight commands based on the text data using artificial intelligence, a communication means for transmitting the flight commands to a drone control platform, and a drone control means for controlling the drone based on the transmitted flight commands. This allows users to easily operate drones using only voice commands, eliminating the need for a piloting device or pilot's license and greatly improving convenience.
[0006] "Voice input means" refers to a device or function for capturing a user's voice instructions.
[0007] A "voice recognition means" is a device or function for converting captured voice into text data.
[0008] A "command generation means" is a device or function that generates flight commands for a drone using artificial intelligence based on text data obtained by a voice recognition means.
[0009] "Communication means" refers to devices or functions for transmitting the generated flight commands to the drone's control board.
[0010] "Drone control means" refers to a device or function for controlling a drone based on flight commands transmitted by a communication means.
[0011] "Generative AI" is an artificial intelligence technology that automatically generates drone flight commands based on user instructions.
[0012] A "flight command" is a series of instructions that instruct a drone to perform specific actions or operations. [Brief explanation of the drawings]
[0013] [Figure 1]1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] To implement the present invention, the following system is required: This system is designed to allow a user to freely control a drone through voice input.
[0035] First, the user uses voice input to give the drone instructions such as "Go to City Hall." This voice input is expected to be implemented on a smartphone or dedicated device.
[0036] Next, the device captures the user's voice and converts the voice data into text data using a speech recognition means. At this stage, a speech recognition engine is used to convert the voice into text, such as "Go to City Hall."
[0037] The converted text data is sent from the device to a server and processed by a command generation means implemented on the server. Specifically, the generative artificial intelligence (generative AI) on the server analyzes the text data and understands the user's instructions. For example, the instruction "Go to city hall" is analyzed, and the location information of city hall is obtained.
[0038] Next, the AI generates flight commands that include a specific flight path to City Hall. These flight commands specify detailed actions from takeoff to the destination, such as "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0039] The generated flight commands are sent to the drone's control platform via the server's communications means using a high-speed communication protocol, ensuring real-time communication.
[0040] Finally, based on the flight commands received by the drone's control board, the drone will fly along the designated path to City Hall, automatically performing a series of operations including takeoff, heading, altitude adjustment, hovering, and landing.
[0041] This allows users to easily control a drone to its destination using only voice input. This system does not require the piloting skills or specialized knowledge that was previously required, and is easy for even general users to use.
[0042] For example, if a user says, "To the supermarket parking lot," the voice input is captured and converted into text using a speech recognition system. The AI then obtains the supermarket's location information, generates a specific flight path, and sends it to the drone, which then automatically flies to the supermarket.
[0043] The processing flow will be explained below.
[0044] Step 1:
[0045] The user uses voice input to give instructions to the drone, such as "Go to City Hall."
[0046] Step 2:
[0047] The device captures the user's voice with a microphone.
[0048] Step 3:
[0049] The device sends the captured voice data to a speech recognition engine, which analyzes the voice data and converts it into corresponding text data ("Go to City Hall").
[0050] Step 4:
[0051] The terminal transmits the converted text data to the server.
[0052] Step 5:
[0053] The server passes the received text data to the generation AI, which analyzes the text data and understands the user's instructions.
[0054] Step 6:
[0055] The server uses a generation AI to obtain the location information of the destination (in this case, city hall) based on the user's instructions.
[0056] Step 7:
[0057] Based on the location information of the destination acquired by the server, the generation AI generates a specific flight path.
[0058] Example: "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0059] Step 8:
[0060] The server transmits the generated flight commands to the drone's control board using a communication means.
[0061] Step 9:
[0062] The drone's control board analyzes the flight commands it receives and converts them into specific operational instructions.
[0063] Step 10:
[0064] The drone's control board controls the drone's motors and propellers, allowing the drone to operate according to flight commands.
[0065] Step 11:
[0066] The drone's control board executes a series of operations from takeoff to the destination, finally reaching and landing at the specified location.
[0067] This allows users to easily control the drone using only voice commands.
[0068] Example 1
[0069] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0070] Conventional drone control systems require specialized piloting skills and complex operations, making them difficult for average users to use. While voice-input systems existed, they faced problems with voice recognition accuracy and real-time response, making it difficult to generate efficient flight commands. Furthermore, communication delays and poor reliability meant that safe and accurate flight was not guaranteed.
[0071] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0072] In this invention, the server includes a voice input unit that captures a user's voice input, a voice recognition unit that converts the voice into text data, an analysis unit that analyzes the voice using generative artificial intelligence to acquire destination information, a command generation unit that generates drone flight commands based on the destination information, a communication unit, and a drone control unit. This allows general users to easily pilot drones to their destinations using only voice input, eliminating the need for piloting skills or specialized knowledge that was previously required. Furthermore, the use of a high-speed communication protocol enables safe and accurate flight in real time.
[0073] "Voice input means" refers to an apparatus or device that a user uses to input voice, and includes smartphones and dedicated devices.
[0074] "Speech recognition means" refers to a technology or function that converts captured speech into text data, such as a speech recognition engine.
[0075] "Analysis means" refers to the technology or function for analyzing the converted text data and obtaining destination information based on the user's instructions, and includes generative artificial intelligence.
[0076] The "command generation means" refers to a technology or function that generates a drone flight command including a specific flight path based on the destination information obtained by the analysis means.
[0077] "Communication means" refers to the technology and functions for transmitting the generated flight commands to the drone's control platform, and is a means that uses high-speed communication protocols, etc.
[0078] "Drone control means" refers to the technology and functions that control drones based on flight commands sent via communication means and perform operations such as takeoff, flight, hovering, and landing.
[0079] "Generative AI" refers to an AI technology that analyzes text data and understands the user's instructions, and generative AI models are examples of this type of AI.
[0080] "High-speed communication protocols" refer to communication protocols and standards for high-speed, reliable data communication, and examples of such protocols include TCP / IP and MQTT.
[0081] The present invention relates to a system that allows a user to freely control a drone using voice input. The system includes a voice input unit, a voice recognition unit, an analysis unit, a command generation unit, a communication unit, and a drone control unit.
[0082] First, the user issues commands to the drone using a voice input method, which is expected to be implemented on a smartphone or dedicated device. For example, the user can issue a command such as "Go to City Hall."
[0083] Next, the device captures the user's voice and converts the voice data into text using a speech recognition engine such as Google Cloud Speech-to-Text, which converts the voice into text such as "Go to City Hall."
[0084] The converted text data is sent from the device to a server. The server analyzes the text data using an analytical method that uses generative artificial intelligence (generative AI) to understand the user's instructions. For example, OpenAI GPT-3 is used as a generative AI model. This analyzes the instruction "Go to city hall" and obtains the location information of city hall.
[0085] Next, the server generates a flight command including a specific flight path based on the destination information obtained by the analysis means, such as "take off -> move north 100 meters -> move east 200 meters -> set altitude to 50 meters -> hover above City Hall -> land."
[0086] The generated flight commands are sent to the drone's control platform via the server's communication means using a high-speed communication protocol (e.g., TCP / IP, MQTT). This communication means uses a high-speed, highly reliable protocol to ensure real-time performance.
[0087] Finally, based on the flight commands received by the drone's control board, the drone will fly along the designated path to City Hall, automatically performing a series of actions including takeoff, changing direction, adjusting altitude, hovering, and landing.
[0088] This allows users to easily control a drone to its destination using only voice input. This system does not require the piloting skills or specialized knowledge that was previously required, and is easy for even general users to use.
[0089] For example, if a user says "to the supermarket parking lot," the voice is captured in the same steps and converted into text by speech recognition. The AI then obtains the supermarket's location information, generates a specific flight path, and sends it to the drone, which then flies to the supermarket automatically.
[0090] Example prompt sentence:
[0091] Generate a specific flight route for the voice command "Go to City Hall." Example: "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0092] or
[0093] Provide example text that generates a flight route when the user requests "to the supermarket parking lot."
[0094] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0095] Step 1:
[0096] The user issues commands to the drone using voice input. Specifically, the user picks up their smartphone, launches the voice input application, presses the record button on the screen, and inputs the command by voice, such as "Go to City Hall." This input is captured as voice data.
[0097] Output example: Voice data (voice saying "Go to City Hall")
[0098] Step 2:
[0099] The device receives the voice data captured by the voice input means and converts this voice data into text data using a voice recognition engine (e.g., Google Cloud Speech-to-Text). The voice data is input, and the text data "Go to City Hall" is output.
[0100] Output example: Text data (the text "Go to City Hall")
[0101] Step 3:
[0102] The terminal sends the converted text data to the server using an HTTP POST request, which takes text data as input and sends the data to the server as a result.
[0103] Output example: Text data sent to the server
[0104] Step 4:
[0105] The server uses a generative AI (e.g., OpenAI GPT-3) to analyze the received text data. The generative AI analyzes the text data and understands the instruction, "Go to city hall." Specifically, the location information of city hall is obtained based on the analyzed instruction. Here, the generative AI receives the prompt text, "Go to city hall," as input and outputs location information data.
[0106] Output example: Destination location data
[0107] Step 5:
[0108] The server generates flight commands including a specific flight path based on the location information obtained by the analysis means. The generative AI model receives the location information as input and generates a detailed flight path (route from takeoff to destination). For example, this command includes the steps: "Take off -> Move north 100 meters -> Move east 200 meters -> Set altitude to 50 meters -> Hover above City Hall -> Land."
[0109] Output example: Flight command data
[0110] Step 6:
[0111] The server sends the generated flight commands to the drone's control platform using a high-speed communication protocol (e.g., TCP / IP, MQTT). This allows the server to receive flight command data as input and output it by sending it to the drone's control platform.
[0112] Example output: Flight commands sent to the drone control board
[0113] Step 7:
[0114] The drone control board controls the drone based on the flight commands it receives. Specifically, the drone automatically takes off and follows the specified path (100 meters north -> 200 meters east -> set altitude to 50 meters), hovering over City Hall before landing. It receives flight commands as input and controls the drone accordingly.
[0115] Example output: A drone reaching City Hall
[0116] summary
[0117] Through the above processing steps, the user can control the drone to the destination using only voice input. The input, data processing, output and specific operations at each step are explained in detail.
[0118] (Application example 1)
[0119] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0120] Conventional drone control systems allow users to control drones with voice input without any specific technical knowledge, but they have limitations in terms of efficiently transporting cargo to various locations within a logistics center. Furthermore, obtaining location information in real time and generating appropriate flight routes requires complex operations and time. Furthermore, there was a need for a system that could respond to the diverse instructions within a logistics center.
[0121] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0122] In this invention, the server includes a voice input means for capturing user voice input, a voice recognition means for converting the voice captured by the voice input means into text data, a command generation means for generating drone flight commands using generative artificial intelligence based on the text data converted by the voice recognition means, a communication means for transmitting the flight commands generated by the command generation means to a drone control platform, a drone control means for controlling the drone based on the flight commands transmitted by the communication means, a logistics management means for analyzing instructions and generating a flight path for transporting cargo at a logistics center, and a location information acquisition means for identifying the location and destination of the cargo based on the user's instructions and generating a flight path to the destination. This enables efficient cargo transportation within a logistics center using only voice input.
[0123] 1. "Voice input means" means a device for capturing a user's voice instructions, and includes a voice capture device such as a mobile terminal such as a smartphone or tablet.
[0124] 2. "Speech recognition means" means a technology for converting captured voice data into text data, and is a device that converts voice into text using a speech recognition API.
[0125] 3. "Command generation means" means a device that analyzes the converted text data using artificial intelligence to generate flight commands that instruct the drone's flight path and operations.
[0126] 4. "Communication means" refers to the devices and protocols used to transmit generated flight commands to the drone's control platform, enabling high-speed communication.
[0127] 5. "Drone control means" means a device for operating a drone based on received flight commands, and is a control platform that executes flight paths, adjusts altitude, etc.
[0128] 6. "Logistics management means" refers to a device that analyzes instructions and generates flight paths when transporting cargo using drones within a logistics center, and is a system that supports efficient cargo transportation.
[0129] 7. "Location information acquisition means" means a device for identifying the location and destination of a package based on user instructions, and a means for providing the location data necessary for generating a flight path.
[0130] The system for implementing this invention automates a series of processes from voice input to drone transportation. This system is designed to allow users to freely control drones through voice input.
[0131] First, the user uses a voice input means to give instructions to the drone, such as "Transport the package from shelf A2 to dock C." This voice input means is expected to be implemented on a mobile device such as a smartphone or tablet. A speech recognition API such as the Google Cloud Speech-to-Text API or the Apple Speech Framework is installed on the smartphone or tablet.
[0132] Next, the terminal captures the user's voice and converts the voice data into text data using a voice recognition means, so the voice is converted into text such as "Please carry the package from shelf A2 to dock C."
[0133] The converted text data is sent from the device to a server, where it is processed by a command generation means implemented on the server. Specifically, a generative artificial intelligence (e.g., a generative model such as GPT-3) on the server analyzes the text data and understands the user's instructions. Based on the user's instructions, a location information acquisition means identifies the location and destination of the package and generates a specific flight path using the Google Maps API.
[0134] The generation AI generates flight commands on the server, including the optimal flight path for cargo transportation. This flight command specifies detailed actions from takeoff to the destination. For example, it may include a path such as "Takeoff -> Get cargo from shelf A2 -> Proceed 100 meters north -> Drop cargo at dock C -> Hover -> Land."
[0135] The generated flight commands are sent to the drone's control platform via the server's communications means using a high-speed communications protocol. Real-time communication is ensured by using a high-speed and highly reliable protocol. Specifically, communications means such as Bluetooth, Wi-Fi, or LTE are considered.
[0136] Finally, based on the flight commands received by the drone's control board, the drone will carry out the cargo delivery along the designated path, automatically performing a series of actions including takeoff, heading, altitude adjustment, hovering, and landing.
[0137] This allows users to easily transport cargo within a logistics center using only voice input. The system does not require the piloting skills or specialized knowledge that was previously required, and can be easily used by general users.
[0138] For example, if a user says, "Please transport the package from shelf B3 to the shipping area," the voice input is captured and converted into text using a speech recognition system. The AI then obtains the location information of the shipping area, generates a specific flight path, and sends it to the drone. Finally, the drone automatically transports the package from shelf B3 to the shipping area.
[0139] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0140] Step 1:
[0141] A user uses a smartphone or tablet to give voice commands to the drone, for example, "Transport the package from shelf A2 to dock C." The input is the user's voice command, which is captured by the smartphone or tablet's microphone. The output is the voice data.
[0142] Step 2:
[0143] A speech recognition API (such as Google Cloud Speech-to-Text or Apple Speech Framework) installed on the device converts the captured voice data into text data. The input is voice data, which is converted into a string using the API. The output is text data such as "Carry the package from shelf A2 to dock C."
[0144] Step 3:
[0145] The converted text data is sent from the terminal to the server. The input is the text data, which is sent to the server via a high-speed communication protocol (Wi-Fi, LTE, etc.). The output is the text data received by the server.
[0146] Step 4:
[0147] The server analyzes the received text data using generative artificial intelligence (generative models such as GPT-3). The input is text data, and the generative AI understands its content and interprets the user's instructions. The output is the analyzed instructions.
[0148] Step 5:
[0149] The server uses a location acquisition method (such as Google Maps API) to determine the location and destination of the package. The input is the parsed instructions, and the API is used to obtain specific location data. The output is the package location and the destination location.
[0150] Step 6:
[0151] The server generates a specific flight path based on the location information, which includes a detailed route from takeoff to the destination. The input is the location information of the baggage and the location information of the destination, and the output is a specific flight command.
[0152] Step 7:
[0153] The server sends the generated flight commands to the drone's control board using a high-speed communication protocol (Wi-Fi, Bluetooth, LTE, etc.). The input is the flight command, which is transmitted to the drone via the communication means. The output is the flight command received by the drone.
[0154] Step 8:
[0155] The drone's control board automatically controls the drone based on the flight commands it receives. The input is the flight commands, which control altitude adjustment, direction of movement, speed, etc. The output is the drone's operation. Specifically, the drone retrieves the package from shelf A2 and transports it to dock C.
[0156] Through the above processing steps, users can efficiently transport drones within a logistics center using only voice input.
[0157] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0158] The present invention provides a system that can set optimal flight parameters according to the user's emotional state by combining an emotion engine with a system that automatically controls a drone based on the user's voice instructions.
[0159] First, the user uses voice input to give instructions to the drone, such as "Go to City Hall."
[0160] Next, the terminal captures the user's voice with a microphone. This captured voice data is converted into text data using a voice recognition means. The voice recognition means analyzes the voice data and generates text data such as "Go to City Hall."
[0161] The converted text data is sent from the device to the server. The command generation means implemented on the server passes the text data to the generation artificial intelligence (generation AI) and analyzes the user's instructions. The generation AI obtains the location information of the destination based on the instructions.
[0162] The server also has an emotion engine built in. This emotion engine recognizes the user's emotional state from the voice input. For example, if the user is nervous, it will recognize that as the corresponding emotional state.
[0163] The generative AI then adjusts flight parameters (speed and altitude) via an emotion engine based on the user's perceived emotional state. For example, if the user is nervous, the flight speed will be slowed down.
[0164] The AI then generates flight commands including a specific flight path, such as "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0165] The generated flight commands are sent to the drone's control platform via the server's communications means using a high-speed communication protocol, ensuring real-time communication.
[0166] Finally, the drone's control board controls the drone based on the flight commands it receives. The drone follows the specified flight path to its destination and automatically performs a series of operations, such as hovering and landing.
[0167] For example, if a user says "to the supermarket parking lot," the voice input is captured and converted into text using speech recognition. The generative AI then recognizes the user's emotional state through its emotion engine and retrieves the destination information. The flight parameters are then adjusted based on the emotional state, a specific flight path is generated, and transmitted to the drone, and the drone then flies automatically to the supermarket.
[0168] This system allows users to fly safely and securely based on their emotional state, and allows them to easily control the drone using only voice commands.
[0169] The processing flow will be explained below.
[0170] Step 1:
[0171] The user can use voice input to give instructions to the drone, for example, "Go to City Hall."
[0172] Step 2:
[0173] The device captures the user's voice with a microphone.
[0174] Step 3:
[0175] The device sends the captured voice data to a speech recognition engine, which analyzes the voice data and converts it into corresponding text data ("Go to City Hall").
[0176] Step 4:
[0177] The terminal transmits the converted text data to the server.
[0178] Step 5:
[0179] The server passes the received text data to the generation AI, which analyzes the text data and understands the instruction, "Go to city hall."
[0180] Step 6:
[0181] The server then passes the text data to the emotion engine, which analyzes and evaluates the user's emotional state from the voice data.
[0182] Step 7:
[0183] The server uses the generated AI to obtain the location information of the destination (city hall) based on the user's instructions.
[0184] Step 8:
[0185] The server uses generated AI to adjust flight parameters (speed, altitude, etc.) based on the acquired location information and the analysis results of the emotion engine.
[0186] Step 9:
[0187] The server instructs the AI to generate a specific flight path including flight parameters, causing it to generate flight commands.
[0188] Example: "Take off -> fly 100 meters north -> fly 200 meters east -> set altitude to 50 meters (or lower if user is nervous) -> hover over City Hall -> land."
[0189] Step 10:
[0190] The server sends flight commands to the drone's control board using a communication means.
[0191] Step 11:
[0192] The drone's control board analyzes the flight commands it receives and converts them into specific operational instructions.
[0193] Step 12:
[0194] The drone's control board controls the drone's motors and propellers, making the drone operate according to flight commands.
[0195] Step 13:
[0196] The drone's control board executes a series of operations from takeoff to the destination, finally reaching and landing at the specified location.
[0197] This allows users to easily control the drone using only voice commands, and also ensures safe and optimal flight based on the user's emotional state.
[0198] Example 2
[0199] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0200] Conventional drone control systems have mechanisms to convert user instructions into text data and generate flight commands based on destination information, but they lack the functionality to optimize flight parameters based on the user's emotional state. As a result, if the user becomes nervous or excited, the drone's flight cannot adapt to the user's emotional state, making it difficult to achieve a safe and secure flight.
[0201] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0202] In this invention, the server includes a voice input means for capturing a user's voice input, a voice recognition means for converting the voice captured by the voice input means into text data, an emotion analysis means for recognizing the user's emotional state from the voice data, a means for adjusting flight parameters based on the emotional state recognized by the emotion analysis means, a means for generating flight commands including the adjusted flight parameters, a communication means for transmitting the flight commands generated by the command generation means to a drone control platform, and a drone control means for controlling the drone based on the flight commands transmitted by the communication means. This makes it possible to set optimal flight parameters based on the user's emotional state, thereby achieving safe and secure flight.
[0203] "Audio input means" refers to a device or system for capturing a user's voice.
[0204] A "speech recognition means" is a device or algorithm that converts captured speech into text data.
[0205] A "command generation means" is a system or device for generating drone flight commands using artificial intelligence based on text data.
[0206] "Emotion analysis means" refers to a system or algorithm for recognizing a user's emotional state from their voice data.
[0207] The "flight parameter adjustment means" is a system or device for adjusting flight parameters based on the emotional state recognized by the emotion analysis means.
[0208] "Communication means" refers to a device or system for transmitting the generated flight commands to the drone's control board.
[0209] A "drone control means" is a system or device for controlling a drone based on received flight commands.
[0210] MODE FOR CARRYING OUT THE INVENTION
[0211] The present invention provides a system that combines an emotion analysis function with a system that automatically controls a drone based on the user's voice instructions, and can set optimal flight parameters according to the user's emotional state.
[0212] First, the user issues commands to the drone using a voice input means. For example, the user might issue a voice command such as "Go to City Hall." Next, the device captures the user's voice with a microphone. This captured voice data is converted into text data using a voice recognition means. Specifically, a voice recognition system such as Google Speech-to-Text or Amazon Transcribe is used. The voice recognition means analyzes the voice data and generates the text data "Go to City Hall."
[0213] The converted text data is sent from the device to the server. Standard communication protocols such as HTTP and WebSocket are used for communication. The command generation means implemented on the server passes the text data to the generation AI, which analyzes the user's instructions. OpenAI GPT and other programs can be used as the generation AI. The generation AI obtains the location information of the destination based on the instructions.
[0214] Additionally, the server is equipped with an emotion analysis engine that recognizes the user's emotional state from the voice input. For example, if the user is nervous, it will be recognized as such. The emotion engine achieves this by using an algorithm that analyzes acoustic features (such as tone and pitch of voice).
[0215] Next, the generation AI adjusts flight parameters (speed and altitude) through the emotion engine based on the recognized emotional state. If the user is nervous, adjustments will be made, such as slowing down the flight speed. Based on the adjusted flight parameters, the generation AI generates a specific flight path and creates flight commands. For example, a series of commands might be "Take off -> Move north 100 meters -> Move right (east) 200 meters -> Set altitude to 50 meters -> Hover above City Hall -> Land."
[0216] The generated flight commands are sent to the drone's control platform via the server's communication means using a high-speed communication protocol (such as MQTT or TCP / IP). The drone's control platform ultimately controls the drone based on the received flight commands. The drone follows the specified flight path to its destination and automatically performs operations such as hovering and landing.
[0217] As a specific example, the operation when the user instructs "to the supermarket parking lot" will be shown.
[0218] 1. The user says, "To the supermarket parking lot."
[0219] 2. The device captures the audio with its microphone.
[0220] 3. The speech recognition system converts the speech data into text data.
[0221] 4. The device sends the converted text data to the server.
[0222] 5. The server's generated AI analyzes the text data and obtains the destination's location information.
[0223] 6. The emotion analysis means built into the server recognizes the user's emotional state from the voice data. For example, if the user is excited, it will be recognized as that emotional state.
[0224] 7. The AI generator adjusts flight parameters based on the user's emotional state. For example, if the user is excited, the AI will increase flight speed.
[0225] 8. The generation AI generates a specific flight path and creates flight commands.
[0226] 9. The server sends flight commands to the drone's control board via communication means.
[0227] 10. The drone's control board controls the drone based on the commands, and it flies automatically to the supermarket parking lot.
[0228] Examples of prompt sentences are as follows:
[0229] "Go to City Hall."
[0230] "Fly to the park fountain"
[0231] "Come back to my garden."
[0232] This system allows users to fly safely and securely based on their emotional state, and allows them to easily control the drone using only voice commands.
[0233] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0234] Step 1:
[0235] The user issues a voice command using the voice input means, for example, "Go to city hall." The input here is the user's voice, and the output is the voice data.
[0236] Step 2:
[0237] The device uses a microphone to capture the user's voice, and in the next step, this voice data is passed to a voice recognition system. The input is the user's voice, and the output is the captured voice data.
[0238] Step 3:
[0239] The voice data captured by the device is input into a voice recognition system (for example, Google Speech-to-Text or Amazon Transcribe) and converted into text data. In this conversion, the voice data "Go to City Hall" is converted into text data "Go to City Hall." The input is the captured voice data, and the output is the converted text data.
[0240] Step 4:
[0241] The terminal sends the converted text data to the server. HTTP or WebSocket is used for communication. The input is the converted text data, and the output is data sent to the server.
[0242] Step 5:
[0243] The server passes the received text data to the generation AI for analysis. The generation AI (for example, OpenAI GPT) analyzes the text data to obtain destination information. For example, it analyzes the text data "Go to city hall" to obtain the location information (latitude, longitude, etc.) of city hall. The input is text data, and the output is the analyzed destination information.
[0244] Step 6:
[0245] The server uses emotion analysis to recognize the user's emotional state from the voice data. The emotion analysis engine analyzes acoustic features (tone of voice, pitch, etc.) and estimates the user's emotional state, such as nervousness or calmness. The input is the voice data, and the output is the recognized emotional state.
[0246] Step 7:
[0247] The server's generation AI adjusts flight parameters (speed, altitude, etc.) based on the recognized emotional state. For example, if the user is nervous, the flight speed will be slowed down. A specific flight path is generated based on these flight parameters, and flight commands are created. The input is the emotional state and destination information, and the output is the adjusted flight parameters and specific flight commands.
[0248] Step 8:
[0249] The flight commands generated by the server are sent to the drone's control platform via a communication method. High-speed communication protocols such as MQTT and TCP / IP are used for communication. The input is the flight command, and the output is the command sent to the drone's control platform.
[0250] Step 9:
[0251] The drone will fly automatically based on flight commands received through the control board. The drone will follow the specified flight path to its destination and automatically perform operations such as landing and hovering. The input is the flight command, and the output is the drone's actual flight behavior.
[0252] (Application example 2)
[0253] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0254] In conventional food delivery systems, it is difficult to control drones based on voice commands, and it is not possible to set optimal flight parameters according to the user's emotional state, which can lead to anxiety and stress during delivery.In addition, it is difficult to adjust flight parameters in real time, leaving issues in terms of safety and efficiency.
[0255] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0256] In this invention, the server includes a voice input means for capturing a user's voice input, a voice recognition means for converting the captured voice into text data, an emotion recognition means for recognizing the user's emotional state using an emotion engine, a parameter adjustment means for setting drone flight parameters according to the user's emotional state, a command generation means for generating drone flight commands using generative artificial intelligence, a communication means for transmitting the generated flight commands to a drone control platform, and a drone control means for controlling the drone based on the transmitted flight commands. This allows the user to simply issue voice commands to set optimal drone flight parameters according to their emotional state, enabling safe and efficient food delivery.
[0257] "Audio input means" is a device or function for capturing a user's voice.
[0258] "Speech recognition means" refers to technology or devices that analyze captured voice data and convert it into text data.
[0259] "Emotion recognition means" refers to technology or devices for recognizing a user's emotional state from voice data or text data.
[0260] "Parameter adjustment means" refers to a function for setting and adjusting the drone's flight parameters (e.g., speed, altitude, etc.) based on the recognized emotional state of the user.
[0261] A "command generation means" is a device or function that generates flight commands for the drone based on the user's voice input and emotional state.
[0262] "Communication means" refers to the method or technology used to transmit the generated flight commands to the drone's control platform.
[0263] "Drone control means" means a device or function for controlling and managing a drone based on received flight commands.
[0264] "Generative AI" is an AI technology that uses large amounts of data to analyze user instructions and generate appropriate commands.
[0265] The "emotion engine" is a technology that detects the user's emotional state from voice and text data and uses it to adjust flight parameters.
[0266] This invention is a drone control system that can capture the user's voice instructions and combine them with an emotion engine to set optimal flight parameters according to the user's emotional state.
[0267] The server first has a voice input means for capturing voice input from the user, which is a voice input device such as a microphone, and captures voice instructions given by the user.
[0268] Once the audio is captured, a speech recognition means processes it and converts it into text data, for example, using the speech_recognition library, which generates text data from the user's voice commands.
[0269] Next, the emotion recognition means recognizes the user's emotional state from the voice. The emotion recognition means uses an emotion engine to extract emotions from the voice. An example of this emotion engine is the EmotionRecognizer module.
[0270] Based on the recognized emotional state, the parameter adjustment means sets and adjusts the optimal drone flight parameters, including the drone's speed and altitude. For example, if the user is nervous, the flight speed may be slowed down.
[0271] The command generation means then uses a generative AI model to generate specific flight commands. This generative AI model embodies the drone's flight route and operations based on the generated text data and the user's emotional state. For example, it generates detailed commands for takeoff, flight path, altitude, hovering, landing, etc.
[0272] The generated flight commands are sent to the drone's control platform via a communication method that uses a high-speed, highly reliable protocol to ensure real-time communication, such as Wi-Fi or 4G / 5G communication.
[0273] Based on the transmitted flight commands, the drone control means controls the drone, allowing the drone to travel to the destination along the specified flight route, achieving safe and efficient flight according to the user's emotional state.
[0274] Specific examples
[0275] Example prompt sentence:
[0276] The user gives a voice command such as "Deliver pizza to my house."
[0277] In the processing flow, voice input is captured and converted into text data such as "Deliver pizza to my house" by the voice recognition means. Then, the emotion recognition means recognizes the user's emotional state, and the parameter adjustment means sets flight parameters. Finally, the generated flight command is sent to the drone, which then delivers the pizza to the house.
[0278] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0279] Step 1: Capturing Audio Input
[0280] The server uses a voice input means (microphone) to capture the user's voice instruction. This input is, for example, the user's voice data saying, "Deliver a pizza to my house." The captured voice data is saved for processing in a later step.
[0281] Step 2: Voice Recognition
[0282] The server uses a speech recognition tool (the speech_recognition library) to convert the captured voice data into text data, which contains specific instructions such as "Deliver a pizza to my house." The converted text data is used in the next processing step.
[0283] Step 3: Emotion Recognition
[0284] The server uses an emotion recognition module (EmotionRecognizer module) to recognize the user's emotional state based on the converted text data. It analyzes the features extracted from the voice data and identifies the user's emotional state, such as "joy." The recognized emotional state is then used to set flight parameters.
[0285] Step 4: Obtaining destination information
[0286] The server uses the command generation means to obtain destination information from the user's voice input. Specifically, it analyzes the text data "Deliver pizza to my house" and obtains the coordinates (latitude and longitude) of "home" using a geographic information acquisition library (such as geopy). The obtained coordinate information is used to generate a flight route.
[0287] Step 5: Adjusting Flight Parameters
[0288] The server uses a parameter adjustment means to set the drone's flight parameters according to the recognized emotional state. For example, if the user is "nervous," the server sets the drone's flight speed slower to increase safety. This adjustment allows the drone to fly in a way that takes the user's emotions into consideration.
[0289] Step 6: Generate Flight Commands
[0290] The server generates specific flight commands using a command generation means and a generative AI model. The generative AI model outputs detailed flight commands, such as takeoff, route, altitude setting, hovering, and landing, based on the user's destination information and emotional state. These flight commands serve as guidelines for the drone's operation.
[0291] Step 7: Sending Flight Commands
[0292] The server then transmits the generated flight commands to the drone's control platform via a communication method using a high-speed, highly reliable protocol (such as Wi-Fi or 4G / 5G communication), enabling accurate command transmission in real time.
[0293] Step 8: Controlling the drone
[0294] The drone operates based on the flight commands received. The drone control means flies the pizza to the customer's home according to the specified flight route and parameters, realizing safe and efficient food delivery that takes into consideration the customer's feelings.
[0295] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0296] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0297] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0298] [Second embodiment]
[0299] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0300] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0301] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0302] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0303] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0304] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0305] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0306] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0307] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0308] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0309] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0310] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0311] To implement the present invention, the following system is required: This system is designed to allow a user to freely control a drone through voice input.
[0312] First, the user uses voice input to give the drone instructions such as "Go to City Hall." This voice input is expected to be implemented on a smartphone or dedicated device.
[0313] Next, the device captures the user's voice and converts the voice data into text data using a speech recognition means. At this stage, a speech recognition engine is used to convert the voice into text, such as "Go to City Hall."
[0314] The converted text data is sent from the device to a server and processed by a command generation means implemented on the server. Specifically, the generative artificial intelligence (generative AI) on the server analyzes the text data and understands the user's instructions. For example, the instruction "Go to city hall" is analyzed, and the location information of city hall is obtained.
[0315] Next, the AI generates flight commands that include a specific flight path to City Hall. These flight commands specify detailed actions from takeoff to the destination, such as "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0316] The generated flight commands are sent to the drone's control platform via the server's communications means using a high-speed communication protocol, ensuring real-time communication.
[0317] Finally, based on the flight commands received by the drone's control board, the drone will fly along the designated path to City Hall, automatically performing a series of operations including takeoff, heading, altitude adjustment, hovering, and landing.
[0318] This allows users to easily control a drone to its destination using only voice input. This system does not require the piloting skills or specialized knowledge that was previously required, and is easy for even general users to use.
[0319] For example, if a user says, "To the supermarket parking lot," the voice input is captured and converted into text using a speech recognition system. The AI then obtains the supermarket's location information, generates a specific flight path, and sends it to the drone, which then automatically flies to the supermarket.
[0320] The processing flow will be explained below.
[0321] Step 1:
[0322] The user uses voice input to give instructions to the drone, such as "Go to City Hall."
[0323] Step 2:
[0324] The device captures the user's voice with a microphone.
[0325] Step 3:
[0326] The device sends the captured voice data to a speech recognition engine, which analyzes the voice data and converts it into corresponding text data ("Go to City Hall").
[0327] Step 4:
[0328] The terminal transmits the converted text data to the server.
[0329] Step 5:
[0330] The server passes the received text data to the generation AI, which analyzes the text data and understands the user's instructions.
[0331] Step 6:
[0332] The server uses a generation AI to obtain the location information of the destination (in this case, city hall) based on the user's instructions.
[0333] Step 7:
[0334] Based on the location information of the destination acquired by the server, the generation AI generates a specific flight path.
[0335] Example: "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0336] Step 8:
[0337] The server transmits the generated flight commands to the drone's control board using a communication means.
[0338] Step 9:
[0339] The drone's control board analyzes the flight commands it receives and converts them into specific operational instructions.
[0340] Step 10:
[0341] The drone's control board controls the drone's motors and propellers, allowing the drone to operate according to flight commands.
[0342] Step 11:
[0343] The drone's control board executes a series of operations from takeoff to the destination, finally reaching and landing at the specified location.
[0344] This allows users to easily control the drone using only voice commands.
[0345] Example 1
[0346] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0347] Conventional drone control systems require specialized piloting skills and complex operations, making them difficult for average users to use. While voice-input systems existed, they faced problems with voice recognition accuracy and real-time response, making it difficult to generate efficient flight commands. Furthermore, communication delays and poor reliability meant that safe and accurate flight was not guaranteed.
[0348] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0349] In this invention, the server includes a voice input unit that captures a user's voice input, a voice recognition unit that converts the voice into text data, an analysis unit that analyzes the voice using generative artificial intelligence to acquire destination information, a command generation unit that generates drone flight commands based on the destination information, a communication unit, and a drone control unit. This allows general users to easily pilot drones to their destinations using only voice input, eliminating the need for piloting skills or specialized knowledge that was previously required. Furthermore, the use of a high-speed communication protocol enables safe and accurate flight in real time.
[0350] "Voice input means" refers to an apparatus or device that a user uses to input voice, and includes smartphones and dedicated devices.
[0351] "Speech recognition means" refers to a technology or function that converts captured speech into text data, such as a speech recognition engine.
[0352] "Analysis means" refers to the technology or function for analyzing the converted text data and obtaining destination information based on the user's instructions, and includes generative artificial intelligence.
[0353] The "command generation means" refers to a technology or function that generates a drone flight command including a specific flight path based on the destination information obtained by the analysis means.
[0354] "Communication means" refers to the technology and functions for transmitting the generated flight commands to the drone's control platform, and is a means that uses high-speed communication protocols, etc.
[0355] "Drone control means" refers to the technology and functions that control drones based on flight commands sent via communication means and perform operations such as takeoff, flight, hovering, and landing.
[0356] "Generative AI" refers to an AI technology that analyzes text data and understands the user's instructions, and generative AI models are examples of this type of AI.
[0357] "High-speed communication protocols" refer to communication protocols and standards for high-speed, reliable data communication, and examples of such protocols include TCP / IP and MQTT.
[0358] The present invention relates to a system that allows a user to freely control a drone using voice input. The system includes a voice input unit, a voice recognition unit, an analysis unit, a command generation unit, a communication unit, and a drone control unit.
[0359] First, the user issues commands to the drone using a voice input method, which is expected to be implemented on a smartphone or dedicated device. For example, the user can issue a command such as "Go to City Hall."
[0360] Next, the device captures the user's voice and converts the voice data into text using a speech recognition engine such as Google Cloud Speech-to-Text, which converts the voice into text such as "Go to City Hall."
[0361] The converted text data is sent from the device to a server. The server analyzes the text data using an analytical method that uses generative artificial intelligence (generative AI) to understand the user's instructions. For example, OpenAI GPT-3 is used as a generative AI model. This analyzes the instruction "Go to city hall" and obtains the location information of city hall.
[0362] Next, the server generates a flight command including a specific flight path based on the destination information obtained by the analysis means, such as "take off -> move north 100 meters -> move east 200 meters -> set altitude to 50 meters -> hover above City Hall -> land."
[0363] The generated flight commands are sent to the drone's control platform via the server's communication means using a high-speed communication protocol (e.g., TCP / IP, MQTT). This communication means uses a high-speed, highly reliable protocol to ensure real-time performance.
[0364] Finally, based on the flight commands received by the drone's control board, the drone will fly along the designated path to City Hall, automatically performing a series of actions including takeoff, changing direction, adjusting altitude, hovering, and landing.
[0365] This allows users to easily control a drone to its destination using only voice input. This system does not require the piloting skills or specialized knowledge that was previously required, and is easy for even general users to use.
[0366] For example, if a user says "to the supermarket parking lot," the voice is captured in the same steps and converted into text by speech recognition. The AI then obtains the supermarket's location information, generates a specific flight path, and sends it to the drone, which then flies to the supermarket automatically.
[0367] Example prompt sentence:
[0368] Generate a specific flight route for the voice command "Go to City Hall." Example: "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0369] or
[0370] Provide example text that generates a flight route when the user requests "to the supermarket parking lot."
[0371] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0372] Step 1:
[0373] The user issues commands to the drone using voice input. Specifically, the user picks up their smartphone, launches the voice input application, presses the record button on the screen, and inputs the command by voice, such as "Go to City Hall." This input is captured as voice data.
[0374] Output example: Voice data (voice saying "Go to City Hall")
[0375] Step 2:
[0376] The device receives the voice data captured by the voice input means and converts this voice data into text data using a voice recognition engine (e.g., Google Cloud Speech-to-Text). The voice data is input, and the text data "Go to City Hall" is output.
[0377] Output example: Text data (the text "Go to City Hall")
[0378] Step 3:
[0379] The terminal sends the converted text data to the server using an HTTP POST request, which takes text data as input and sends the data to the server as a result.
[0380] Output example: Text data sent to the server
[0381] Step 4:
[0382] The server uses a generative AI (e.g., OpenAI GPT-3) to analyze the received text data. The generative AI analyzes the text data and understands the instruction, "Go to city hall." Specifically, the location information of city hall is obtained based on the analyzed instruction. Here, the generative AI receives the prompt text, "Go to city hall," as input and outputs location information data.
[0383] Output example: Destination location data
[0384] Step 5:
[0385] The server generates flight commands including a specific flight path based on the location information obtained by the analysis means. The generative AI model receives the location information as input and generates a detailed flight path (route from takeoff to destination). For example, this command includes the steps: "Take off -> Move north 100 meters -> Move east 200 meters -> Set altitude to 50 meters -> Hover above City Hall -> Land."
[0386] Output example: Flight command data
[0387] Step 6:
[0388] The server sends the generated flight commands to the drone's control platform using a high-speed communication protocol (e.g., TCP / IP, MQTT). This allows the server to receive flight command data as input and output it by sending it to the drone's control platform.
[0389] Example output: Flight commands sent to the drone control board
[0390] Step 7:
[0391] The drone control board controls the drone based on the flight commands it receives. Specifically, the drone automatically takes off and follows the specified path (100 meters north -> 200 meters east -> set altitude to 50 meters), hovering over City Hall before landing. It receives flight commands as input and controls the drone accordingly.
[0392] Example output: A drone reaching City Hall
[0393] summary
[0394] Through the above processing steps, the user can control the drone to the destination using only voice input. The input, data processing, output and specific operations at each step are explained in detail.
[0395] (Application example 1)
[0396] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0397] Conventional drone control systems allow users to control drones with voice input without any specific technical knowledge, but they have limitations in terms of efficiently transporting cargo to various locations within a logistics center. Furthermore, obtaining location information in real time and generating appropriate flight routes requires complex operations and time. Furthermore, there was a need for a system that could respond to the diverse instructions within a logistics center.
[0398] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0399] In this invention, the server includes a voice input means for capturing user voice input, a voice recognition means for converting the voice captured by the voice input means into text data, a command generation means for generating drone flight commands using generative artificial intelligence based on the text data converted by the voice recognition means, a communication means for transmitting the flight commands generated by the command generation means to a drone control platform, a drone control means for controlling the drone based on the flight commands transmitted by the communication means, a logistics management means for analyzing instructions and generating a flight path for transporting cargo at a logistics center, and a location information acquisition means for identifying the location and destination of the cargo based on the user's instructions and generating a flight path to the destination. This enables efficient cargo transportation within a logistics center using only voice input.
[0400] 1. "Voice input means" means a device for capturing a user's voice instructions, and includes a voice capture device such as a mobile terminal such as a smartphone or tablet.
[0401] 2. "Speech recognition means" means a technology for converting captured voice data into text data, and is a device that converts voice into text using a speech recognition API.
[0402] 3. "Command generation means" means a device that analyzes the converted text data using artificial intelligence to generate flight commands that instruct the drone's flight path and operations.
[0403] 4. "Communication means" refers to the devices and protocols used to transmit generated flight commands to the drone's control platform, enabling high-speed communication.
[0404] 5. "Drone control means" means a device for operating a drone based on received flight commands, and is a control platform that executes flight paths, adjusts altitude, etc.
[0405] 6. "Logistics management means" refers to a device that analyzes instructions and generates flight paths when transporting cargo using drones within a logistics center, and is a system that supports efficient cargo transportation.
[0406] 7. "Location information acquisition means" means a device for identifying the location and destination of a package based on user instructions, and a means for providing the location data necessary for generating a flight path.
[0407] The system for implementing this invention automates a series of processes from voice input to drone transportation. This system is designed to allow users to freely control drones through voice input.
[0408] First, the user uses a voice input means to give instructions to the drone, such as "Transport the package from shelf A2 to dock C." This voice input means is expected to be implemented on a mobile device such as a smartphone or tablet. A speech recognition API such as the Google Cloud Speech-to-Text API or the Apple Speech Framework is installed on the smartphone or tablet.
[0409] Next, the terminal captures the user's voice and converts the voice data into text data using a voice recognition means, so the voice is converted into text such as "Please carry the package from shelf A2 to dock C."
[0410] The converted text data is sent from the device to a server, where it is processed by a command generation means implemented on the server. Specifically, a generative artificial intelligence (e.g., a generative model such as GPT-3) on the server analyzes the text data and understands the user's instructions. Based on the user's instructions, a location information acquisition means identifies the location and destination of the package and generates a specific flight path using the Google Maps API.
[0411] The generation AI generates flight commands on the server, including the optimal flight path for cargo transportation. This flight command specifies detailed actions from takeoff to the destination. For example, it may include a path such as "Takeoff -> Get cargo from shelf A2 -> Proceed 100 meters north -> Drop cargo at dock C -> Hover -> Land."
[0412] The generated flight commands are sent to the drone's control platform via the server's communications means using a high-speed communications protocol. Real-time communication is ensured by using a high-speed and highly reliable protocol. Specifically, communications means such as Bluetooth, Wi-Fi, or LTE are considered.
[0413] Finally, based on the flight commands received by the drone's control board, the drone will carry out the cargo delivery along the designated path, automatically performing a series of actions including takeoff, heading, altitude adjustment, hovering, and landing.
[0414] This allows users to easily transport cargo within a logistics center using only voice input. The system does not require the piloting skills or specialized knowledge that was previously required, and can be easily used by general users.
[0415] For example, if a user says, "Please transport the package from shelf B3 to the shipping area," the voice input is captured and converted into text using a speech recognition system. The AI then obtains the location information of the shipping area, generates a specific flight path, and sends it to the drone. Finally, the drone automatically transports the package from shelf B3 to the shipping area.
[0416] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0417] Step 1:
[0418] A user uses a smartphone or tablet to give voice commands to the drone, for example, "Transport the package from shelf A2 to dock C." The input is the user's voice command, which is captured by the smartphone or tablet's microphone. The output is the voice data.
[0419] Step 2:
[0420] A speech recognition API (such as Google Cloud Speech-to-Text or Apple Speech Framework) installed on the device converts the captured voice data into text data. The input is voice data, which is converted into a string using the API. The output is text data such as "Carry the package from shelf A2 to dock C."
[0421] Step 3:
[0422] The converted text data is sent from the terminal to the server. The input is the text data, which is sent to the server via a high-speed communication protocol (Wi-Fi, LTE, etc.). The output is the text data received by the server.
[0423] Step 4:
[0424] The server analyzes the received text data using generative artificial intelligence (generative models such as GPT-3). The input is text data, and the generative AI understands its content and interprets the user's instructions. The output is the analyzed instructions.
[0425] Step 5:
[0426] The server uses a location acquisition method (such as Google Maps API) to determine the location and destination of the package. The input is the parsed instructions, and the API is used to obtain specific location data. The output is the package location and the destination location.
[0427] Step 6:
[0428] The server generates a specific flight path based on the location information, which includes a detailed route from takeoff to the destination. The input is the location information of the baggage and the location information of the destination, and the output is a specific flight command.
[0429] Step 7:
[0430] The server sends the generated flight commands to the drone's control board using a high-speed communication protocol (Wi-Fi, Bluetooth, LTE, etc.). The input is the flight command, which is transmitted to the drone via the communication means. The output is the flight command received by the drone.
[0431] Step 8:
[0432] The drone's control board automatically controls the drone based on the flight commands it receives. The input is the flight commands, which control altitude adjustment, direction of movement, speed, etc. The output is the drone's operation. Specifically, the drone retrieves the package from shelf A2 and transports it to dock C.
[0433] Through the above processing steps, users can efficiently transport drones within a logistics center using only voice input.
[0434] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0435] The present invention provides a system that can set optimal flight parameters according to the user's emotional state by combining an emotion engine with a system that automatically controls a drone based on the user's voice instructions.
[0436] First, the user uses voice input to give instructions to the drone, such as "Go to City Hall."
[0437] Next, the terminal captures the user's voice with a microphone. This captured voice data is converted into text data using a voice recognition means. The voice recognition means analyzes the voice data and generates text data such as "Go to City Hall."
[0438] The converted text data is sent from the device to the server. The command generation means implemented on the server passes the text data to the generation artificial intelligence (generation AI) and analyzes the user's instructions. The generation AI obtains the location information of the destination based on the instructions.
[0439] The server also has an emotion engine built in. This emotion engine recognizes the user's emotional state from the voice input. For example, if the user is nervous, it will recognize that as the corresponding emotional state.
[0440] The generative AI then adjusts flight parameters (speed and altitude) via an emotion engine based on the user's perceived emotional state. For example, if the user is nervous, the flight speed will be slowed down.
[0441] The AI then generates flight commands including a specific flight path, such as "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0442] The generated flight commands are sent to the drone's control platform via the server's communications means using a high-speed communication protocol, ensuring real-time communication.
[0443] Finally, the drone's control board controls the drone based on the flight commands it receives. The drone follows the specified flight path to its destination and automatically performs a series of operations, such as hovering and landing.
[0444] For example, if a user says "to the supermarket parking lot," the voice input is captured and converted into text using speech recognition. The generative AI then recognizes the user's emotional state through its emotion engine and retrieves the destination information. The flight parameters are then adjusted based on the emotional state, a specific flight path is generated, and transmitted to the drone, and the drone then flies automatically to the supermarket.
[0445] This system allows users to fly safely and securely based on their emotional state, and allows them to easily control the drone using only voice commands.
[0446] The processing flow will be explained below.
[0447] Step 1:
[0448] The user can use voice input to give instructions to the drone, for example, "Go to City Hall."
[0449] Step 2:
[0450] The device captures the user's voice with a microphone.
[0451] Step 3:
[0452] The device sends the captured voice data to a speech recognition engine, which analyzes the voice data and converts it into corresponding text data ("Go to City Hall").
[0453] Step 4:
[0454] The terminal transmits the converted text data to the server.
[0455] Step 5:
[0456] The server passes the received text data to the generation AI, which analyzes the text data and understands the instruction, "Go to city hall."
[0457] Step 6:
[0458] The server then passes the text data to the emotion engine, which analyzes and evaluates the user's emotional state from the voice data.
[0459] Step 7:
[0460] The server uses the generated AI to obtain the location information of the destination (city hall) based on the user's instructions.
[0461] Step 8:
[0462] The server uses generated AI to adjust flight parameters (speed, altitude, etc.) based on the acquired location information and the analysis results of the emotion engine.
[0463] Step 9:
[0464] The server instructs the AI to generate a specific flight path including flight parameters, causing it to generate flight commands.
[0465] Example: "Take off -> fly 100 meters north -> fly 200 meters east -> set altitude to 50 meters (or lower if user is nervous) -> hover over City Hall -> land."
[0466] Step 10:
[0467] The server sends flight commands to the drone's control board using a communication means.
[0468] Step 11:
[0469] The drone's control board analyzes the flight commands it receives and converts them into specific operational instructions.
[0470] Step 12:
[0471] The drone's control board controls the drone's motors and propellers, making the drone operate according to flight commands.
[0472] Step 13:
[0473] The drone's control board executes a series of operations from takeoff to the destination, finally reaching and landing at the specified location.
[0474] This allows users to easily control the drone using only voice commands, and also ensures safe and optimal flight based on the user's emotional state.
[0475] Example 2
[0476] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0477] Conventional drone control systems have mechanisms to convert user instructions into text data and generate flight commands based on destination information, but they lack the functionality to optimize flight parameters based on the user's emotional state. As a result, if the user becomes nervous or excited, the drone's flight cannot adapt to the user's emotional state, making it difficult to achieve a safe and secure flight.
[0478] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0479] In this invention, the server includes a voice input means for capturing a user's voice input, a voice recognition means for converting the voice captured by the voice input means into text data, an emotion analysis means for recognizing the user's emotional state from the voice data, a means for adjusting flight parameters based on the emotional state recognized by the emotion analysis means, a means for generating flight commands including the adjusted flight parameters, a communication means for transmitting the flight commands generated by the command generation means to a drone control platform, and a drone control means for controlling the drone based on the flight commands transmitted by the communication means. This makes it possible to set optimal flight parameters based on the user's emotional state, thereby achieving safe and secure flight.
[0480] "Audio input means" refers to a device or system for capturing a user's voice.
[0481] A "speech recognition means" is a device or algorithm that converts captured speech into text data.
[0482] A "command generation means" is a system or device for generating drone flight commands using artificial intelligence based on text data.
[0483] "Emotion analysis means" refers to a system or algorithm for recognizing a user's emotional state from their voice data.
[0484] The "flight parameter adjustment means" is a system or device for adjusting flight parameters based on the emotional state recognized by the emotion analysis means.
[0485] "Communication means" refers to a device or system for transmitting the generated flight commands to the drone's control board.
[0486] A "drone control means" is a system or device for controlling a drone based on received flight commands.
[0487] MODE FOR CARRYING OUT THE INVENTION
[0488] The present invention provides a system that combines an emotion analysis function with a system that automatically controls a drone based on the user's voice instructions, and can set optimal flight parameters according to the user's emotional state.
[0489] First, the user issues commands to the drone using a voice input means. For example, the user might issue a voice command such as "Go to City Hall." Next, the device captures the user's voice with a microphone. This captured voice data is converted into text data using a voice recognition means. Specifically, a voice recognition system such as Google Speech-to-Text or Amazon Transcribe is used. The voice recognition means analyzes the voice data and generates the text data "Go to City Hall."
[0490] The converted text data is sent from the device to the server. Standard communication protocols such as HTTP and WebSocket are used for communication. The command generation means implemented on the server passes the text data to the generation AI, which analyzes the user's instructions. OpenAI GPT and other programs can be used as the generation AI. The generation AI obtains the location information of the destination based on the instructions.
[0491] Additionally, the server is equipped with an emotion analysis engine that recognizes the user's emotional state from the voice input. For example, if the user is nervous, it will be recognized as such. The emotion engine achieves this by using an algorithm that analyzes acoustic features (such as tone and pitch of voice).
[0492] Next, the generation AI adjusts flight parameters (speed and altitude) through the emotion engine based on the recognized emotional state. If the user is nervous, adjustments will be made, such as slowing down the flight speed. Based on the adjusted flight parameters, the generation AI generates a specific flight path and creates flight commands. For example, a series of commands might be "Take off -> Move north 100 meters -> Move right (east) 200 meters -> Set altitude to 50 meters -> Hover above City Hall -> Land."
[0493] The generated flight commands are sent to the drone's control platform via the server's communication means using a high-speed communication protocol (such as MQTT or TCP / IP). The drone's control platform ultimately controls the drone based on the received flight commands. The drone follows the specified flight path to its destination and automatically performs operations such as hovering and landing.
[0494] As a specific example, the operation when the user instructs "to the supermarket parking lot" will be shown.
[0495] 1. The user says, "To the supermarket parking lot."
[0496] 2. The device captures the audio with its microphone.
[0497] 3. The speech recognition system converts the speech data into text data.
[0498] 4. The device sends the converted text data to the server.
[0499] 5. The server's generated AI analyzes the text data and obtains the destination's location information.
[0500] 6. The emotion analysis means built into the server recognizes the user's emotional state from the voice data. For example, if the user is excited, it will be recognized as that emotional state.
[0501] 7. The AI generator adjusts flight parameters based on the user's emotional state. For example, if the user is excited, the AI will increase flight speed.
[0502] 8. The generation AI generates a specific flight path and creates flight commands.
[0503] 9. The server sends flight commands to the drone's control board via communication means.
[0504] 10. The drone's control board controls the drone based on the commands, and it flies automatically to the supermarket parking lot.
[0505] Examples of prompt sentences are as follows:
[0506] "Go to City Hall."
[0507] "Fly to the park fountain"
[0508] "Come back to my garden."
[0509] This system allows users to fly safely and securely based on their emotional state, and allows them to easily control the drone using only voice commands.
[0510] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0511] Step 1:
[0512] The user issues a voice command using the voice input means, for example, "Go to city hall." The input here is the user's voice, and the output is the voice data.
[0513] Step 2:
[0514] The device uses a microphone to capture the user's voice, and in the next step, this voice data is passed to a voice recognition system. The input is the user's voice, and the output is the captured voice data.
[0515] Step 3:
[0516] The voice data captured by the device is input into a voice recognition system (for example, Google Speech-to-Text or Amazon Transcribe) and converted into text data. In this conversion, the voice data "Go to City Hall" is converted into text data "Go to City Hall." The input is the captured voice data, and the output is the converted text data.
[0517] Step 4:
[0518] The terminal sends the converted text data to the server. HTTP or WebSocket is used for communication. The input is the converted text data, and the output is data sent to the server.
[0519] Step 5:
[0520] The server passes the received text data to the generation AI for analysis. The generation AI (for example, OpenAI GPT) analyzes the text data to obtain destination information. For example, it analyzes the text data "Go to city hall" to obtain the location information (latitude, longitude, etc.) of city hall. The input is text data, and the output is the analyzed destination information.
[0521] Step 6:
[0522] The server uses emotion analysis to recognize the user's emotional state from the voice data. The emotion analysis engine analyzes acoustic features (tone of voice, pitch, etc.) and estimates the user's emotional state, such as nervousness or calmness. The input is the voice data, and the output is the recognized emotional state.
[0523] Step 7:
[0524] The server's generation AI adjusts flight parameters (speed, altitude, etc.) based on the recognized emotional state. For example, if the user is nervous, the flight speed will be slowed down. A specific flight path is generated based on these flight parameters, and flight commands are created. The input is the emotional state and destination information, and the output is the adjusted flight parameters and specific flight commands.
[0525] Step 8:
[0526] The flight commands generated by the server are sent to the drone's control platform via a communication method. High-speed communication protocols such as MQTT and TCP / IP are used for communication. The input is the flight command, and the output is the command sent to the drone's control platform.
[0527] Step 9:
[0528] The drone will fly automatically based on flight commands received through the control board. The drone will follow the specified flight path to its destination and automatically perform operations such as landing and hovering. The input is the flight command, and the output is the drone's actual flight behavior.
[0529] (Application example 2)
[0530] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0531] In conventional food delivery systems, it is difficult to control drones based on voice commands, and it is not possible to set optimal flight parameters according to the user's emotional state, which can lead to anxiety and stress during delivery.In addition, it is difficult to adjust flight parameters in real time, leaving issues in terms of safety and efficiency.
[0532] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0533] In this invention, the server includes a voice input means for capturing a user's voice input, a voice recognition means for converting the captured voice into text data, an emotion recognition means for recognizing the user's emotional state using an emotion engine, a parameter adjustment means for setting drone flight parameters according to the user's emotional state, a command generation means for generating drone flight commands using generative artificial intelligence, a communication means for transmitting the generated flight commands to a drone control platform, and a drone control means for controlling the drone based on the transmitted flight commands. This allows the user to simply issue voice commands to set optimal drone flight parameters according to their emotional state, enabling safe and efficient food delivery.
[0534] "Audio input means" is a device or function for capturing a user's voice.
[0535] "Speech recognition means" refers to technology or devices that analyze captured voice data and convert it into text data.
[0536] "Emotion recognition means" refers to technology or devices for recognizing a user's emotional state from voice data or text data.
[0537] "Parameter adjustment means" refers to a function for setting and adjusting the drone's flight parameters (e.g., speed, altitude, etc.) based on the recognized emotional state of the user.
[0538] A "command generation means" is a device or function that generates flight commands for the drone based on the user's voice input and emotional state.
[0539] "Communication means" refers to the method or technology used to transmit the generated flight commands to the drone's control platform.
[0540] "Drone control means" means a device or function for controlling and managing a drone based on received flight commands.
[0541] "Generative AI" is an AI technology that uses large amounts of data to analyze user instructions and generate appropriate commands.
[0542] The "emotion engine" is a technology that detects the user's emotional state from voice and text data and uses it to adjust flight parameters.
[0543] This invention is a drone control system that can capture the user's voice instructions and combine them with an emotion engine to set optimal flight parameters according to the user's emotional state.
[0544] The server first has a voice input means for capturing voice input from the user, which is a voice input device such as a microphone, and captures voice instructions given by the user.
[0545] Once the audio is captured, a speech recognition means processes it and converts it into text data, for example, using the speech_recognition library, which generates text data from the user's voice commands.
[0546] Next, the emotion recognition means recognizes the user's emotional state from the voice. The emotion recognition means uses an emotion engine to extract emotions from the voice. An example of this emotion engine is the EmotionRecognizer module.
[0547] Based on the recognized emotional state, the parameter adjustment means sets and adjusts the optimal drone flight parameters, including the drone's speed and altitude. For example, if the user is nervous, the flight speed may be slowed down.
[0548] The command generation means then uses a generative AI model to generate specific flight commands. This generative AI model embodies the drone's flight route and operations based on the generated text data and the user's emotional state. For example, it generates detailed commands for takeoff, flight path, altitude, hovering, landing, etc.
[0549] The generated flight commands are sent to the drone's control platform via a communication method that uses a high-speed, highly reliable protocol to ensure real-time communication, such as Wi-Fi or 4G / 5G communication.
[0550] Based on the transmitted flight commands, the drone control means controls the drone, allowing the drone to travel to the destination along the specified flight route, achieving safe and efficient flight according to the user's emotional state.
[0551] Specific examples
[0552] Example prompt sentence:
[0553] The user gives a voice command such as "Deliver pizza to my house."
[0554] In the processing flow, voice input is captured and converted into text data such as "Deliver pizza to my house" by the voice recognition means. Then, the emotion recognition means recognizes the user's emotional state, and the parameter adjustment means sets flight parameters. Finally, the generated flight command is sent to the drone, which then delivers the pizza to the house.
[0555] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0556] Step 1: Capturing Audio Input
[0557] The server uses a voice input means (microphone) to capture the user's voice instruction. This input is, for example, the user's voice data saying, "Deliver a pizza to my house." The captured voice data is saved for processing in a later step.
[0558] Step 2: Voice Recognition
[0559] The server uses a speech recognition tool (the speech_recognition library) to convert the captured voice data into text data, which contains specific instructions such as "Deliver a pizza to my house." The converted text data is used in the next processing step.
[0560] Step 3: Emotion Recognition
[0561] The server uses an emotion recognition module (EmotionRecognizer module) to recognize the user's emotional state based on the converted text data. It analyzes the features extracted from the voice data and identifies the user's emotional state, such as "joy." The recognized emotional state is then used to set flight parameters.
[0562] Step 4: Obtaining destination information
[0563] The server uses the command generation means to obtain destination information from the user's voice input. Specifically, it analyzes the text data "Deliver pizza to my house" and obtains the coordinates (latitude and longitude) of "home" using a geographic information acquisition library (such as geopy). The obtained coordinate information is used to generate a flight route.
[0564] Step 5: Adjusting Flight Parameters
[0565] The server uses a parameter adjustment means to set the drone's flight parameters according to the recognized emotional state. For example, if the user is "nervous," the server sets the drone's flight speed slower to increase safety. This adjustment allows the drone to fly in a way that takes the user's emotions into consideration.
[0566] Step 6: Generate Flight Commands
[0567] The server generates specific flight commands using a command generation means and a generative AI model. The generative AI model outputs detailed flight commands, such as takeoff, route, altitude setting, hovering, and landing, based on the user's destination information and emotional state. These flight commands serve as guidelines for the drone's operation.
[0568] Step 7: Sending Flight Commands
[0569] The server then transmits the generated flight commands to the drone's control platform via a communication method using a high-speed, highly reliable protocol (such as Wi-Fi or 4G / 5G communication), enabling accurate command transmission in real time.
[0570] Step 8: Controlling the drone
[0571] The drone operates based on the flight commands received. The drone control means flies the pizza to the customer's home according to the specified flight route and parameters, realizing safe and efficient food delivery that takes into consideration the customer's feelings.
[0572] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0573] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0574] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0575] [Third embodiment]
[0576] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0577] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0578] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0579] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0580] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0581] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0582] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0583] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0584] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0585] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0586] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0587] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0588] To implement the present invention, the following system is required: This system is designed to allow a user to freely control a drone through voice input.
[0589] First, the user uses voice input to give the drone instructions such as "Go to City Hall." This voice input is expected to be implemented on a smartphone or dedicated device.
[0590] Next, the device captures the user's voice and converts the voice data into text data using a speech recognition means. At this stage, a speech recognition engine is used to convert the voice into text, such as "Go to City Hall."
[0591] The converted text data is sent from the device to a server and processed by a command generation means implemented on the server. Specifically, the generative artificial intelligence (generative AI) on the server analyzes the text data and understands the user's instructions. For example, the instruction "Go to city hall" is analyzed, and the location information of city hall is obtained.
[0592] Next, the AI generates flight commands that include a specific flight path to City Hall. These flight commands specify detailed actions from takeoff to the destination, such as "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0593] The generated flight commands are sent to the drone's control platform via the server's communications means using a high-speed communication protocol, ensuring real-time communication.
[0594] Finally, based on the flight commands received by the drone's control board, the drone will fly along the designated path to City Hall, automatically performing a series of operations including takeoff, heading, altitude adjustment, hovering, and landing.
[0595] This allows users to easily control a drone to its destination using only voice input. This system does not require the piloting skills or specialized knowledge that was previously required, and is easy for even general users to use.
[0596] For example, if a user says, "To the supermarket parking lot," the voice input is captured and converted into text using a speech recognition system. The AI then obtains the supermarket's location information, generates a specific flight path, and sends it to the drone, which then automatically flies to the supermarket.
[0597] The processing flow will be explained below.
[0598] Step 1:
[0599] The user uses voice input to give instructions to the drone, such as "Go to City Hall."
[0600] Step 2:
[0601] The device captures the user's voice with a microphone.
[0602] Step 3:
[0603] The device sends the captured voice data to a speech recognition engine, which analyzes the voice data and converts it into corresponding text data ("Go to City Hall").
[0604] Step 4:
[0605] The terminal transmits the converted text data to the server.
[0606] Step 5:
[0607] The server passes the received text data to the generation AI, which analyzes the text data and understands the user's instructions.
[0608] Step 6:
[0609] The server uses a generation AI to obtain the location information of the destination (in this case, city hall) based on the user's instructions.
[0610] Step 7:
[0611] Based on the location information of the destination acquired by the server, the generation AI generates a specific flight path.
[0612] Example: "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0613] Step 8:
[0614] The server transmits the generated flight commands to the drone's control board using a communication means.
[0615] Step 9:
[0616] The drone's control board analyzes the flight commands it receives and converts them into specific operational instructions.
[0617] Step 10:
[0618] The drone's control board controls the drone's motors and propellers, allowing the drone to operate according to flight commands.
[0619] Step 11:
[0620] The drone's control board executes a series of operations from takeoff to the destination, finally reaching and landing at the specified location.
[0621] This allows users to easily control the drone using only voice commands.
[0622] Example 1
[0623] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0624] Conventional drone control systems require specialized piloting skills and complex operations, making them difficult for average users to use. While voice-input systems existed, they faced problems with voice recognition accuracy and real-time response, making it difficult to generate efficient flight commands. Furthermore, communication delays and poor reliability meant that safe and accurate flight was not guaranteed.
[0625] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0626] In this invention, the server includes a voice input unit that captures a user's voice input, a voice recognition unit that converts the voice into text data, an analysis unit that analyzes the voice using generative artificial intelligence to acquire destination information, a command generation unit that generates drone flight commands based on the destination information, a communication unit, and a drone control unit. This allows general users to easily pilot drones to their destinations using only voice input, eliminating the need for piloting skills or specialized knowledge that was previously required. Furthermore, the use of a high-speed communication protocol enables safe and accurate flight in real time.
[0627] "Voice input means" refers to an apparatus or device that a user uses to input voice, and includes smartphones and dedicated devices.
[0628] "Speech recognition means" refers to a technology or function that converts captured speech into text data, such as a speech recognition engine.
[0629] "Analysis means" refers to the technology or function for analyzing the converted text data and obtaining destination information based on the user's instructions, and includes generative artificial intelligence.
[0630] The "command generation means" refers to a technology or function that generates a drone flight command including a specific flight path based on the destination information obtained by the analysis means.
[0631] "Communication means" refers to the technology and functions for transmitting the generated flight commands to the drone's control platform, and is a means that uses high-speed communication protocols, etc.
[0632] "Drone control means" refers to the technology and functions that control drones based on flight commands sent via communication means and perform operations such as takeoff, flight, hovering, and landing.
[0633] "Generative AI" refers to an AI technology that analyzes text data and understands the user's instructions, and generative AI models are examples of this type of AI.
[0634] "High-speed communication protocols" refer to communication protocols and standards for high-speed, reliable data communication, and examples of such protocols include TCP / IP and MQTT.
[0635] The present invention relates to a system that allows a user to freely control a drone using voice input. The system includes a voice input unit, a voice recognition unit, an analysis unit, a command generation unit, a communication unit, and a drone control unit.
[0636] First, the user issues commands to the drone using a voice input method, which is expected to be implemented on a smartphone or dedicated device. For example, the user can issue a command such as "Go to City Hall."
[0637] Next, the device captures the user's voice and converts the voice data into text using a speech recognition engine such as Google Cloud Speech-to-Text, which converts the voice into text such as "Go to City Hall."
[0638] The converted text data is sent from the device to a server. The server analyzes the text data using an analytical method that uses generative artificial intelligence (generative AI) to understand the user's instructions. For example, OpenAI GPT-3 is used as a generative AI model. This analyzes the instruction "Go to city hall" and obtains the location information of city hall.
[0639] Next, the server generates a flight command including a specific flight path based on the destination information obtained by the analysis means, such as "take off -> move north 100 meters -> move east 200 meters -> set altitude to 50 meters -> hover above City Hall -> land."
[0640] The generated flight commands are sent to the drone's control platform via the server's communication means using a high-speed communication protocol (e.g., TCP / IP, MQTT). This communication means uses a high-speed, highly reliable protocol to ensure real-time performance.
[0641] Finally, based on the flight commands received by the drone's control board, the drone will fly along the designated path to City Hall, automatically performing a series of actions including takeoff, changing direction, adjusting altitude, hovering, and landing.
[0642] This allows users to easily control a drone to its destination using only voice input. This system does not require the piloting skills or specialized knowledge that was previously required, and is easy for even general users to use.
[0643] For example, if a user says "to the supermarket parking lot," the voice is captured in the same steps and converted into text by speech recognition. The AI then obtains the supermarket's location information, generates a specific flight path, and sends it to the drone, which then flies to the supermarket automatically.
[0644] Example prompt sentence:
[0645] Generate a specific flight route for the voice command "Go to City Hall." Example: "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0646] or
[0647] Provide example text that generates a flight route when the user requests "to the supermarket parking lot."
[0648] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0649] Step 1:
[0650] The user issues commands to the drone using voice input. Specifically, the user picks up their smartphone, launches the voice input application, presses the record button on the screen, and inputs the command by voice, such as "Go to City Hall." This input is captured as voice data.
[0651] Output example: Voice data (voice saying "Go to City Hall")
[0652] Step 2:
[0653] The device receives the voice data captured by the voice input means and converts this voice data into text data using a voice recognition engine (e.g., Google Cloud Speech-to-Text). The voice data is input, and the text data "Go to City Hall" is output.
[0654] Output example: Text data (the text "Go to City Hall")
[0655] Step 3:
[0656] The terminal sends the converted text data to the server using an HTTP POST request, which takes text data as input and sends the data to the server as a result.
[0657] Output example: Text data sent to the server
[0658] Step 4:
[0659] The server uses a generative AI (e.g., OpenAI GPT-3) to analyze the received text data. The generative AI analyzes the text data and understands the instruction, "Go to city hall." Specifically, the location information of city hall is obtained based on the analyzed instruction. Here, the generative AI receives the prompt text, "Go to city hall," as input and outputs location information data.
[0660] Output example: Destination location data
[0661] Step 5:
[0662] The server generates flight commands including a specific flight path based on the location information obtained by the analysis means. The generative AI model receives the location information as input and generates a detailed flight path (route from takeoff to destination). For example, this command includes the steps: "Take off -> Move north 100 meters -> Move east 200 meters -> Set altitude to 50 meters -> Hover above City Hall -> Land."
[0663] Output example: Flight command data
[0664] Step 6:
[0665] The server sends the generated flight commands to the drone's control platform using a high-speed communication protocol (e.g., TCP / IP, MQTT). This allows the server to receive flight command data as input and output it by sending it to the drone's control platform.
[0666] Example output: Flight commands sent to the drone control board
[0667] Step 7:
[0668] The drone control board controls the drone based on the flight commands it receives. Specifically, the drone automatically takes off and follows the specified path (100 meters north -> 200 meters east -> set altitude to 50 meters), hovering over City Hall before landing. It receives flight commands as input and controls the drone accordingly.
[0669] Example output: A drone reaching City Hall
[0670] summary
[0671] Through the above processing steps, the user can control the drone to the destination using only voice input. The input, data processing, output and specific operations at each step are explained in detail.
[0672] (Application example 1)
[0673] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0674] Conventional drone control systems allow users to control drones with voice input without any specific technical knowledge, but they have limitations in terms of efficiently transporting cargo to various locations within a logistics center. Furthermore, obtaining location information in real time and generating appropriate flight routes requires complex operations and time. Furthermore, there was a need for a system that could respond to the diverse instructions within a logistics center.
[0675] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0676] In this invention, the server includes a voice input means for capturing user voice input, a voice recognition means for converting the voice captured by the voice input means into text data, a command generation means for generating drone flight commands using generative artificial intelligence based on the text data converted by the voice recognition means, a communication means for transmitting the flight commands generated by the command generation means to a drone control platform, a drone control means for controlling the drone based on the flight commands transmitted by the communication means, a logistics management means for analyzing instructions and generating a flight path for transporting cargo at a logistics center, and a location information acquisition means for identifying the location and destination of the cargo based on the user's instructions and generating a flight path to the destination. This enables efficient cargo transportation within a logistics center using only voice input.
[0677] 1. "Voice input means" means a device for capturing a user's voice instructions, and includes a voice capture device such as a mobile terminal such as a smartphone or tablet.
[0678] 2. "Speech recognition means" means a technology for converting captured voice data into text data, and is a device that converts voice into text using a speech recognition API.
[0679] 3. "Command generation means" means a device that analyzes the converted text data using artificial intelligence to generate flight commands that instruct the drone's flight path and operations.
[0680] 4. "Communication means" refers to the devices and protocols used to transmit generated flight commands to the drone's control platform, enabling high-speed communication.
[0681] 5. "Drone control means" means a device for operating a drone based on received flight commands, and is a control platform that executes flight paths, adjusts altitude, etc.
[0682] 6. "Logistics management means" refers to a device that analyzes instructions and generates flight paths when transporting cargo using drones within a logistics center, and is a system that supports efficient cargo transportation.
[0683] 7. "Location information acquisition means" means a device for identifying the location and destination of a package based on user instructions, and a means for providing the location data necessary for generating a flight path.
[0684] The system for implementing this invention automates a series of processes from voice input to drone transportation. This system is designed to allow users to freely control drones through voice input.
[0685] First, the user uses a voice input means to give instructions to the drone, such as "Transport the package from shelf A2 to dock C." This voice input means is expected to be implemented on a mobile device such as a smartphone or tablet. A speech recognition API such as the Google Cloud Speech-to-Text API or the Apple Speech Framework is installed on the smartphone or tablet.
[0686] Next, the terminal captures the user's voice and converts the voice data into text data using a voice recognition means, so the voice is converted into text such as "Please carry the package from shelf A2 to dock C."
[0687] The converted text data is sent from the device to a server, where it is processed by a command generation means implemented on the server. Specifically, a generative artificial intelligence (e.g., a generative model such as GPT-3) on the server analyzes the text data and understands the user's instructions. Based on the user's instructions, a location information acquisition means identifies the location and destination of the package and generates a specific flight path using the Google Maps API.
[0688] The generation AI generates flight commands on the server, including the optimal flight path for cargo transportation. This flight command specifies detailed actions from takeoff to the destination. For example, it may include a path such as "Takeoff -> Get cargo from shelf A2 -> Proceed 100 meters north -> Drop cargo at dock C -> Hover -> Land."
[0689] The generated flight commands are sent to the drone's control platform via the server's communications means using a high-speed communications protocol. Real-time communication is ensured by using a high-speed and highly reliable protocol. Specifically, communications means such as Bluetooth, Wi-Fi, or LTE are considered.
[0690] Finally, based on the flight commands received by the drone's control board, the drone will carry out the cargo delivery along the designated path, automatically performing a series of actions including takeoff, heading, altitude adjustment, hovering, and landing.
[0691] This allows users to easily transport cargo within a logistics center using only voice input. The system does not require the piloting skills or specialized knowledge that was previously required, and can be easily used by general users.
[0692] For example, if a user says, "Please transport the package from shelf B3 to the shipping area," the voice input is captured and converted into text using a speech recognition system. The AI then obtains the location information of the shipping area, generates a specific flight path, and sends it to the drone. Finally, the drone automatically transports the package from shelf B3 to the shipping area.
[0693] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0694] Step 1:
[0695] A user uses a smartphone or tablet to give voice commands to the drone, for example, "Transport the package from shelf A2 to dock C." The input is the user's voice command, which is captured by the smartphone or tablet's microphone. The output is the voice data.
[0696] Step 2:
[0697] A speech recognition API (such as Google Cloud Speech-to-Text or Apple Speech Framework) installed on the device converts the captured voice data into text data. The input is voice data, which is converted into a string using the API. The output is text data such as "Carry the package from shelf A2 to dock C."
[0698] Step 3:
[0699] The converted text data is sent from the terminal to the server. The input is the text data, which is sent to the server via a high-speed communication protocol (Wi-Fi, LTE, etc.). The output is the text data received by the server.
[0700] Step 4:
[0701] The server analyzes the received text data using generative artificial intelligence (generative models such as GPT-3). The input is text data, and the generative AI understands its content and interprets the user's instructions. The output is the analyzed instructions.
[0702] Step 5:
[0703] The server uses a location acquisition method (such as Google Maps API) to determine the location and destination of the package. The input is the parsed instructions, and the API is used to obtain specific location data. The output is the package location and the destination location.
[0704] Step 6:
[0705] The server generates a specific flight path based on the location information, which includes a detailed route from takeoff to the destination. The input is the location information of the baggage and the location information of the destination, and the output is a specific flight command.
[0706] Step 7:
[0707] The server sends the generated flight commands to the drone's control board using a high-speed communication protocol (Wi-Fi, Bluetooth, LTE, etc.). The input is the flight command, which is transmitted to the drone via the communication means. The output is the flight command received by the drone.
[0708] Step 8:
[0709] The drone's control board automatically controls the drone based on the flight commands it receives. The input is the flight commands, which control altitude adjustment, direction of movement, speed, etc. The output is the drone's operation. Specifically, the drone retrieves the package from shelf A2 and transports it to dock C.
[0710] Through the above processing steps, users can efficiently transport drones within a logistics center using only voice input.
[0711] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0712] The present invention provides a system that can set optimal flight parameters according to the user's emotional state by combining an emotion engine with a system that automatically controls a drone based on the user's voice instructions.
[0713] First, the user uses voice input to give instructions to the drone, such as "Go to City Hall."
[0714] Next, the terminal captures the user's voice with a microphone. This captured voice data is converted into text data using a voice recognition means. The voice recognition means analyzes the voice data and generates text data such as "Go to City Hall."
[0715] The converted text data is sent from the device to the server. The command generation means implemented on the server passes the text data to the generation artificial intelligence (generation AI) and analyzes the user's instructions. The generation AI obtains the location information of the destination based on the instructions.
[0716] The server also has an emotion engine built in. This emotion engine recognizes the user's emotional state from the voice input. For example, if the user is nervous, it will recognize that as the corresponding emotional state.
[0717] The generative AI then adjusts flight parameters (speed and altitude) via an emotion engine based on the user's perceived emotional state. For example, if the user is nervous, the flight speed will be slowed down.
[0718] The AI then generates flight commands including a specific flight path, such as "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0719] The generated flight commands are sent to the drone's control platform via the server's communications means using a high-speed communication protocol, ensuring real-time communication.
[0720] Finally, the drone's control board controls the drone based on the flight commands it receives. The drone follows the specified flight path to its destination and automatically performs a series of operations, such as hovering and landing.
[0721] For example, if a user says "to the supermarket parking lot," the voice input is captured and converted into text using speech recognition. The generative AI then recognizes the user's emotional state through its emotion engine and retrieves the destination information. The flight parameters are then adjusted based on the emotional state, a specific flight path is generated, and transmitted to the drone, and the drone then flies automatically to the supermarket.
[0722] This system allows users to fly safely and securely based on their emotional state, and allows them to easily control the drone using only voice commands.
[0723] The processing flow will be explained below.
[0724] Step 1:
[0725] The user can use voice input to give instructions to the drone, for example, "Go to City Hall."
[0726] Step 2:
[0727] The device captures the user's voice with a microphone.
[0728] Step 3:
[0729] The device sends the captured voice data to a speech recognition engine, which analyzes the voice data and converts it into corresponding text data ("Go to City Hall").
[0730] Step 4:
[0731] The terminal transmits the converted text data to the server.
[0732] Step 5:
[0733] The server passes the received text data to the generation AI, which analyzes the text data and understands the instruction, "Go to city hall."
[0734] Step 6:
[0735] The server then passes the text data to the emotion engine, which analyzes and evaluates the user's emotional state from the voice data.
[0736] Step 7:
[0737] The server uses the generated AI to obtain the location information of the destination (city hall) based on the user's instructions.
[0738] Step 8:
[0739] The server uses generated AI to adjust flight parameters (speed, altitude, etc.) based on the acquired location information and the analysis results of the emotion engine.
[0740] Step 9:
[0741] The server instructs the AI to generate a specific flight path including flight parameters, causing it to generate flight commands.
[0742] Example: "Take off -> fly 100 meters north -> fly 200 meters east -> set altitude to 50 meters (or lower if user is nervous) -> hover over City Hall -> land."
[0743] Step 10:
[0744] The server sends flight commands to the drone's control board using a communication means.
[0745] Step 11:
[0746] The drone's control board analyzes the flight commands it receives and converts them into specific operational instructions.
[0747] Step 12:
[0748] The drone's control board controls the drone's motors and propellers, making the drone operate according to flight commands.
[0749] Step 13:
[0750] The drone's control board executes a series of operations from takeoff to the destination, finally reaching and landing at the specified location.
[0751] This allows users to easily control the drone using only voice commands, and also ensures safe and optimal flight based on the user's emotional state.
[0752] Example 2
[0753] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0754] Conventional drone control systems have mechanisms to convert user instructions into text data and generate flight commands based on destination information, but they lack the functionality to optimize flight parameters based on the user's emotional state. As a result, if the user becomes nervous or excited, the drone's flight cannot adapt to the user's emotional state, making it difficult to achieve a safe and secure flight.
[0755] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0756] In this invention, the server includes a voice input means for capturing a user's voice input, a voice recognition means for converting the voice captured by the voice input means into text data, an emotion analysis means for recognizing the user's emotional state from the voice data, a means for adjusting flight parameters based on the emotional state recognized by the emotion analysis means, a means for generating flight commands including the adjusted flight parameters, a communication means for transmitting the flight commands generated by the command generation means to a drone control platform, and a drone control means for controlling the drone based on the flight commands transmitted by the communication means. This makes it possible to set optimal flight parameters based on the user's emotional state, thereby achieving safe and secure flight.
[0757] "Audio input means" refers to a device or system for capturing a user's voice.
[0758] A "speech recognition means" is a device or algorithm that converts captured speech into text data.
[0759] A "command generation means" is a system or device for generating drone flight commands using artificial intelligence based on text data.
[0760] "Emotion analysis means" refers to a system or algorithm for recognizing a user's emotional state from their voice data.
[0761] The "flight parameter adjustment means" is a system or device for adjusting flight parameters based on the emotional state recognized by the emotion analysis means.
[0762] "Communication means" refers to a device or system for transmitting the generated flight commands to the drone's control board.
[0763] A "drone control means" is a system or device for controlling a drone based on received flight commands.
[0764] MODE FOR CARRYING OUT THE INVENTION
[0765] The present invention provides a system that combines an emotion analysis function with a system that automatically controls a drone based on the user's voice instructions, and can set optimal flight parameters according to the user's emotional state.
[0766] First, the user issues commands to the drone using a voice input means. For example, the user might issue a voice command such as "Go to City Hall." Next, the device captures the user's voice with a microphone. This captured voice data is converted into text data using a voice recognition means. Specifically, a voice recognition system such as Google Speech-to-Text or Amazon Transcribe is used. The voice recognition means analyzes the voice data and generates the text data "Go to City Hall."
[0767] The converted text data is sent from the device to the server. Standard communication protocols such as HTTP and WebSocket are used for communication. The command generation means implemented on the server passes the text data to the generation AI, which analyzes the user's instructions. OpenAI GPT and other programs can be used as the generation AI. The generation AI obtains the location information of the destination based on the instructions.
[0768] Additionally, the server is equipped with an emotion analysis engine that recognizes the user's emotional state from the voice input. For example, if the user is nervous, it will be recognized as such. The emotion engine achieves this by using an algorithm that analyzes acoustic features (such as tone and pitch of voice).
[0769] Next, the generation AI adjusts flight parameters (speed and altitude) through the emotion engine based on the recognized emotional state. If the user is nervous, adjustments will be made, such as slowing down the flight speed. Based on the adjusted flight parameters, the generation AI generates a specific flight path and creates flight commands. For example, a series of commands might be "Take off -> Move north 100 meters -> Move right (east) 200 meters -> Set altitude to 50 meters -> Hover above City Hall -> Land."
[0770] The generated flight commands are sent to the drone's control platform via the server's communication means using a high-speed communication protocol (such as MQTT or TCP / IP). The drone's control platform ultimately controls the drone based on the received flight commands. The drone follows the specified flight path to its destination and automatically performs operations such as hovering and landing.
[0771] As a specific example, the operation when the user instructs "to the supermarket parking lot" will be shown.
[0772] 1. The user says, "To the supermarket parking lot."
[0773] 2. The device captures the audio with its microphone.
[0774] 3. The speech recognition system converts the speech data into text data.
[0775] 4. The device sends the converted text data to the server.
[0776] 5. The server's generated AI analyzes the text data and obtains the destination's location information.
[0777] 6. The emotion analysis means built into the server recognizes the user's emotional state from the voice data. For example, if the user is excited, it will be recognized as that emotional state.
[0778] 7. The AI generator adjusts flight parameters based on the user's emotional state. For example, if the user is excited, the AI will increase flight speed.
[0779] 8. The generation AI generates a specific flight path and creates flight commands.
[0780] 9. The server sends flight commands to the drone's control board via communication means.
[0781] 10. The drone's control board controls the drone based on the commands, and it flies automatically to the supermarket parking lot.
[0782] Examples of prompt sentences are as follows:
[0783] "Go to City Hall."
[0784] "Fly to the park fountain"
[0785] "Come back to my garden."
[0786] This system allows users to fly safely and securely based on their emotional state, and allows them to easily control the drone using only voice commands.
[0787] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0788] Step 1:
[0789] The user issues a voice command using the voice input means, for example, "Go to city hall." The input here is the user's voice, and the output is the voice data.
[0790] Step 2:
[0791] The device uses a microphone to capture the user's voice, and in the next step, this voice data is passed to a voice recognition system. The input is the user's voice, and the output is the captured voice data.
[0792] Step 3:
[0793] The voice data captured by the device is input into a voice recognition system (for example, Google Speech-to-Text or Amazon Transcribe) and converted into text data. In this conversion, the voice data "Go to City Hall" is converted into text data "Go to City Hall." The input is the captured voice data, and the output is the converted text data.
[0794] Step 4:
[0795] The terminal sends the converted text data to the server. HTTP or WebSocket is used for communication. The input is the converted text data, and the output is data sent to the server.
[0796] Step 5:
[0797] The server passes the received text data to the generation AI for analysis. The generation AI (for example, OpenAI GPT) analyzes the text data to obtain destination information. For example, it analyzes the text data "Go to city hall" to obtain the location information (latitude, longitude, etc.) of city hall. The input is text data, and the output is the analyzed destination information.
[0798] Step 6:
[0799] The server uses emotion analysis to recognize the user's emotional state from the voice data. The emotion analysis engine analyzes acoustic features (tone of voice, pitch, etc.) and estimates the user's emotional state, such as nervousness or calmness. The input is the voice data, and the output is the recognized emotional state.
[0800] Step 7:
[0801] The server's generation AI adjusts flight parameters (speed, altitude, etc.) based on the recognized emotional state. For example, if the user is nervous, the flight speed will be slowed down. A specific flight path is generated based on these flight parameters, and flight commands are created. The input is the emotional state and destination information, and the output is the adjusted flight parameters and specific flight commands.
[0802] Step 8:
[0803] The flight commands generated by the server are sent to the drone's control platform via a communication method. High-speed communication protocols such as MQTT and TCP / IP are used for communication. The input is the flight command, and the output is the command sent to the drone's control platform.
[0804] Step 9:
[0805] The drone will fly automatically based on flight commands received through the control board. The drone will follow the specified flight path to its destination and automatically perform operations such as landing and hovering. The input is the flight command, and the output is the drone's actual flight behavior.
[0806] (Application example 2)
[0807] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0808] In conventional food delivery systems, it is difficult to control drones based on voice commands, and it is not possible to set optimal flight parameters according to the user's emotional state, which can lead to anxiety and stress during delivery.In addition, it is difficult to adjust flight parameters in real time, leaving issues in terms of safety and efficiency.
[0809] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0810] In this invention, the server includes a voice input means for capturing a user's voice input, a voice recognition means for converting the captured voice into text data, an emotion recognition means for recognizing the user's emotional state using an emotion engine, a parameter adjustment means for setting drone flight parameters according to the user's emotional state, a command generation means for generating drone flight commands using generative artificial intelligence, a communication means for transmitting the generated flight commands to a drone control platform, and a drone control means for controlling the drone based on the transmitted flight commands. This allows the user to simply issue voice commands to set optimal drone flight parameters according to their emotional state, enabling safe and efficient food delivery.
[0811] "Audio input means" is a device or function for capturing a user's voice.
[0812] "Speech recognition means" refers to technology or devices that analyze captured voice data and convert it into text data.
[0813] "Emotion recognition means" refers to technology or devices for recognizing a user's emotional state from voice data or text data.
[0814] "Parameter adjustment means" refers to a function for setting and adjusting the drone's flight parameters (e.g., speed, altitude, etc.) based on the recognized emotional state of the user.
[0815] A "command generation means" is a device or function that generates flight commands for the drone based on the user's voice input and emotional state.
[0816] "Communication means" refers to the method or technology used to transmit the generated flight commands to the drone's control platform.
[0817] "Drone control means" means a device or function for controlling and managing a drone based on received flight commands.
[0818] "Generative AI" is an AI technology that uses large amounts of data to analyze user instructions and generate appropriate commands.
[0819] The "emotion engine" is a technology that detects the user's emotional state from voice and text data and uses it to adjust flight parameters.
[0820] This invention is a drone control system that can capture the user's voice instructions and combine them with an emotion engine to set optimal flight parameters according to the user's emotional state.
[0821] The server first has a voice input means for capturing voice input from the user, which is a voice input device such as a microphone, and captures voice instructions given by the user.
[0822] Once the audio is captured, a speech recognition means processes it and converts it into text data, for example, using the speech_recognition library, which generates text data from the user's voice commands.
[0823] Next, the emotion recognition means recognizes the user's emotional state from the voice. The emotion recognition means uses an emotion engine to extract emotions from the voice. An example of this emotion engine is the EmotionRecognizer module.
[0824] Based on the recognized emotional state, the parameter adjustment means sets and adjusts the optimal drone flight parameters, including the drone's speed and altitude. For example, if the user is nervous, the flight speed may be slowed down.
[0825] The command generation means then uses a generative AI model to generate specific flight commands. This generative AI model embodies the drone's flight route and operations based on the generated text data and the user's emotional state. For example, it generates detailed commands for takeoff, flight path, altitude, hovering, landing, etc.
[0826] The generated flight commands are sent to the drone's control platform via a communication method that uses a high-speed, highly reliable protocol to ensure real-time communication, such as Wi-Fi or 4G / 5G communication.
[0827] Based on the transmitted flight commands, the drone control means controls the drone, allowing the drone to travel to the destination along the specified flight route, achieving safe and efficient flight according to the user's emotional state.
[0828] Specific examples
[0829] Example prompt sentence:
[0830] The user gives a voice command such as "Deliver pizza to my house."
[0831] In the processing flow, voice input is captured and converted into text data such as "Deliver pizza to my house" by the voice recognition means. Then, the emotion recognition means recognizes the user's emotional state, and the parameter adjustment means sets flight parameters. Finally, the generated flight command is sent to the drone, which then delivers the pizza to the house.
[0832] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0833] Step 1: Capturing Audio Input
[0834] The server uses a voice input means (microphone) to capture the user's voice instruction. This input is, for example, the user's voice data saying, "Deliver a pizza to my house." The captured voice data is saved for processing in a later step.
[0835] Step 2: Voice Recognition
[0836] The server uses a speech recognition tool (the speech_recognition library) to convert the captured voice data into text data, which contains specific instructions such as "Deliver a pizza to my house." The converted text data is used in the next processing step.
[0837] Step 3: Emotion Recognition
[0838] The server uses an emotion recognition module (EmotionRecognizer module) to recognize the user's emotional state based on the converted text data. It analyzes the features extracted from the voice data and identifies the user's emotional state, such as "joy." The recognized emotional state is then used to set flight parameters.
[0839] Step 4: Obtaining destination information
[0840] The server uses the command generation means to obtain destination information from the user's voice input. Specifically, it analyzes the text data "Deliver pizza to my house" and obtains the coordinates (latitude and longitude) of "home" using a geographic information acquisition library (such as geopy). The obtained coordinate information is used to generate a flight route.
[0841] Step 5: Adjusting Flight Parameters
[0842] The server uses a parameter adjustment means to set the drone's flight parameters according to the recognized emotional state. For example, if the user is "nervous," the server sets the drone's flight speed slower to increase safety. This adjustment allows the drone to fly in a way that takes the user's emotions into consideration.
[0843] Step 6: Generate Flight Commands
[0844] The server generates specific flight commands using a command generation means and a generative AI model. The generative AI model outputs detailed flight commands, such as takeoff, route, altitude setting, hovering, and landing, based on the user's destination information and emotional state. These flight commands serve as guidelines for the drone's operation.
[0845] Step 7: Sending Flight Commands
[0846] The server then transmits the generated flight commands to the drone's control platform via a communication method using a high-speed, highly reliable protocol (such as Wi-Fi or 4G / 5G communication), enabling accurate command transmission in real time.
[0847] Step 8: Controlling the drone
[0848] The drone operates based on the flight commands received. The drone control means flies the pizza to the customer's home according to the specified flight route and parameters, realizing safe and efficient food delivery that takes into consideration the customer's feelings.
[0849] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0850] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0851] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0852] [Fourth embodiment]
[0853] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0854] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0855] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0856] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0857] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0858] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0859] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0860] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0861] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0862] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0863] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0864] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0865] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0866] To implement the present invention, the following system is required: This system is designed to allow a user to freely control a drone through voice input.
[0867] First, the user uses voice input to give the drone instructions such as "Go to City Hall." This voice input is expected to be implemented on a smartphone or dedicated device.
[0868] Next, the device captures the user's voice and converts the voice data into text data using a speech recognition means. At this stage, a speech recognition engine is used to convert the voice into text, such as "Go to City Hall."
[0869] The converted text data is sent from the device to a server and processed by a command generation means implemented on the server. Specifically, the generative artificial intelligence (generative AI) on the server analyzes the text data and understands the user's instructions. For example, the instruction "Go to city hall" is analyzed, and the location information of city hall is obtained.
[0870] Next, the AI generates flight commands that include a specific flight path to City Hall. These flight commands specify detailed actions from takeoff to the destination, such as "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0871] The generated flight commands are sent to the drone's control platform via the server's communications means using a high-speed communication protocol, ensuring real-time communication.
[0872] Finally, based on the flight commands received by the drone's control board, the drone will fly along the designated path to City Hall, automatically performing a series of operations including takeoff, heading, altitude adjustment, hovering, and landing.
[0873] This allows users to easily control a drone to its destination using only voice input. This system does not require the piloting skills or specialized knowledge that was previously required, and is easy for even general users to use.
[0874] For example, if a user says, "To the supermarket parking lot," the voice input is captured and converted into text using a speech recognition system. The AI then obtains the supermarket's location information, generates a specific flight path, and sends it to the drone, which then automatically flies to the supermarket.
[0875] The processing flow will be explained below.
[0876] Step 1:
[0877] The user uses voice input to give instructions to the drone, such as "Go to City Hall."
[0878] Step 2:
[0879] The device captures the user's voice with a microphone.
[0880] Step 3:
[0881] The device sends the captured voice data to a speech recognition engine, which analyzes the voice data and converts it into corresponding text data ("Go to City Hall").
[0882] Step 4:
[0883] The terminal transmits the converted text data to the server.
[0884] Step 5:
[0885] The server passes the received text data to the generation AI, which analyzes the text data and understands the user's instructions.
[0886] Step 6:
[0887] The server uses a generation AI to obtain the location information of the destination (in this case, city hall) based on the user's instructions.
[0888] Step 7:
[0889] Based on the location information of the destination acquired by the server, the generation AI generates a specific flight path.
[0890] Example: "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0891] Step 8:
[0892] The server transmits the generated flight commands to the drone's control board using a communication means.
[0893] Step 9:
[0894] The drone's control board analyzes the flight commands it receives and converts them into specific operational instructions.
[0895] Step 10:
[0896] The drone's control board controls the drone's motors and propellers, allowing the drone to operate according to flight commands.
[0897] Step 11:
[0898] The drone's control board executes a series of operations from takeoff to the destination, finally reaching and landing at the specified location.
[0899] This allows users to easily control the drone using only voice commands.
[0900] Example 1
[0901] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0902] Conventional drone control systems require specialized piloting skills and complex operations, making them difficult for average users to use. While voice-input systems existed, they faced problems with voice recognition accuracy and real-time response, making it difficult to generate efficient flight commands. Furthermore, communication delays and poor reliability meant that safe and accurate flight was not guaranteed.
[0903] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0904] In this invention, the server includes a voice input unit that captures a user's voice input, a voice recognition unit that converts the voice into text data, an analysis unit that analyzes the voice using generative artificial intelligence to acquire destination information, a command generation unit that generates drone flight commands based on the destination information, a communication unit, and a drone control unit. This allows general users to easily pilot drones to their destinations using only voice input, eliminating the need for piloting skills or specialized knowledge that was previously required. Furthermore, the use of a high-speed communication protocol enables safe and accurate flight in real time.
[0905] "Voice input means" refers to an apparatus or device that a user uses to input voice, and includes smartphones and dedicated devices.
[0906] "Speech recognition means" refers to a technology or function that converts captured speech into text data, such as a speech recognition engine.
[0907] "Analysis means" refers to the technology or function for analyzing the converted text data and obtaining destination information based on the user's instructions, and includes generative artificial intelligence.
[0908] The "command generation means" refers to a technology or function that generates a drone flight command including a specific flight path based on the destination information obtained by the analysis means.
[0909] "Communication means" refers to the technology and functions for transmitting the generated flight commands to the drone's control platform, and is a means that uses high-speed communication protocols, etc.
[0910] "Drone control means" refers to the technology and functions that control drones based on flight commands sent via communication means and perform operations such as takeoff, flight, hovering, and landing.
[0911] "Generative AI" refers to an AI technology that analyzes text data and understands the user's instructions, and generative AI models are examples of this type of AI.
[0912] "High-speed communication protocols" refer to communication protocols and standards for high-speed, reliable data communication, and examples of such protocols include TCP / IP and MQTT.
[0913] The present invention relates to a system that allows a user to freely control a drone using voice input. The system includes a voice input unit, a voice recognition unit, an analysis unit, a command generation unit, a communication unit, and a drone control unit.
[0914] First, the user issues commands to the drone using a voice input method, which is expected to be implemented on a smartphone or dedicated device. For example, the user can issue a command such as "Go to City Hall."
[0915] Next, the device captures the user's voice and converts the voice data into text using a speech recognition engine such as Google Cloud Speech-to-Text, which converts the voice into text such as "Go to City Hall."
[0916] The converted text data is sent from the device to a server. The server analyzes the text data using an analytical method that uses generative artificial intelligence (generative AI) to understand the user's instructions. For example, OpenAI GPT-3 is used as a generative AI model. This analyzes the instruction "Go to city hall" and obtains the location information of city hall.
[0917] Next, the server generates a flight command including a specific flight path based on the destination information obtained by the analysis means, such as "take off -> move north 100 meters -> move east 200 meters -> set altitude to 50 meters -> hover above City Hall -> land."
[0918] The generated flight commands are sent to the drone's control platform via the server's communication means using a high-speed communication protocol (e.g., TCP / IP, MQTT). This communication means uses a high-speed, highly reliable protocol to ensure real-time performance.
[0919] Finally, based on the flight commands received by the drone's control board, the drone will fly along the designated path to City Hall, automatically performing a series of actions including takeoff, changing direction, adjusting altitude, hovering, and landing.
[0920] This allows users to easily control a drone to its destination using only voice input. This system does not require the piloting skills or specialized knowledge that was previously required, and is easy for even general users to use.
[0921] For example, if a user says "to the supermarket parking lot," the voice is captured in the same steps and converted into text by speech recognition. The AI then obtains the supermarket's location information, generates a specific flight path, and sends it to the drone, which then flies to the supermarket automatically.
[0922] Example prompt sentence:
[0923] Generate a specific flight route for the voice command "Go to City Hall." Example: "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0924] or
[0925] Provide example text that generates a flight route when the user requests "to the supermarket parking lot."
[0926] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0927] Step 1:
[0928] The user issues commands to the drone using voice input. Specifically, the user picks up their smartphone, launches the voice input application, presses the record button on the screen, and inputs the command by voice, such as "Go to City Hall." This input is captured as voice data.
[0929] Output example: Voice data (voice saying "Go to City Hall")
[0930] Step 2:
[0931] The device receives the voice data captured by the voice input means and converts this voice data into text data using a voice recognition engine (e.g., Google Cloud Speech-to-Text). The voice data is input, and the text data "Go to City Hall" is output.
[0932] Output example: Text data (the text "Go to City Hall")
[0933] Step 3:
[0934] The terminal sends the converted text data to the server using an HTTP POST request, which takes text data as input and sends the data to the server as a result.
[0935] Output example: Text data sent to the server
[0936] Step 4:
[0937] The server uses a generative AI (e.g., OpenAI GPT-3) to analyze the received text data. The generative AI analyzes the text data and understands the instruction, "Go to city hall." Specifically, the location information of city hall is obtained based on the analyzed instruction. Here, the generative AI receives the prompt text, "Go to city hall," as input and outputs location information data.
[0938] Output example: Destination location data
[0939] Step 5:
[0940] The server generates flight commands including a specific flight path based on the location information obtained by the analysis means. The generative AI model receives the location information as input and generates a detailed flight path (route from takeoff to destination). For example, this command includes the steps: "Take off -> Move north 100 meters -> Move east 200 meters -> Set altitude to 50 meters -> Hover above City Hall -> Land."
[0941] Output example: Flight command data
[0942] Step 6:
[0943] The server sends the generated flight commands to the drone's control platform using a high-speed communication protocol (e.g., TCP / IP, MQTT). This allows the server to receive flight command data as input and output it by sending it to the drone's control platform.
[0944] Example output: Flight commands sent to the drone control board
[0945] Step 7:
[0946] The drone control board controls the drone based on the flight commands it receives. Specifically, the drone automatically takes off and follows the specified path (100 meters north -> 200 meters east -> set altitude to 50 meters), hovering over City Hall before landing. It receives flight commands as input and controls the drone accordingly.
[0947] Example output: A drone reaching City Hall
[0948] summary
[0949] Through the above processing steps, the user can control the drone to the destination using only voice input. The input, data processing, output and specific operations at each step are explained in detail.
[0950] (Application example 1)
[0951] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0952] Conventional drone control systems allow users to control drones with voice input without any specific technical knowledge, but they have limitations in terms of efficiently transporting cargo to various locations within a logistics center. Furthermore, obtaining location information in real time and generating appropriate flight routes requires complex operations and time. Furthermore, there was a need for a system that could respond to the diverse instructions within a logistics center.
[0953] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0954] In this invention, the server includes a voice input means for capturing user voice input, a voice recognition means for converting the voice captured by the voice input means into text data, a command generation means for generating drone flight commands using generative artificial intelligence based on the text data converted by the voice recognition means, a communication means for transmitting the flight commands generated by the command generation means to a drone control platform, a drone control means for controlling the drone based on the flight commands transmitted by the communication means, a logistics management means for analyzing instructions and generating a flight path for transporting cargo at a logistics center, and a location information acquisition means for identifying the location and destination of the cargo based on the user's instructions and generating a flight path to the destination. This enables efficient cargo transportation within a logistics center using only voice input.
[0955] 1. "Voice input means" means a device for capturing a user's voice instructions, and includes a voice capture device such as a mobile terminal such as a smartphone or tablet.
[0956] 2. "Speech recognition means" means a technology for converting captured voice data into text data, and is a device that converts voice into text using a speech recognition API.
[0957] 3. "Command generation means" means a device that analyzes the converted text data using artificial intelligence to generate flight commands that instruct the drone's flight path and operations.
[0958] 4. "Communication means" refers to the devices and protocols used to transmit generated flight commands to the drone's control platform, enabling high-speed communication.
[0959] 5. "Drone control means" means a device for operating a drone based on received flight commands, and is a control platform that executes flight paths, adjusts altitude, etc.
[0960] 6. "Logistics management means" refers to a device that analyzes instructions and generates flight paths when transporting cargo using drones within a logistics center, and is a system that supports efficient cargo transportation.
[0961] 7. "Location information acquisition means" means a device for identifying the location and destination of a package based on user instructions, and a means for providing the location data necessary for generating a flight path.
[0962] The system for implementing this invention automates a series of processes from voice input to drone transportation. This system is designed to allow users to freely control drones through voice input.
[0963] First, the user uses a voice input means to give instructions to the drone, such as "Transport the package from shelf A2 to dock C." This voice input means is expected to be implemented on a mobile device such as a smartphone or tablet. A speech recognition API such as the Google Cloud Speech-to-Text API or the Apple Speech Framework is installed on the smartphone or tablet.
[0964] Next, the terminal captures the user's voice and converts the voice data into text data using a voice recognition means, so the voice is converted into text such as "Please carry the package from shelf A2 to dock C."
[0965] The converted text data is sent from the device to a server, where it is processed by a command generation means implemented on the server. Specifically, a generative artificial intelligence (e.g., a generative model such as GPT-3) on the server analyzes the text data and understands the user's instructions. Based on the user's instructions, a location information acquisition means identifies the location and destination of the package and generates a specific flight path using the Google Maps API.
[0966] The generation AI generates flight commands on the server, including the optimal flight path for cargo transportation. This flight command specifies detailed actions from takeoff to the destination. For example, it may include a path such as "Takeoff -> Get cargo from shelf A2 -> Proceed 100 meters north -> Drop cargo at dock C -> Hover -> Land."
[0967] The generated flight commands are sent to the drone's control platform via the server's communications means using a high-speed communications protocol. Real-time communication is ensured by using a high-speed and highly reliable protocol. Specifically, communications means such as Bluetooth, Wi-Fi, or LTE are considered.
[0968] Finally, based on the flight commands received by the drone's control board, the drone will carry out the cargo delivery along the designated path, automatically performing a series of actions including takeoff, heading, altitude adjustment, hovering, and landing.
[0969] This allows users to easily transport cargo within a logistics center using only voice input. The system does not require the piloting skills or specialized knowledge that was previously required, and can be easily used by general users.
[0970] For example, if a user says, "Please transport the package from shelf B3 to the shipping area," the voice input is captured and converted into text using a speech recognition system. The AI then obtains the location information of the shipping area, generates a specific flight path, and sends it to the drone. Finally, the drone automatically transports the package from shelf B3 to the shipping area.
[0971] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0972] Step 1:
[0973] A user uses a smartphone or tablet to give voice commands to the drone, for example, "Transport the package from shelf A2 to dock C." The input is the user's voice command, which is captured by the smartphone or tablet's microphone. The output is the voice data.
[0974] Step 2:
[0975] A speech recognition API (such as Google Cloud Speech-to-Text or Apple Speech Framework) installed on the device converts the captured voice data into text data. The input is voice data, which is converted into a string using the API. The output is text data such as "Carry the package from shelf A2 to dock C."
[0976] Step 3:
[0977] The converted text data is sent from the terminal to the server. The input is the text data, which is sent to the server via a high-speed communication protocol (Wi-Fi, LTE, etc.). The output is the text data received by the server.
[0978] Step 4:
[0979] The server analyzes the received text data using generative artificial intelligence (generative models such as GPT-3). The input is text data, and the generative AI understands its content and interprets the user's instructions. The output is the analyzed instructions.
[0980] Step 5:
[0981] The server uses a location acquisition method (such as Google Maps API) to determine the location and destination of the package. The input is the parsed instructions, and the API is used to obtain specific location data. The output is the package location and the destination location.
[0982] Step 6:
[0983] The server generates a specific flight path based on the location information, which includes a detailed route from takeoff to the destination. The input is the location information of the baggage and the location information of the destination, and the output is a specific flight command.
[0984] Step 7:
[0985] The server sends the generated flight commands to the drone's control board using a high-speed communication protocol (Wi-Fi, Bluetooth, LTE, etc.). The input is the flight command, which is transmitted to the drone via the communication means. The output is the flight command received by the drone.
[0986] Step 8:
[0987] The drone's control board automatically controls the drone based on the flight commands it receives. The input is the flight commands, which control altitude adjustment, direction of movement, speed, etc. The output is the drone's operation. Specifically, the drone retrieves the package from shelf A2 and transports it to dock C.
[0988] Through the above processing steps, users can efficiently transport drones within a logistics center using only voice input.
[0989] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0990] The present invention provides a system that can set optimal flight parameters according to the user's emotional state by combining an emotion engine with a system that automatically controls a drone based on the user's voice instructions.
[0991] First, the user uses voice input to give instructions to the drone, such as "Go to City Hall."
[0992] Next, the terminal captures the user's voice with a microphone. This captured voice data is converted into text data using a voice recognition means. The voice recognition means analyzes the voice data and generates text data such as "Go to City Hall."
[0993] The converted text data is sent from the device to the server. The command generation means implemented on the server passes the text data to the generation artificial intelligence (generation AI) and analyzes the user's instructions. The generation AI obtains the location information of the destination based on the instructions.
[0994] The server also has an emotion engine built in. This emotion engine recognizes the user's emotional state from the voice input. For example, if the user is nervous, it will recognize that as the corresponding emotional state.
[0995] The generative AI then adjusts flight parameters (speed and altitude) via an emotion engine based on the user's perceived emotional state. For example, if the user is nervous, the flight speed will be slowed down.
[0996] The AI then generates flight commands including a specific flight path, such as "Take off -> Head north 100 meters -> Head right (east) 200 meters -> Set altitude to 50 meters -> Hover over City Hall -> Land."
[0997] The generated flight commands are sent to the drone's control platform via the server's communications means using a high-speed communication protocol, ensuring real-time communication.
[0998] Finally, the drone's control board controls the drone based on the flight commands it receives. The drone follows the specified flight path to its destination and automatically performs a series of operations, such as hovering and landing.
[0999] For example, if a user says "to the supermarket parking lot," the voice input is captured and converted into text using speech recognition. The generative AI then recognizes the user's emotional state through its emotion engine and retrieves the destination information. The flight parameters are then adjusted based on the emotional state, a specific flight path is generated, and transmitted to the drone, and the drone then flies automatically to the supermarket.
[1000] This system allows users to fly safely and securely based on their emotional state, and allows them to easily control the drone using only voice commands.
[1001] The processing flow will be explained below.
[1002] Step 1:
[1003] The user can use voice input to give instructions to the drone, for example, "Go to City Hall."
[1004] Step 2:
[1005] The device captures the user's voice with a microphone.
[1006] Step 3:
[1007] The device sends the captured voice data to a speech recognition engine, which analyzes the voice data and converts it into corresponding text data ("Go to City Hall").
[1008] Step 4:
[1009] The terminal transmits the converted text data to the server.
[1010] Step 5:
[1011] The server passes the received text data to the generation AI, which analyzes the text data and understands the instruction, "Go to city hall."
[1012] Step 6:
[1013] The server then passes the text data to the emotion engine, which analyzes and evaluates the user's emotional state from the voice data.
[1014] Step 7:
[1015] The server uses the generated AI to obtain the location information of the destination (city hall) based on the user's instructions.
[1016] Step 8:
[1017] The server uses generated AI to adjust flight parameters (speed, altitude, etc.) based on the acquired location information and the analysis results of the emotion engine.
[1018] Step 9:
[1019] The server instructs the AI to generate a specific flight path including flight parameters, causing it to generate flight commands.
[1020] Example: "Take off -> fly 100 meters north -> fly 200 meters east -> set altitude to 50 meters (or lower if user is nervous) -> hover over City Hall -> land."
[1021] Step 10:
[1022] The server sends flight commands to the drone's control board using a communication means.
[1023] Step 11:
[1024] The drone's control board analyzes the flight commands it receives and converts them into specific operational instructions.
[1025] Step 12:
[1026] The drone's control board controls the drone's motors and propellers, making the drone operate according to flight commands.
[1027] Step 13:
[1028] The drone's control board executes a series of operations from takeoff to the destination, finally reaching and landing at the specified location.
[1029] This allows users to easily control the drone using only voice commands, and also ensures safe and optimal flight based on the user's emotional state.
[1030] Example 2
[1031] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1032] Conventional drone control systems have mechanisms to convert user instructions into text data and generate flight commands based on destination information, but they lack the functionality to optimize flight parameters based on the user's emotional state. As a result, if the user becomes nervous or excited, the drone's flight cannot adapt to the user's emotional state, making it difficult to achieve a safe and secure flight.
[1033] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1034] In this invention, the server includes a voice input means for capturing a user's voice input, a voice recognition means for converting the voice captured by the voice input means into text data, an emotion analysis means for recognizing the user's emotional state from the voice data, a means for adjusting flight parameters based on the emotional state recognized by the emotion analysis means, a means for generating flight commands including the adjusted flight parameters, a communication means for transmitting the flight commands generated by the command generation means to a drone control platform, and a drone control means for controlling the drone based on the flight commands transmitted by the communication means. This makes it possible to set optimal flight parameters based on the user's emotional state, thereby achieving safe and secure flight.
[1035] "Audio input means" refers to a device or system for capturing a user's voice.
[1036] A "speech recognition means" is a device or algorithm that converts captured speech into text data.
[1037] A "command generation means" is a system or device for generating drone flight commands using artificial intelligence based on text data.
[1038] "Emotion analysis means" refers to a system or algorithm for recognizing a user's emotional state from their voice data.
[1039] The "flight parameter adjustment means" is a system or device for adjusting flight parameters based on the emotional state recognized by the emotion analysis means.
[1040] "Communication means" refers to a device or system for transmitting the generated flight commands to the drone's control board.
[1041] A "drone control means" is a system or device for controlling a drone based on received flight commands.
[1042] MODE FOR CARRYING OUT THE INVENTION
[1043] The present invention provides a system that combines an emotion analysis function with a system that automatically controls a drone based on the user's voice instructions, and can set optimal flight parameters according to the user's emotional state.
[1044] First, the user issues commands to the drone using a voice input means. For example, the user might issue a voice command such as "Go to City Hall." Next, the device captures the user's voice with a microphone. This captured voice data is converted into text data using a voice recognition means. Specifically, a voice recognition system such as Google Speech-to-Text or Amazon Transcribe is used. The voice recognition means analyzes the voice data and generates the text data "Go to City Hall."
[1045] The converted text data is sent from the device to the server. Standard communication protocols such as HTTP and WebSocket are used for communication. The command generation means implemented on the server passes the text data to the generation AI, which analyzes the user's instructions. OpenAI GPT and other programs can be used as the generation AI. The generation AI obtains the location information of the destination based on the instructions.
[1046] Additionally, the server is equipped with an emotion analysis engine that recognizes the user's emotional state from the voice input. For example, if the user is nervous, it will be recognized as such. The emotion engine achieves this by using an algorithm that analyzes acoustic features (such as tone and pitch of voice).
[1047] Next, the generation AI adjusts flight parameters (speed and altitude) through the emotion engine based on the recognized emotional state. If the user is nervous, adjustments will be made, such as slowing down the flight speed. Based on the adjusted flight parameters, the generation AI generates a specific flight path and creates flight commands. For example, a series of commands might be "Take off -> Move north 100 meters -> Move right (east) 200 meters -> Set altitude to 50 meters -> Hover above City Hall -> Land."
[1048] The generated flight commands are sent to the drone's control platform via the server's communication means using a high-speed communication protocol (such as MQTT or TCP / IP). The drone's control platform ultimately controls the drone based on the received flight commands. The drone follows the specified flight path to its destination and automatically performs operations such as hovering and landing.
[1049] As a specific example, the operation when the user instructs "to the supermarket parking lot" will be shown.
[1050] 1. The user says, "To the supermarket parking lot."
[1051] 2. The device captures the audio with its microphone.
[1052] 3. The speech recognition system converts the speech data into text data.
[1053] 4. The device sends the converted text data to the server.
[1054] 5. The server's generated AI analyzes the text data and obtains the destination's location information.
[1055] 6. The emotion analysis means built into the server recognizes the user's emotional state from the voice data. For example, if the user is excited, it will be recognized as that emotional state.
[1056] 7. The AI generator adjusts flight parameters based on the user's emotional state. For example, if the user is excited, the AI will increase flight speed.
[1057] 8. The generation AI generates a specific flight path and creates flight commands.
[1058] 9. The server sends flight commands to the drone's control board via communication means.
[1059] 10. The drone's control board controls the drone based on the commands, and it flies automatically to the supermarket parking lot.
[1060] Examples of prompt sentences are as follows:
[1061] "Go to City Hall."
[1062] "Fly to the park fountain"
[1063] "Come back to my garden."
[1064] This system allows users to fly safely and securely based on their emotional state, and allows them to easily control the drone using only voice commands.
[1065] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1066] Step 1:
[1067] The user issues a voice command using the voice input means, for example, "Go to city hall." The input here is the user's voice, and the output is the voice data.
[1068] Step 2:
[1069] The device uses a microphone to capture the user's voice, and in the next step, this voice data is passed to a voice recognition system. The input is the user's voice, and the output is the captured voice data.
[1070] Step 3:
[1071] The voice data captured by the device is input into a voice recognition system (for example, Google Speech-to-Text or Amazon Transcribe) and converted into text data. In this conversion, the voice data "Go to City Hall" is converted into text data "Go to City Hall." The input is the captured voice data, and the output is the converted text data.
[1072] Step 4:
[1073] The terminal sends the converted text data to the server. HTTP or WebSocket is used for communication. The input is the converted text data, and the output is data sent to the server.
[1074] Step 5:
[1075] The server passes the received text data to the generation AI for analysis. The generation AI (for example, OpenAI GPT) analyzes the text data to obtain destination information. For example, it analyzes the text data "Go to city hall" to obtain the location information (latitude, longitude, etc.) of city hall. The input is text data, and the output is the analyzed destination information.
[1076] Step 6:
[1077] The server uses emotion analysis to recognize the user's emotional state from the voice data. The emotion analysis engine analyzes acoustic features (tone of voice, pitch, etc.) and estimates the user's emotional state, such as nervousness or calmness. The input is the voice data, and the output is the recognized emotional state.
[1078] Step 7:
[1079] The server's generation AI adjusts flight parameters (speed, altitude, etc.) based on the recognized emotional state. For example, if the user is nervous, the flight speed will be slowed down. A specific flight path is generated based on these flight parameters, and flight commands are created. The input is the emotional state and destination information, and the output is the adjusted flight parameters and specific flight commands.
[1080] Step 8:
[1081] The flight commands generated by the server are sent to the drone's control platform via a communication method. High-speed communication protocols such as MQTT and TCP / IP are used for communication. The input is the flight command, and the output is the command sent to the drone's control platform.
[1082] Step 9:
[1083] The drone will fly automatically based on flight commands received through the control board. The drone will follow the specified flight path to its destination and automatically perform operations such as landing and hovering. The input is the flight command, and the output is the drone's actual flight behavior.
[1084] (Application example 2)
[1085] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1086] In conventional food delivery systems, it is difficult to control drones based on voice commands, and it is not possible to set optimal flight parameters according to the user's emotional state, which can lead to anxiety and stress during delivery.In addition, it is difficult to adjust flight parameters in real time, leaving issues in terms of safety and efficiency.
[1087] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1088] In this invention, the server includes a voice input means for capturing a user's voice input, a voice recognition means for converting the captured voice into text data, an emotion recognition means for recognizing the user's emotional state using an emotion engine, a parameter adjustment means for setting drone flight parameters according to the user's emotional state, a command generation means for generating drone flight commands using generative artificial intelligence, a communication means for transmitting the generated flight commands to a drone control platform, and a drone control means for controlling the drone based on the transmitted flight commands. This allows the user to simply issue voice commands to set optimal drone flight parameters according to their emotional state, enabling safe and efficient food delivery.
[1089] "Audio input means" is a device or function for capturing a user's voice.
[1090] "Speech recognition means" refers to technology or devices that analyze captured voice data and convert it into text data.
[1091] "Emotion recognition means" refers to technology or devices for recognizing a user's emotional state from voice data or text data.
[1092] "Parameter adjustment means" refers to a function for setting and adjusting the drone's flight parameters (e.g., speed, altitude, etc.) based on the recognized emotional state of the user.
[1093] A "command generation means" is a device or function that generates flight commands for the drone based on the user's voice input and emotional state.
[1094] "Communication means" refers to the method or technology used to transmit the generated flight commands to the drone's control platform.
[1095] "Drone control means" means a device or function for controlling and managing a drone based on received flight commands.
[1096] "Generative AI" is an AI technology that uses large amounts of data to analyze user instructions and generate appropriate commands.
[1097] The "emotion engine" is a technology that detects the user's emotional state from voice and text data and uses it to adjust flight parameters.
[1098] This invention is a drone control system that can capture the user's voice instructions and combine them with an emotion engine to set optimal flight parameters according to the user's emotional state.
[1099] The server first has a voice input means for capturing voice input from the user, which is a voice input device such as a microphone, and captures voice instructions given by the user.
[1100] Once the audio is captured, a speech recognition means processes it and converts it into text data, for example, using the speech_recognition library, which generates text data from the user's voice commands.
[1101] Next, the emotion recognition means recognizes the user's emotional state from the voice. The emotion recognition means uses an emotion engine to extract emotions from the voice. An example of this emotion engine is the EmotionRecognizer module.
[1102] Based on the recognized emotional state, the parameter adjustment means sets and adjusts the optimal drone flight parameters, including the drone's speed and altitude. For example, if the user is nervous, the flight speed may be slowed down.
[1103] The command generation means then uses a generative AI model to generate specific flight commands. This generative AI model embodies the drone's flight route and operations based on the generated text data and the user's emotional state. For example, it generates detailed commands for takeoff, flight path, altitude, hovering, landing, etc.
[1104] The generated flight commands are sent to the drone's control platform via a communication method that uses a high-speed, highly reliable protocol to ensure real-time communication, such as Wi-Fi or 4G / 5G communication.
[1105] Based on the transmitted flight commands, the drone control means controls the drone, allowing the drone to travel to the destination along the specified flight route, achieving safe and efficient flight according to the user's emotional state.
[1106] Specific examples
[1107] Example prompt sentence:
[1108] The user gives a voice command such as "Deliver pizza to my house."
[1109] In the processing flow, voice input is captured and converted into text data such as "Deliver pizza to my house" by the voice recognition means. Then, the emotion recognition means recognizes the user's emotional state, and the parameter adjustment means sets flight parameters. Finally, the generated flight command is sent to the drone, which then delivers the pizza to the house.
[1110] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1111] Step 1: Capturing Audio Input
[1112] The server uses a voice input means (microphone) to capture the user's voice instruction. This input is, for example, the user's voice data saying, "Deliver a pizza to my house." The captured voice data is saved for processing in a later step.
[1113] Step 2: Voice Recognition
[1114] The server uses a speech recognition tool (the speech_recognition library) to convert the captured voice data into text data, which contains specific instructions such as "Deliver a pizza to my house." The converted text data is used in the next processing step.
[1115] Step 3: Emotion Recognition
[1116] The server uses an emotion recognition module (EmotionRecognizer module) to recognize the user's emotional state based on the converted text data. It analyzes the features extracted from the voice data and identifies the user's emotional state, such as "joy." The recognized emotional state is then used to set flight parameters.
[1117] Step 4: Obtaining destination information
[1118] The server uses the command generation means to obtain destination information from the user's voice input. Specifically, it analyzes the text data "Deliver pizza to my house" and obtains the coordinates (latitude and longitude) of "home" using a geographic information acquisition library (such as geopy). The obtained coordinate information is used to generate a flight route.
[1119] Step 5: Adjusting Flight Parameters
[1120] The server uses a parameter adjustment means to set the drone's flight parameters according to the recognized emotional state. For example, if the user is "nervous," the server sets the drone's flight speed slower to increase safety. This adjustment allows the drone to fly in a way that takes the user's emotions into consideration.
[1121] Step 6: Generate Flight Commands
[1122] The server generates specific flight commands using a command generation means and a generative AI model. The generative AI model outputs detailed flight commands, such as takeoff, route, altitude setting, hovering, and landing, based on the user's destination information and emotional state. These flight commands serve as guidelines for the drone's operation.
[1123] Step 7: Sending Flight Commands
[1124] The server then transmits the generated flight commands to the drone's control platform via a communication method using a high-speed, highly reliable protocol (such as Wi-Fi or 4G / 5G communication), enabling accurate command transmission in real time.
[1125] Step 8: Controlling the drone
[1126] The drone operates based on the flight commands received. The drone control means flies the pizza to the customer's home according to the specified flight route and parameters, realizing safe and efficient food delivery that takes into consideration the customer's feelings.
[1127] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1128] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1129] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1130] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1131] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1132] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1133] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1134] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1135] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1136] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1137] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1138] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1139] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1140] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1141] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1142] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1143] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1144] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1145] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1146] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1147] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1148] The following is further disclosed regarding the above embodiment.
[1149] (Claim 1)
[1150] a voice input means for capturing a user's voice input;
[1151] a voice recognition means for converting the voice captured by the voice input means into text data;
[1152] a command generation means for generating flight commands for the drone using a generation artificial intelligence based on the text data converted by the voice recognition means;
[1153] a communication means for transmitting the flight command generated by the command generation means to a control board of the drone;
[1154] a drone control means for controlling the drone based on the flight command transmitted by the communication means;
[1155] A system including:
[1156] (Claim 2)
[1157] The system of claim 1 , wherein the command generation means obtains destination information based on a user's voice input and generates a flight command for the drone including a flight path to the destination.
[1158] (Claim 3)
[1159] 10. The system of claim 1, wherein the communication means transmits flight commands to the drone's control board using a high-speed communication protocol.
[1160] "Example 1"
[1161] (Claim 1)
[1162] a voice input means for capturing a user's voice input;
[1163] a voice recognition means for converting the voice captured by the voice input means into text data;
[1164] an analysis means for analyzing the text data converted by the speech recognition means using a generation artificial intelligence to acquire destination information;
[1165] a command generation means for generating a flight command for the drone based on the destination information acquired by the analysis means;
[1166] a communication means for transmitting the flight command generated by the command generation means to a control board of the drone;
[1167] a drone control means for controlling the drone based on the flight command transmitted by the communication means;
[1168] A system including:
[1169] (Claim 2)
[1170] 2. The system according to claim 1, wherein a specific flight path is generated based on the destination information acquired by the analysis means and included in the flight command.
[1171] (Claim 3)
[1172] 10. The system of claim 1, wherein the communication means transmits flight commands to the drone's control board using a high-speed communication protocol.
[1173] "Application Example 1"
[1174] (Claim 1)
[1175] a voice input means for capturing a user's voice input;
[1176] a voice recognition means for converting the voice captured by the voice input means into text data;
[1177] a command generation means for generating flight commands for the drone using a generation artificial intelligence based on the text data converted by the voice recognition means;
[1178] a communication means for transmitting the flight command generated by the command generation means to a control board of the drone;
[1179] a drone control means for controlling the drone based on the flight command transmitted by the communication means;
[1180] a logistics management means for analyzing instructions and generating flight paths for transporting cargo in a logistics center;
[1181] a location information acquisition means for identifying the location and destination of the package based on instructions from the user and generating a flight path to the destination;
[1182] A system including:
[1183] (Claim 2)
[1184] The system of claim 1 , wherein the command generation means obtains destination information based on a user's voice input and generates a flight command for the drone including a flight path to the destination.
[1185] (Claim 3)
[1186] 10. The system of claim 1, wherein the communication means transmits flight commands to the drone's control board using a high-speed communication protocol.
[1187] "Example 2: Combining Emotion Engines"
[1188] (Claim 1)
[1189] a voice input means for capturing a user's voice input;
[1190] a voice recognition means for converting the voice captured by the voice input means into text data;
[1191] a command generation means for generating flight commands for the drone using a generation artificial intelligence based on the text data converted by the voice recognition means;
[1192] an emotion analysis means for recognizing an emotional state of a user from the user's voice data;
[1193] means for adjusting flight parameters based on the emotional state recognized by the emotion analysis means;
[1194] means for generating flight commands including the adjusted flight parameters;
[1195] a communication means for transmitting the flight command generated by the command generation means to a control board of the drone;
[1196] a drone control means for controlling the drone based on the flight command transmitted by the communication means;
[1197] A system including:
[1198] (Claim 2)
[1199] The system of claim 1 , wherein the command generation means obtains destination information based on a user's voice input and generates a flight command for the drone including a flight path to the destination.
[1200] (Claim 3)
[1201] 10. The system of claim 1, wherein the communication means transmits flight commands to the drone's control board using a high-speed communication protocol.
[1202] "Application example 2 when combining emotion engines"
[1203] (Claim 1)
[1204] a voice input means for capturing a user's voice input;
[1205] a voice recognition means for converting the voice captured by the voice input means into text data;
[1206] emotion recognition means for recognizing an emotional state of a user using an emotion engine based on the text data converted by the speech recognition means;
[1207] a parameter adjustment means for setting flight parameters of the drone according to the emotional state of the user based on the emotion recognition means;
[1208] a command generation means for generating flight commands for the drone using a generation artificial intelligence based on the flight parameters set by the parameter adjustment means;
[1209] a communication means for transmitting the flight command generated by the command generation means to a control board of the drone;
[1210] a drone control means for controlling the drone based on the flight command transmitted by the communication means;
[1211] A system including:
[1212] (Claim 2)
[1213] The system of claim 1 , wherein the command generation means obtains destination information based on a user's voice input and generates a flight command for the drone including a flight path to the destination.
[1214] (Claim 3)
[1215] 10. The system of claim 1, wherein the communication means transmits flight commands to the drone's control board using a high-speed communication protocol. [Explanation of symbols]
[1216] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a voice input means for capturing a user's voice input; a voice recognition means for converting the voice captured by the voice input means into text data; a command generation means for generating flight commands for the drone using a generation artificial intelligence based on the text data converted by the voice recognition means; a communication means for transmitting the flight command generated by the command generation means to a control board of the drone; a drone control means for controlling the drone based on the flight command transmitted by the communication means; A system including:
2. The system of claim 1 , wherein the command generation means acquires destination information based on a user's voice input and generates a flight command for the drone including a flight path to the destination.
3. The system of claim 1 , wherein the communication means transmits flight commands to the drone's control board using a high-speed communication protocol.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A