system

A system using voice recognition, data processing, and augmented reality addresses the challenges faced by elderly agricultural workers, improving efficiency and communication within agricultural communities.

JP2026073356APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Elderly agricultural workers face challenges in using digital devices, and there is a lack of effective communication and knowledge sharing among agricultural communities, leading to decreased efficiency and interest in agriculture among younger generations.

Method used

A system utilizing voice recognition, data processing, speech synthesis, augmented reality, and IoT devices for environmental data management, along with community building tools, to facilitate information sharing and educational support.

Benefits of technology

Enables elderly users to intuitively access agricultural information, improves work efficiency, and promotes interaction and learning among agricultural workers, bridging the technological gap and enhancing interest in agriculture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073356000001_ABST
    Figure 2026073356000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A voice recognition means that converts voice input into a digital signal, A data processing means for analyzing the converted digital signal and obtaining related information, A speech synthesis means that converts acquired information into speech output and provides it to the user, An interface means that simplifies operation through a user interface and supports use by elderly users, Educational support tools for providing agricultural educational content in augmented reality format, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] To eliminate technical barriers for elderly agricultural workers and facilitate the use of digital devices, and to improve the efficiency of work in the agricultural field. Further, it aims to provide knowledge and technology to increase the interest of the younger generation in agriculture and to solve the problem of activating communication among those engaged in agriculture.

Means for Solving the Problems

[0005] By using voice recognition to easily convert voice input into digital signals and employing data processing to acquire related information, even elderly users can intuitively obtain information. Furthermore, the acquired information is output as voice using speech synthesis to provide users with appropriate information. In addition, the implementation of educational support using augmented reality (AR) technology enables the visual and effective delivery of educational content to young people. Moreover, environmental data management using IoT devices collects and analyzes agricultural environmental data in real time, thereby improving work efficiency. Finally, community building means are included to promote interaction among users and support information sharing.

[0006] "Voice recognition means" refers to technology that has the function of converting the user's voice into a digital signal.

[0007] "Data processing means" refers to technology that has the function of analyzing converted digital signals and obtaining related information.

[0008] "Speech synthesis means" refers to a technology that has the function of outputting acquired information as speech and providing it to the user.

[0009] "Interface means" refers to technologies that have the function of simplifying operations through a user interface and supporting elderly users so that they can use the system easily.

[0010] "Educational support tools" refer to technologies that provide agricultural education content in augmented reality format and have the function of improving learning effectiveness.

[0011] "Environmental data management means" refers to a technology that has the function of acquiring and analyzing environmental data in real time from IoT devices via an information and communication network.

[0012] "Community building tools" are technologies that have functions to promote information sharing among users and revitalize communication. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of the data processing device and smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a tagged processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, a tagged RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, a tagged storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention is implemented by a system comprising speech recognition means, data processing means, speech synthesis means, interface means, educational support means, environmental data management means, and community formation means.

[0035] The voice recognition system has the function of converting the user's voice input into a digital signal. When the user asks a question to the voice assistant, the device captures and digitizes the voice. This digital data is sent to a server and analyzed by data processing equipment.

[0036] The server uses data processing equipment to analyze the received digital signals and extract relevant information. This information may include specific agricultural data or data for work support. This information is then converted into a speech format via speech synthesis equipment and provided to the user.

[0037] The interface provides a simple user interface, making it easy for elderly users to operate. This allows users to operate the device intuitively.

[0038] The educational support measures will provide educational programs for young people that utilize augmented reality (AR) technology. The devices will use their camera functions to visually display agricultural equipment and work procedures overlaid on real-world scenery, providing an interactive learning experience.

[0039] The environmental data management system supports the efficiency of agricultural work by collecting agricultural environmental data in real time via IoT devices and analyzing it on a server. This enables timely work instructions.

[0040] The community-building tools provide users with the ability to share information online and communicate with other farmers. This enables knowledge sharing and community revitalization.

[0041] For example, when an elderly user asks the device, "What's the weather like today?", the device recognizes the voice and sends data to the server. The server collects the current weather information and responds to the device verbally. Furthermore, when a younger user uses an AR educational program on their smartphone, the system can provide visual instruction by overlaying instructions on how to use virtual agricultural machinery onto the real-world environment. In this way, the system bridges the technical gap among agricultural workers and enables more efficient farming.

[0042] The following describes the processing flow.

[0043] Step 1:

[0044] The user speaks aloud to the voice assistant, giving questions or instructions.

[0045] Step 2:

[0046] The device uses a microphone to capture the user's voice and converts the voice into a digital signal using speech recognition technology.

[0047] Step 3:

[0048] The terminal sends the converted digital signal to the server. The transmitted data includes the user's voice content as well as necessary contextual information.

[0049] Step 4:

[0050] The server uses data processing equipment to analyze the received digital signals and identify relevant information. For example, if a user requests weather information, the server will refer to a weather information database to obtain the necessary data.

[0051] Step 5:

[0052] The server converts the acquired information into text data for audio output and sends it to the terminal as a response.

[0053] Step 6:

[0054] The terminal converts the received text data into speech using a speech synthesis system and conveys the information to the user through the speaker.

[0055] Step 7:

[0056] The user takes appropriate action based on the provided voice information. If necessary, they can ask additional questions and restart the process.

[0057] (Example 1)

[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0059] In modern agriculture, older generations find it difficult to utilize digital technology, creating a technological gap with younger generations. Furthermore, insufficient information sharing and utilization of environmental data are leading to decreased efficiency in farm work. To address these challenges, a system is needed that effectively utilizes voice recognition and data analysis while maintaining user-friendly operation for the elderly.

[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0061] In this invention, the server includes a speech recognition means for converting voice input into a digital signal, a data processing means for analyzing the converted digital signal and obtaining relevant information, and a speech synthesis means for converting the obtained information into voice output and providing it to the user. This enables real-time information acquisition and appropriate agricultural work support through a user interface that can be easily operated even by the elderly.

[0062] "Voice recognition means" refers to an element that has the function of converting voice input into a digital signal.

[0063] A "digital signal" is a signal in digital format converted by speech recognition means and analyzed by data processing means.

[0064] A "data processing means" is an element that has the function of analyzing the received digital signal and extracting relevant information.

[0065] A "speech synthesis means" is an element that has the function of converting information acquired by a data processing means into speech output and providing it to the user.

[0066] "Interface means" refers to elements that simplify operation through a user interface and support the use of the device by elderly users.

[0067] An "educational support tool" is an element that has the function of providing agricultural education content in an augmented reality format.

[0068] A "generative model" refers to algorithms and techniques used to generate appropriate responses based on received data.

[0069] "Environmental data management means" refers to elements for acquiring environmental data from sensor devices via an information and communication network and analyzing it using data processing means.

[0070] "Information sharing means" are elements that provide collaborative functions, allow users to share information with each other, and facilitate communication.

[0071] This invention is a comprehensive system for effectively utilizing voice input to provide agricultural information and educational support. The system primarily comprises voice recognition means, data processing means, voice synthesis means, interface means, educational support means, environmental data management means, and information sharing means. This enables both the elderly and young people to receive agricultural support using digital technology.

[0072] The device captures voice input from the user and converts it into a digital signal using speech recognition. This speech recognition uses a general-purpose speech recognition engine (e.g., a speech API) to convert the voice into text data while removing noise.

[0073] The server receives digital data obtained through speech recognition and analyzes it using data processing tools. Here, natural language processing techniques are used to interpret the user's intent, and relevant information is extracted by a generating AI model (e.g., a language model API). The generated information is then converted into natural-sounding speech by speech synthesis tools. For speech synthesis, a speech synthesis engine is used to generate easy-to-understand speech.

[0074] Users can utilize an interface that allows for easy operation through the system. In particular, an intuitive user interface is provided for elderly users to enhance convenience.

[0075] Furthermore, as a means of educational support, augmented reality technology will be used to provide young people with visually and interactively educational content about agriculture. This will be achieved by using the camera function of a smartphone to overlay agriculture-related information onto real-world scenery.

[0076] Furthermore, the environmental data management system acquires environmental data in real time from sensor devices via an information and communication network and analyzes it on a server. Based on this information, it becomes possible to provide instructions that support efficient agricultural work.

[0077] Information sharing tools allow users to share information with other users online and facilitate communication. This enables knowledge exchange and collaboration among agricultural workers.

[0078] As a concrete example, when an elderly person speaks to the terminal saying, "Tell me today's weather," the terminal captures the voice and converts it into digital data. The server then analyzes this data to obtain weather information and responds in voice. As an example of a prompt, providing the input "Create a description of the voice recognition system for improving agricultural efficiency" to the generating AI model will yield relevant information.

[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0080] Step 1:

[0081] The user speaks into the device, providing voice input. The device captures this voice data with its microphone and filters out background noise to produce clearer audio.

[0082] Step 2:

[0083] The device uses a speech recognition engine to convert the captured audio into a digital signal. Using a speech API, the audio data is converted into text data, making the audio content analyzable. The output of this step is the converted text data.

[0084] Step 3:

[0085] The terminal sends text data to the server. A secure protocol (e.g., HTTPS) is used for transmission, and the data is encrypted to ensure the security of the transmitted content.

[0086] Step 4:

[0087] The server analyzes the received text data using data processing tools. A natural language processing engine is used to extract the intent of the question and keywords, and to identify relevant information. Once this information extraction is complete, data for the next step is generated.

[0088] Step 5:

[0089] The server uses a generative AI model to generate an appropriate response based on the extracted information. The generated response is stored on the server in text format.

[0090] Step 6:

[0091] The server passes the generated text response to the speech synthesis engine, which converts it into speech format. Using the speech synthesis API, it generates natural, human-like speech. Speech format data is then generated.

[0092] Step 7:

[0093] The terminal receives audio data from the server and provides audio output to the user through its speaker. This allows the user to hear the answer to their question.

[0094] Step 8:

[0095] Users can ask additional questions as needed, repeating the process from speech recognition to speech synthesis to obtain the necessary information.

[0096] (Application Example 1)

[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0098] Improving work efficiency and bridging the skills gap among workers with varying levels of experience are key challenges in the manufacturing industry. In particular, there is a lack of intuitive work support and training opportunities for older workers and new recruits; therefore, a flexible system is needed to address these issues.

[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0100] In this invention, the server includes speech recognition means for converting acoustic input into an encoded signal, data processing means for analyzing the encoded signal and obtaining related information, and speech synthesis means for converting the obtained information into an acoustic output and providing it to the user. This enables users to operate and obtain information using their voice, allowing elderly people and newcomers to perform factory work efficiently.

[0101] "Auditory input" refers to sound information obtained through hearing, such as the user's voice or ambient sounds.

[0102] An "encoded signal" is signal data obtained by converting a continuous acoustic input into a digital format.

[0103] "Speech recognition means" refers to a device or program for converting an acoustic input into an encoded signal.

[0104] "Data processing means" refers to a device or program that analyzes an encoded signal and extracts or calculates related information.

[0105] "Speech synthesis means" refers to a device or program that converts acquired information back into an acoustic format and provides it to the user.

[0106] "User interface" is a concept that refers to the means of displaying or inputting information that a user uses to operate a device or system.

[0107] "Elderly people" refers to people of an age group who require support, particularly considering physical or cognitive changes that result from aging.

[0108] Augmented reality is a technological format that overlays virtual information onto the real environment.

[0109] An "information transmission network" is a communication infrastructure used for exchanging data with remote locations.

[0110] "Internet of Things devices" refer to physical devices and systems that are interconnected via the internet.

[0111] "Status data" refers to data about the operating status and environmental conditions collected from Internet of Things devices.

[0112] "Cooperative features" are part of a system that allows users to support each other and share information.

[0113] "Support measures" refer to devices, methods, or tools that enable efficient work.

[0114] In order to implement this invention, it is necessary to use a combination of various hardware and software. Specifically, these include a server, a speech recognition device for encoding acoustic input, a programming library for data analysis, and software for synthesizing sound.

[0115] The server receives an encoded signal from a speech recognition device that receives acoustic input. This is done using a common speech recognition API. For example, the Google® Cloud Speech-to-Text API can be used to convert acoustic input into text data. The server then analyzes this encoded signal using data processing tools and collects relevant information. Python programs and data processing libraries such as NumPy and Pandas can be used for data analysis.

[0116] The acquired information is converted back into an audio format using speech synthesis technology. This can be done using speech synthesis APIs such as Amazon Polly, allowing the information to be presented in a user-friendly format.

[0117] The user interface should be designed with ease of use for the elderly in mind. It should incorporate an intuitive touch interface and a system that accepts voice commands. For example, if a user says, "Tell me today's work schedule," the system should display the scheduled tasks and provide voice guidance.

[0118] Status data collected via Internet of Things (IoT) devices allows for real-time management of environmental conditions and equipment operating status within the factory, and data analysis enables the proposal of efficient work schedules. This allows for immediate and specific work instructions to be given to qualified workers.

[0119] As a concrete example, let's consider a scenario where a user says, "I want to check the current status of the production line." An example of a prompt for the generating AI model would be, "Please simulate a voice assistant aimed at improving the efficiency of factory operations. This assistant will use speech recognition, data analysis, and speech synthesis to provide the user with real-time production status."

[0120] This invention aims to efficiently and intuitively support work in manufacturing environments, and is designed to be easily operated by all workers, including elderly and new employees.

[0121] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0122] Step 1:

[0123] The device captures the user's audio input and sends it to a speech recognition device. The input is the user's voice, and the output is a digitized encoded signal. This signal is parsed using the Google Cloud Speech-to-Text API, and the audio signal is converted into text.

[0124] Step 2:

[0125] The server receives encoded signals and performs analysis using data processing tools. The input is encoded text data, and the output is relational information. A Python program analyzes the text data and efficiently extracts information using NumPy and Pandas.

[0126] Step 3:

[0127] The server converts the analyzed information back into speech format using speech synthesis technology. The input is the analyzed information, and the output is the synthesized speech. Amazon Polly is used to generate and provide speech in a format that is easy for the user to understand.

[0128] Step 4:

[0129] The device uses speech synthesis output to notify the user of information verbally. Here, the user can easily obtain information through the voice interface. The information is conveyed at a speed and volume that is easy for elderly people to hear.

[0130] Step 5:

[0131] The user requests status data from an Internet of Things (IoT) device to the server. The input is a voice command from the user, and the server collects this data and performs integrated data analysis.

[0132] Step 6:

[0133] Based on data collected from IoT devices, the server provides users with detailed information about the current environmental conditions and status of the factory. The output is the user's environmental monitoring information, which is used to facilitate efficient operations. An example of its use in a generated AI model is the prompt message: "Simulate a voice assistant aimed at improving the efficiency of factory work. This assistant uses speech recognition, data analysis, and speech synthesis to provide users with real-time production status."

[0134] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0135] This invention is implemented by a system that includes speech recognition means, data processing means, speech synthesis means, interface means, educational support means, environmental data management means, community formation means, and emotion engine means.

[0136] The emotion engine has the function of recognizing emotions from the user's voice input. When the user asks a question or gives an instruction to the voice assistant, the terminal converts the voice into a digital signal. The converted digital signal is sent to a server and analyzed by the emotion engine. The server identifies the user's emotion based on the characteristics of the voice and generates a response optimized for that emotion through a data processing means. The obtained information is output as voice using a speech synthesis means and delivered to the user.

[0137] The interface is designed to enable intuitive operation, making it easy for elderly users to utilize the system. The educational support system provides new experiences based on the user's emotional state. Educational programs for young people are displayed in augmented reality format via the device's camera function, enabling contextually adaptive learning.

[0138] The environmental data management system analyzes agricultural environmental data collected from IoT devices in real time using an information and communication network, improving user work efficiency. Furthermore, the community building system facilitates online information exchange and communication, strengthening collaboration among agricultural workers.

[0139] For example, if a user says, "I'm not feeling very good right now," the device analyzes that emotion, and the server generates appropriate advice to reduce stress. Similarly, in an educational support scenario, if a user feels frustrated with learning, the learning content is adjusted through an emotion engine to make it more engaging. This system makes it possible to provide efficient agricultural support while being mindful of the user's emotions.

[0140] The following describes the processing flow.

[0141] Step 1:

[0142] Users speak to the voice assistant, asking questions, giving instructions, and expressing their current feelings.

[0143] Step 2:

[0144] The device uses a microphone to capture the user's voice and uses speech recognition to convert the voice into a digital signal.

[0145] Step 3:

[0146] The terminal sends the converted digital signal to the server. In addition to the audio content, it also includes user context information.

[0147] Step 4:

[0148] The server utilizes data processing tools to analyze the voice data. Furthermore, an emotion engine identifies the user's emotions from the voice.

[0149] Step 5:

[0150] The server generates a user-optimized response based on the emotions it identifies. For example, if the user is feeling stressed, it might incorporate relaxation advice.

[0151] Step 6:

[0152] The response generated by the server is converted into audio data using speech synthesis technology and sent to the terminal.

[0153] Step 7:

[0154] The device plays the received audio data and provides a response to the user through the speaker.

[0155] Step 8:

[0156] When users access educational content, the device uses educational support tools to deliver the educational program in an augmented reality format. The content is adjusted to the user's emotions to ensure effective learning.

[0157] (Example 2)

[0158] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0159] Systems utilizing speech recognition technology require more user-friendly interfaces for the elderly and flexible responses that can accommodate the emotions of individual users. Furthermore, improving learning efficiency through the provision of agricultural-related educational content is also crucial. The need for means to enhance information sharing and communication with others is also increasing.

[0160] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0161] In this invention, the server includes processing means for converting acoustic data into digital information, computing means for analyzing the converted digital information and extracting relevant information, and acoustic synthesis means for converting the extracted information into acoustic output and providing it to the user. This enables the recognition of emotions from voice input, the provision of responses optimized for individual users, and the promotion of educational content and communication utilizing augmented reality.

[0162] The "processing department" is the department responsible for converting audio data into digital information.

[0163] The "Computation Department" is a department that has the function of analyzing converted digital information and extracting relevant information.

[0164] The "Sound Synthesis Department" is the department responsible for converting extracted information into sound output and providing it to users.

[0165] The "User Interface Department" is a department that simplifies operations through the user interface and provides support for elderly users.

[0166] The "Educational Support Department" is a department that has the function of providing educational materials related to agriculture in augmented reality format.

[0167] The "Emotion Analysis Department" is responsible for recognizing emotions from user voice input and optimizing responses.

[0168] The "Environmental Data Management Department" is a department that has the function of acquiring and analyzing environmental information from sensor devices via an information and communication network.

[0169] The "Collaboration Building Department" is a department that has the function of sharing information with other users and promoting communication.

[0170] This invention is a system that specifically analyzes user voice input, recognizes emotions, and responds appropriately. Specific embodiments for implementing this system are described below.

[0171] The user provides questions or instructions to the system through a voice input device. The terminal receives this voice and, in the first step, converts the voice into a digital signal using speech recognition technology. Speech recognition software is typically used for this purpose. For example, when converting audio data into text data, a speech recognition API can be used.

[0172] The converted digital signal is sent to a server. The server uses an emotion analysis engine to analyze this audio data and identify the user's emotions. Machine learning models and natural language processing are used for the analysis, such as the BERT model. Based on the analysis results, the server understands the content of the voice spoken by the user and generates the optimal response through its data processing functions.

[0173] The generated response is converted back from text to speech using speech synthesis technology. The terminal outputs this speech in a format audible to the user. An example of a prompt is, "Explain how to recognize emotions and generate a response when the user asks about their current feelings." This prompt supports output from a generative AI model that takes user reactions into account.

[0174] For example, if a user says, "I'm feeling down today," the server's emotion analysis engine receives this and generates advice to alleviate stress, outputting it as a voice message from the terminal, such as, "Try taking some deep breaths to relax." This system provides responses that take emotions into account, showing empathy to the user and achieving a more comfortable interaction.

[0175] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0176] Step 1:

[0177] The user speaks into the voice input device, asking questions or giving instructions. The device receives this voice input and uses speech recognition technology to convert the speech into a digital signal. The input is analog audio data, and the output is digital data in text format. This includes the specific operation of converting speech to text using a speech recognition API.

[0178] Step 2:

[0179] The terminal sends the converted digital data to the server. The server sends the received text data to an emotion analysis engine to analyze the user's emotions. It receives text data as input and generates analyzed emotion information as output. Machine learning models are used, and data analysis is performed, particularly using natural language processing techniques.

[0180] Step 3:

[0181] The server generates an appropriate response using data processing functions based on the analysis results. The input is analyzed emotional information, and the output is the text data of the response. The server uses a generative AI model to perform specific actions to construct a natural response that is appropriate to the emotion and scenario.

[0182] Step 4:

[0183] The text data of the response returned from the server to the terminal is converted back into speech by a speech synthesis system. The terminal then plays this speech data to the user. The input is the text data of the response, and the output is the speech data. The specific operation in this step is the conversion from text to speech using a speech synthesis engine.

[0184] (Application Example 2)

[0185] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0186] In modern brick-and-mortar stores, customer service that can immediately respond to the diverse needs and emotions of customers is required. However, conventional systems struggle to provide responsive service that is tailored to the customer's emotions and situation, necessitating further efficiency and personalization. To solve this problem, a system is needed that can recognize customer emotions and provide the most appropriate response.

[0187] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0188] In this invention, the server includes speech recognition means for converting voice input into a digital signal, data processing means for analyzing the converted digital signal and obtaining relevant information, and speech synthesis means for converting the obtained information into voice output and providing it to the user. This makes it possible to analyze the emotions of a customer in real time from their voice and provide an optimal customer service response that corresponds to those emotions.

[0189] "Speech recognition means" refers to a device or process that converts speech into a digital signal.

[0190] "Data processing means" refers to a function or device for analyzing digital signals and obtaining related information.

[0191] "Speech synthesis means" refers to a technology or device for converting acquired information into speech output and providing it to the user.

[0192] "Interface means" refers to a design or device that simplifies operation through a user interface and assists elderly users in using the device.

[0193] "Customer service support tools" refer to functions or systems that support customer service in stores and generate optimal responses according to the customer's emotions.

[0194] "Environmental data management means" refers to a technology or device for acquiring environmental data from IoT devices via an information and communication network and analyzing it using data processing means.

[0195] "Community building means" refers to functions or systems that provide community features, allow users to share information with each other, and facilitate communication.

[0196] The system realizing this invention primarily utilizes means for speech recognition, data processing, speech synthesis, interface, customer service support, environmental data management, and community building. It acquires the user's voice through smart glasses or other voice input devices and converts it into a digital signal in real time. Cloud-based speech recognition services such as Google Cloud Speech-to-Text are used for speech recognition. This digital signal is sent to a server and analyzed by sentiment analysis software such as IBM Watson® Natural Language Understanding. Based on the analyzed sentiment data, an appropriate response or customer service method is generated using OpenAI®'s GPT-based model. This response is then provided to the user via speech synthesis.

[0197] Furthermore, if a user is wearing smart glasses in the store, the display will show customer service recommendations. This allows staff to provide service while considering the customer's emotions. For example, if a customer appears tired when asking about a product, the emotion engine will recognize that the customer is tired and suggest products that promote relaxation or encourage a gentler tone of voice.

[0198] For example, when a user says, "I'm looking for something to help me relax," the server can use the information that "the customer seems to want to relax" to generate recommendations for products with relaxation effects. An example of a prompt to the generating AI model would be, "The customer is asking about products with relaxation effects. What do you recommend?"

[0199] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0200] Step 1:

[0201] The device acquires the user's voice. The input is the user's voice, which is collected using the microphone of smart glasses or a smartphone. The output is real-time audio data. This audio data is temporarily stored within the device.

[0202] Step 2:

[0203] The device converts the audio data into a digital signal. The input is the audio data collected in step 1, which is then converted into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text). The output is text data. This converted digital signal is immediately sent to the server.

[0204] Step 3:

[0205] The server analyzes the text data to identify the user's emotions. The input is the text data generated in step 2, which is analyzed using emotion analysis software (e.g., IBM Watson Natural Language Understanding). The output is data indicating the user's emotions. Further analysis is performed based on this data.

[0206] Step 4:

[0207] The server generates optimal responses and product recommendations based on the analysis results. The input is the sentiment data obtained in step 3, and a generative AI model (e.g., OpenAI's GPT-based model) is used to generate the response. The output is the content and conversation suggested to the user. This generated response is used in the next step.

[0208] Step 5:

[0209] The device synthesizes the generated response into speech and provides it to the user. The input is the response data generated in step 4, which is converted into speech output using speech synthesis technology. The output is the content provided to the user aurally. As a result, the user can receive a response that aligns with their own emotions.

[0210] Step 6:

[0211] Visual information is displayed on the user's display (such as the screen of smart glasses). The input is the information generated in step 4, and the content is displayed on the display screen at the appropriate time. The output is the visually presented information. This display allows the user to enjoy a deeper customer service experience.

[0212] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0213] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0214] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0215] [Second Embodiment]

[0216] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0217] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0218] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0219] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0220] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0221] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0222] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0223] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0224] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0225] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0226] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0227] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0228] This invention is implemented by a system comprising speech recognition means, data processing means, speech synthesis means, interface means, educational support means, environmental data management means, and community formation means.

[0229] The voice recognition system has the function of converting the user's voice input into a digital signal. When the user asks a question to the voice assistant, the device captures and digitizes the voice. This digital data is sent to a server and analyzed by data processing equipment.

[0230] The server uses data processing equipment to analyze the received digital signals and extract relevant information. This information may include specific agricultural data or data for work support. This information is then converted into a speech format via speech synthesis equipment and provided to the user.

[0231] The interface provides a simple user interface, making it easy for elderly users to operate. This allows users to operate the device intuitively.

[0232] The educational support measures will provide educational programs for young people that utilize augmented reality (AR) technology. The devices will use their camera functions to visually display agricultural equipment and work procedures overlaid on real-world scenery, providing an interactive learning experience.

[0233] The environmental data management system supports the efficiency of agricultural work by collecting agricultural environmental data in real time via IoT devices and analyzing it on a server. This enables timely work instructions.

[0234] The community-building tools provide users with the ability to share information online and communicate with other farmers. This enables knowledge sharing and community revitalization.

[0235] For example, if an elderly user asks the device, "What's the weather like today?", the device recognizes the voice and sends data to the server. The server collects the current weather information and responds to the device verbally. Furthermore, when younger users use an AR educational program on their smartphones, the system can provide visual instruction by overlaying instructions on how to use virtual agricultural machinery onto the real-world environment. In this way, the system bridges the technical gap among agricultural workers and enables more efficient farming.

[0236] The following describes the processing flow.

[0237] Step 1:

[0238] The user speaks aloud to the voice assistant, giving questions or instructions.

[0239] Step 2:

[0240] The device uses a microphone to capture the user's voice and converts the voice into a digital signal using speech recognition technology.

[0241] Step 3:

[0242] The terminal sends the converted digital signal to the server. The transmitted data includes the user's voice content as well as necessary contextual information.

[0243] Step 4:

[0244] The server uses data processing equipment to analyze the received digital signals and identify relevant information. For example, if a user requests weather information, the server will refer to a weather information database to obtain the necessary data.

[0245] Step 5:

[0246] The server converts the acquired information into text data for audio output and sends it to the terminal as a response.

[0247] Step 6:

[0248] The terminal converts the received text data into speech using a speech synthesis system and conveys the information to the user through the speaker.

[0249] Step 7:

[0250] The user takes appropriate action based on the provided voice information. If necessary, they can ask additional questions and restart the process.

[0251] (Example 1)

[0252] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0253] In modern agriculture, older generations find it difficult to utilize digital technology, creating a technological gap with younger generations. Furthermore, insufficient information sharing and utilization of environmental data are leading to decreased efficiency in farm work. To address these challenges, a system is needed that effectively utilizes voice recognition and data analysis while maintaining user-friendly operation for the elderly.

[0254] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0255] In this invention, the server includes a speech recognition means for converting voice input into a digital signal, a data processing means for analyzing the converted digital signal and obtaining relevant information, and a speech synthesis means for converting the obtained information into voice output and providing it to the user. This enables real-time information acquisition and appropriate agricultural work support through a user interface that can be easily operated even by the elderly.

[0256] "Voice recognition means" refers to an element that has the function of converting voice input into a digital signal.

[0257] A "digital signal" is a signal in digital format converted by speech recognition means and analyzed by data processing means.

[0258] A "data processing means" is an element that has the function of analyzing the received digital signal and extracting relevant information.

[0259] A "speech synthesis means" is an element that has the function of converting information acquired by a data processing means into speech output and providing it to the user.

[0260] "Interface means" refers to elements that simplify operation through a user interface and support the use of the device by elderly users.

[0261] An "educational support tool" is an element that has the function of providing agricultural education content in an augmented reality format.

[0262] A "generative model" refers to algorithms and techniques used to generate appropriate responses based on received data.

[0263] "Environmental data management means" refers to elements for acquiring environmental data from sensor devices via an information and communication network and analyzing it using data processing means.

[0264] "Information sharing means" are elements that provide collaborative functions, allow users to share information with each other, and facilitate communication.

[0265] This invention is a comprehensive system for effectively utilizing voice input to provide agricultural information and educational support. The system primarily comprises voice recognition means, data processing means, voice synthesis means, interface means, educational support means, environmental data management means, and information sharing means. This enables both the elderly and young people to receive agricultural support using digital technology.

[0266] The device captures voice input from the user and converts it into a digital signal using speech recognition. This speech recognition uses a general-purpose speech recognition engine (e.g., a speech API) to convert the voice into text data while removing noise.

[0267] The server receives digital data obtained through speech recognition and analyzes it using data processing tools. Here, natural language processing techniques are used to interpret the user's intent, and relevant information is extracted by a generating AI model (e.g., a language model API). The generated information is then converted into natural-sounding speech by speech synthesis tools. For speech synthesis, a speech synthesis engine is used to generate easy-to-understand speech.

[0268] Users can utilize an interface that allows for easy operation through the system. In particular, an intuitive user interface is provided for elderly users to enhance convenience.

[0269] Furthermore, in terms of educational support, augmented reality technology will be used to provide young people with visually and interactively presented educational content about agriculture. This will be achieved by using the camera function of a smartphone to overlay agriculture-related information onto real-world scenery.

[0270] Furthermore, the environmental data management system acquires environmental data in real time from sensor devices via an information and communication network and analyzes it on a server. Based on this information, it becomes possible to provide instructions that support efficient agricultural work.

[0271] Information sharing tools allow users to share information with other users online and facilitate communication. This enables knowledge exchange and collaboration among agricultural workers.

[0272] As a concrete example, when an elderly person speaks to the terminal saying, "Tell me today's weather," the terminal captures the voice and converts it into digital data. The server then analyzes this data to obtain weather information and responds in voice. As an example of a prompt, providing the input "Create a description of the voice recognition system for improving agricultural efficiency" to the generating AI model yields relevant information.

[0273] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0274] Step 1:

[0275] The user speaks into the device, providing voice input. The device captures this voice data with its microphone and filters out background noise to produce clearer audio.

[0276] Step 2:

[0277] The terminal utilizes a speech recognition engine to convert the captured voice into a digital signal. By using the speech API to convert the speech data into text data, the content of the speech is put into a form that can be analyzed. The output of this step is the converted text data.

[0278] Step 3:

[0279] The terminal sends the text data to the server. A secure protocol (e.g., HTTPS) is used for the transmission, and the data is encrypted to ensure the security of the transmitted content.

[0280] Step 4:

[0281] The server analyzes the received text data using data processing means. Using a natural language processing engine, it extracts the intent of the question and keywords, and identifies relevant information. When this information extraction is completed, data for the next step is generated.

[0282] Step 5:

[0283] The server uses a generation AI model to generate an appropriate response based on the extracted information. The generated response is held in the server in text form.

[0284] Step 6:

[0285] The server passes the generated text response to a speech synthesis engine to convert it into audio form. Using the speech synthesis API, a natural and human-like voice is generated. Audio-form data is generated.

[0286] Step 7:

[0287] The terminal receives the audio data from the server and provides an audio output to the user through the speaker. As a result, the user can listen to the answer to the question.

[0288] Step 8:

[0289] Users can ask additional questions as needed, repeating the process from speech recognition to speech synthesis to obtain the necessary information.

[0290] (Application Example 1)

[0291] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0292] Improving work efficiency and bridging the skills gap among workers with varying levels of experience are key challenges in the manufacturing industry. In particular, there is a lack of intuitive work support and training opportunities for older workers and new recruits; therefore, a flexible system is needed to address these issues.

[0293] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0294] In this invention, the server includes speech recognition means for converting acoustic input into an encoded signal, data processing means for analyzing the encoded signal and obtaining related information, and speech synthesis means for converting the obtained information into an acoustic output and providing it to the user. This enables users to operate and obtain information using their voice, allowing elderly people and newcomers to perform factory work efficiently.

[0295] "Auditory input" refers to sound information obtained through hearing, such as the user's voice or ambient sounds.

[0296] An "encoded signal" is signal data obtained by converting a continuous acoustic input into a digital format.

[0297] "Speech recognition means" refers to a device or program for converting an acoustic input into an encoded signal.

[0298] The "data processing means" is a device or program that analyzes an encoded signal and extracts or calculates related information.

[0299] The "voice synthesis means" is a device or program that converts the acquired information back into an acoustic form and provides it to the user.

[0300] The "user interface" is a concept that refers to information display and input means for the user to operate the device or system.

[0301] The "elderly" refers to people in the age group who require support considering physical or cognitive changes as a result of aging.

[0302] The "augmented reality format" is a technical format that overlays virtual information on the real environment for display.

[0303] The "information transmission network" refers to the communication infrastructure for data exchange with remote locations.

[0304] The "Internet of Things device" refers to physical devices and systems interconnected via the Internet.

[0305] The "status data" refers to data related to the operating status and environmental conditions collected from the Internet of Things device.

[0306] The "cooperation function" is part of a system that enables users to support each other and share information.

[0307] The "support means" is a device, method, or tool for realizing efficient work.

[0308] In order to implement this invention, it is necessary to use a combination of various hardware and software. Specifically, this includes a server, a speech recognition device for encoding acoustic input, a programming library for data analysis, and software for synthesizing sound.

[0309] The server receives an encoded signal from a speech recognition device that receives acoustic input. This is done using a common speech recognition API. For example, the Google Cloud Speech-to-Text API can be used to convert acoustic input into text data. The server then analyzes this encoded signal using data processing tools and collects relevant information. Data analysis can be performed using Python programs or data processing libraries such as NumPy and Pandas.

[0310] The acquired information is converted back into an audio format using speech synthesis technology. This can be done using speech synthesis APIs such as Amazon Polly, allowing the information to be presented in a way that is easily understood by the user.

[0311] The user interface should be designed with ease of use for the elderly in mind. It should incorporate an intuitive touch interface and a system that accepts voice commands. For example, if a user says, "Tell me today's work schedule," the system should display the scheduled tasks and provide voice guidance.

[0312] Status data collected via Internet of Things (IoT) devices allows for real-time management of environmental conditions and equipment operating status within the factory, and data analysis enables the proposal of efficient work schedules. This allows for immediate and specific work instructions to be given to qualified workers.

[0313] As a concrete example, let's consider a scenario where a user says, "I want to check the current status of the production line." An example of a prompt for the generating AI model would be, "Please simulate a voice assistant aimed at improving the efficiency of factory operations. This assistant will use speech recognition, data analysis, and speech synthesis to provide the user with real-time production status."

[0314] This invention aims to efficiently and intuitively support work in manufacturing environments, and is designed to be easily operated by all workers, including elderly and new employees.

[0315] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0316] Step 1:

[0317] The device captures the user's audio input and sends it to a speech recognition device. The input is the user's voice, and the output is a digitized encoded signal. This signal is parsed using the Google Cloud Speech-to-Text API, and the audio signal is converted into text.

[0318] Step 2:

[0319] The server receives encoded signals and performs analysis using data processing tools. The input is encoded text data, and the output is relational information. A Python program analyzes the text data and efficiently extracts information using NumPy and Pandas.

[0320] Step 3:

[0321] The server converts the analyzed information back into speech format using speech synthesis technology. The input is the analyzed information, and the output is the synthesized speech. Amazon Polly is used to generate and provide speech in a format that is easy for the user to understand.

[0322] Step 4:

[0323] The device uses speech synthesis output to notify the user of information verbally. Here, the user easily obtains information through the voice interface. The information is delivered at a speed and volume that is easy for elderly people to hear.

[0324] Step 5:

[0325] The user requests status data from an Internet of Things (IoT) device to the server. The input is a voice command from the user, and the server collects this data and performs integrated data analysis.

[0326] Step 6:

[0327] Based on data collected from IoT devices, the server provides users with detailed information about the current environmental conditions and status of the factory. The output is the user's environmental monitoring information, which is used to facilitate efficient operations. An example of its use in a generated AI model is the prompt message: "Simulate a voice assistant aimed at improving the efficiency of factory work. This assistant uses speech recognition, data analysis, and speech synthesis to provide users with real-time production status."

[0328] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0329] This invention is implemented by a system that includes speech recognition means, data processing means, speech synthesis means, interface means, educational support means, environmental data management means, community formation means, and emotion engine means.

[0330] The emotion engine has the function of recognizing emotions from the user's voice input. When the user asks a question or gives an instruction to the voice assistant, the terminal converts the voice into a digital signal. The converted digital signal is sent to a server and analyzed by the emotion engine. The server identifies the user's emotion based on the characteristics of the voice and generates a response optimized for that emotion through a data processing means. The obtained information is output as voice using a speech synthesis means and delivered to the user.

[0331] The interface is designed to enable intuitive operation, making it easy for elderly users to utilize the system. The educational support system provides new experiences based on the user's emotional state. Educational programs for young people are displayed in augmented reality format via the device's camera function, enabling contextually adaptive learning.

[0332] The environmental data management system analyzes agricultural environmental data collected from IoT devices in real time using an information and communication network, improving user work efficiency. Furthermore, the community building system facilitates online information exchange and communication, strengthening collaboration among agricultural workers.

[0333] For example, if a user says, "I'm not feeling very good right now," the device analyzes that emotion, and the server generates appropriate advice to reduce stress. Similarly, in an educational support scenario, if a user feels frustrated with learning, the learning content is adjusted through an emotion engine to make it more engaging. This system makes it possible to provide efficient agricultural support while being mindful of the user's emotions.

[0334] The following describes the processing flow.

[0335] Step 1:

[0336] Users speak to the voice assistant, asking questions, giving instructions, and expressing their current feelings.

[0337] Step 2:

[0338] The device uses a microphone to capture the user's voice and uses speech recognition to convert the voice into a digital signal.

[0339] Step 3:

[0340] The terminal sends the converted digital signal to the server. In addition to the audio content, it also includes user context information.

[0341] Step 4:

[0342] The server utilizes data processing tools to analyze the voice data. Furthermore, an emotion engine identifies the user's emotions from the voice.

[0343] Step 5:

[0344] The server generates a user-optimized response based on the emotions it identifies. For example, if the user is feeling stressed, it might incorporate relaxation advice.

[0345] Step 6:

[0346] The response generated by the server is converted into audio data using speech synthesis technology and sent to the terminal.

[0347] Step 7:

[0348] The device plays the received audio data and provides a response to the user through the speaker.

[0349] Step 8:

[0350] When users access educational content, the device uses educational support tools to deliver the educational program in an augmented reality format. The content is adjusted to the user's emotions to ensure effective learning.

[0351] (Example 2)

[0352] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0353] Systems utilizing speech recognition technology require more user-friendly interfaces for the elderly and flexible responses that can accommodate the emotions of individual users. Furthermore, improving learning efficiency through the provision of agricultural-related educational content is also crucial. The need for means to enhance information sharing and communication with others is also increasing.

[0354] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0355] In this invention, the server includes processing means for converting acoustic data into digital information, computing means for analyzing the converted digital information and extracting relevant information, and acoustic synthesis means for converting the extracted information into acoustic output and providing it to the user. This enables the recognition of emotions from voice input, the provision of responses optimized for individual users, and the promotion of educational content and communication utilizing augmented reality.

[0356] The "processing department" is the department responsible for converting audio data into digital information.

[0357] The "Computation Department" is a department that has the function of analyzing converted digital information and extracting relevant information.

[0358] The "Sound Synthesis Department" is the department responsible for converting extracted information into sound output and providing it to users.

[0359] The "User Interface Department" is a department that simplifies operations through the user interface and provides support for elderly users.

[0360] The "Educational Support Department" is a department that has the function of providing educational materials related to agriculture in augmented reality format.

[0361] The "Emotion Analysis Department" is responsible for recognizing emotions from user voice input and optimizing responses.

[0362] The "Environmental Data Management Department" is a department that has the function of acquiring and analyzing environmental information from sensor devices via an information and communication network.

[0363] The "Collaboration Building Department" is a department that has the function of sharing information with other users and promoting communication.

[0364] This invention is a system that specifically analyzes user voice input, recognizes emotions, and responds appropriately. Specific embodiments for implementing this system are described below.

[0365] The user provides questions or instructions to the system through a voice input device. The terminal receives this voice and, in the first step, converts the voice into a digital signal using speech recognition technology. Speech recognition software is typically used for this purpose. For example, when converting audio data into text data, a speech recognition API can be used.

[0366] The converted digital signal is sent to a server. The server uses an emotion analysis engine to analyze this audio data and identify the user's emotions. Machine learning models and natural language processing are used for the analysis, such as the BERT model. Based on the analysis results, the server understands the content of the voice spoken by the user and generates the optimal response through its data processing functions.

[0367] The generated response is converted back from text to speech using speech synthesis technology. The terminal outputs this speech in a format audible to the user. An example of a prompt is, "Explain how to recognize emotions and generate a response when the user asks about their current feelings." This prompt supports output from a generative AI model that takes user reactions into account.

[0368] For example, if a user says, "I'm feeling down today," the server's emotion analysis engine receives this and generates advice to alleviate stress, outputting it as a voice message from the terminal, such as, "Try taking some deep breaths to relax." This system provides responses that take emotions into account, showing empathy to the user and achieving a more comfortable interaction.

[0369] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0370] Step 1:

[0371] The user speaks into the voice input device, asking questions or giving instructions. The device receives this voice input and uses speech recognition technology to convert the speech into a digital signal. The input is analog audio data, and the output is digital data in text format. This includes the specific operation of converting speech to text using a speech recognition API.

[0372] Step 2:

[0373] The terminal sends the converted digital data to the server. The server sends the received text data to an emotion analysis engine to analyze the user's emotions. It receives text data as input and generates analyzed emotion information as output. Machine learning models are used, and data analysis is performed, particularly using natural language processing techniques.

[0374] Step 3:

[0375] The server generates an appropriate response using data processing functions based on the analysis results. The input is analyzed emotional information, and the output is the text data of the response. The server uses a generative AI model to perform specific actions to construct a natural response that is appropriate to the emotion and scenario.

[0376] Step 4:

[0377] The text data of the response returned from the server to the terminal is converted back into speech by a speech synthesis system. The terminal then plays this speech data to the user. The input is the text data of the response, and the output is the speech data. The specific operation in this step is the conversion from text to speech using a speech synthesis engine.

[0378] (Application Example 2)

[0379] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0380] In modern brick-and-mortar stores, customer service that can immediately respond to the diverse needs and emotions of customers is required. However, conventional systems struggle to provide responsive service that is tailored to the customer's emotions and situation, necessitating further efficiency and personalization. To solve this problem, a system is needed that can recognize customer emotions and provide the most appropriate response.

[0381] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0382] In this invention, the server includes speech recognition means for converting voice input into a digital signal, data processing means for analyzing the converted digital signal and obtaining relevant information, and speech synthesis means for converting the obtained information into voice output and providing it to the user. This makes it possible to analyze the emotions of a customer in real time from their voice and provide an optimal customer service response that corresponds to those emotions.

[0383] "Speech recognition means" refers to a device or process that converts speech into a digital signal.

[0384] "Data processing means" refers to a function or device for analyzing digital signals and obtaining related information.

[0385] "Speech synthesis means" refers to a technology or device for converting acquired information into speech output and providing it to the user.

[0386] "Interface means" refers to a design or device that simplifies operation through a user interface and assists elderly users in using the device.

[0387] "Customer service support tools" refer to functions or systems that support customer service in stores and generate optimal responses according to the customer's emotions.

[0388] "Environmental data management means" refers to a technology or apparatus for acquiring environmental data from IoT devices via an information and communication network and analyzing it using data processing means.

[0389] "Community building means" refers to functions or systems that provide community features, facilitate information sharing among users, and promote communication.

[0390] The system realizing this invention primarily utilizes means for speech recognition, data processing, speech synthesis, interface, customer service support, environmental data management, and community building. It acquires the user's voice through smart glasses or other voice input devices and converts it into a digital signal in real time. Cloud-based speech recognition services such as Google Cloud Speech-to-Text are used for speech recognition. This digital signal is sent to a server and analyzed by sentiment analysis software such as IBM Watson Natural Language Understanding. Based on the analyzed sentiment data, an appropriate response or customer service method is generated using an OpenAI GPT-based model. This response is then provided to the user via speech synthesis.

[0391] Furthermore, if a user is wearing smart glasses in the store, the display will show customer service recommendations. This allows staff to provide service while considering the customer's emotions. For example, if a customer appears tired when asking about a product, the emotion engine will recognize that the customer is tired and suggest products that promote relaxation or encourage a gentler tone of voice.

[0392] For example, when a user says, "I'm looking for something to help me relax," the server can use the information that "the customer seems to want to relax" to generate recommendations for products with relaxation effects. An example of a prompt to the generating AI model would be, "The customer is asking about products with relaxation effects. What do you recommend?"

[0393] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0394] Step 1:

[0395] The device acquires the user's voice. The input is the user's voice, which is collected using the microphone of smart glasses or a smartphone. The output is real-time audio data. This audio data is temporarily stored within the device.

[0396] Step 2:

[0397] The device converts the audio data into a digital signal. The input is the audio data collected in step 1, which is then converted into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text). The output is text data. This converted digital signal is immediately sent to the server.

[0398] Step 3:

[0399] The server analyzes the text data to identify the user's emotions. The input is the text data generated in step 2, which is analyzed using emotion analysis software (e.g., IBM Watson Natural Language Understanding). The output is data indicating the user's emotions. Further analysis is performed based on this data.

[0400] Step 4:

[0401] The server generates optimal responses and product recommendations based on the analysis results. The input is the sentiment data obtained in step 3, and a generative AI model (e.g., OpenAI's GPT-based model) is used to generate the response. The output is the content and conversation suggested to the user. This generated response is used in the next step.

[0402] Step 5:

[0403] The device synthesizes the generated response into speech and provides it to the user. The input is the response data generated in step 4, which is converted into speech output using speech synthesis technology. The output is the content provided to the user aurally. As a result, the user can receive a response that aligns with their own emotions.

[0404] Step 6:

[0405] Visual information is displayed on the user's display (such as the screen of smart glasses). The input is the information generated in step 4, and the content is displayed on the display screen at the appropriate time. The output is the visually presented information. This display allows the user to enjoy a deeper customer service experience.

[0406] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0407] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0408] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0409] [Third Embodiment]

[0410] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0411] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0412] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0413] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0414] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0415] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0416] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0417] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0418] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0419] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0420] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0421] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0422] This invention is implemented by a system comprising speech recognition means, data processing means, speech synthesis means, interface means, educational support means, environmental data management means, and community formation means.

[0423] The voice recognition system has the function of converting the user's voice input into a digital signal. When the user asks a question to the voice assistant, the device captures and digitizes the voice. This digital data is sent to a server and analyzed by data processing equipment.

[0424] The server uses data processing equipment to analyze the received digital signals and extract relevant information. This information may include specific agricultural data or data for work support. This information is then converted into a speech format via speech synthesis equipment and provided to the user.

[0425] The interface provides a simple user interface, making it easy for elderly users to operate. This allows users to operate the device intuitively.

[0426] The educational support measures will provide educational programs for young people that utilize augmented reality (AR) technology. The devices will use their camera functions to visually display agricultural equipment and work procedures overlaid on real-world scenery, providing an interactive learning experience.

[0427] The environmental data management system supports the efficiency of agricultural work by collecting agricultural environmental data in real time via IoT devices and analyzing it on a server. This enables timely work instructions.

[0428] The community-building tools provide users with the ability to share information online and communicate with other farmers. This enables knowledge sharing and community revitalization.

[0429] For example, if an elderly user asks the device, "What's the weather like today?", the device recognizes the voice and sends data to the server. The server collects the current weather information and responds to the device verbally. Furthermore, when younger users use an AR educational program on their smartphones, the system can provide visual instruction by overlaying instructions on how to use virtual agricultural machinery onto the real-world environment. In this way, the system bridges the technical gap among agricultural workers and enables more efficient farming.

[0430] The following describes the processing flow.

[0431] Step 1:

[0432] The user speaks aloud to the voice assistant, giving questions or instructions.

[0433] Step 2:

[0434] The device uses a microphone to capture the user's voice and converts the voice into a digital signal using speech recognition technology.

[0435] Step 3:

[0436] The terminal sends the converted digital signal to the server. The transmitted data includes the user's voice content as well as necessary contextual information.

[0437] Step 4:

[0438] The server uses data processing equipment to analyze the received digital signals and identify relevant information. For example, if a user requests weather information, the server will refer to a weather information database to obtain the necessary data.

[0439] Step 5:

[0440] The server converts the acquired information into text data for audio output and sends it to the terminal as a response.

[0441] Step 6:

[0442] The terminal converts the received text data into speech using a speech synthesis system and conveys the information to the user through the speaker.

[0443] Step 7:

[0444] The user takes appropriate action based on the provided voice information. If necessary, they can ask additional questions and restart the process.

[0445] (Example 1)

[0446] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0447] In modern agriculture, older generations find it difficult to utilize digital technology, creating a technological gap with younger generations. Furthermore, insufficient information sharing and utilization of environmental data are leading to decreased efficiency in farm work. To address these challenges, a system is needed that effectively utilizes voice recognition and data analysis while maintaining user-friendly operation for the elderly.

[0448] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0449] In this invention, the server includes a speech recognition means for converting voice input into a digital signal, a data processing means for analyzing the converted digital signal and obtaining relevant information, and a speech synthesis means for converting the obtained information into voice output and providing it to the user. This enables real-time information acquisition and appropriate agricultural work support through a user interface that can be easily operated even by the elderly.

[0450] "Voice recognition means" refers to an element that has the function of converting voice input into a digital signal.

[0451] A "digital signal" is a signal in digital format converted by speech recognition means and analyzed by data processing means.

[0452] A "data processing means" is an element that has the function of analyzing the received digital signal and extracting relevant information.

[0453] A "speech synthesis means" is an element that has the function of converting information acquired by a data processing means into speech output and providing it to the user.

[0454] "Interface means" refers to elements that simplify operation through a user interface and support the use of the device by elderly users.

[0455] An "educational support tool" is an element that has the function of providing agricultural education content in an augmented reality format.

[0456] A "generative model" refers to algorithms and techniques used to generate appropriate responses based on received data.

[0457] "Environmental data management means" refers to elements for acquiring environmental data from sensor devices via an information and communication network and analyzing it using data processing means.

[0458] "Information sharing means" are elements that provide collaborative functions, allow users to share information with each other, and facilitate communication.

[0459] This invention is a comprehensive system for effectively utilizing voice input to provide agricultural information and educational support. The system primarily comprises voice recognition means, data processing means, voice synthesis means, interface means, educational support means, environmental data management means, and information sharing means. This enables both the elderly and young people to receive agricultural support using digital technology.

[0460] The device captures voice input from the user and converts it into a digital signal using speech recognition. This speech recognition uses a general-purpose speech recognition engine (e.g., a speech API) to convert the voice into text data while removing noise.

[0461] The server receives digital data obtained through speech recognition and analyzes it using data processing tools. Here, natural language processing techniques are used to interpret the user's intent, and relevant information is extracted by a generating AI model (e.g., a language model API). The generated information is then converted into natural-sounding speech by speech synthesis tools. For speech synthesis, a speech synthesis engine is used to generate easy-to-understand speech.

[0462] Users can utilize an interface that allows for easy operation through the system. In particular, an intuitive user interface is provided for elderly users to enhance convenience.

[0463] Furthermore, in terms of educational support, augmented reality technology will be used to provide young people with visually and interactively presented educational content about agriculture. This will be achieved by using the camera function of a smartphone to overlay agriculture-related information onto real-world scenery.

[0464] Furthermore, the environmental data management system acquires environmental data in real time from sensor devices via an information and communication network and analyzes it on a server. Based on this information, it becomes possible to provide instructions that support efficient agricultural work.

[0465] Information sharing tools allow users to share information with other users online and facilitate communication. This enables knowledge exchange and collaboration among agricultural workers.

[0466] As a concrete example, when an elderly person speaks to the terminal saying, "Tell me today's weather," the terminal captures the voice and converts it into digital data. The server then analyzes this data to obtain weather information and responds in voice. As an example of a prompt, providing the input "Create a description of the voice recognition system for improving agricultural efficiency" to the generating AI model yields relevant information.

[0467] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0468] Step 1:

[0469] The user speaks into the device, providing voice input. The device captures this voice data with its microphone and filters out background noise to produce clearer audio.

[0470] Step 2:

[0471] The device uses a speech recognition engine to convert the captured audio into a digital signal. Using a speech API, the audio data is converted into text data, making the audio content analyzable. The output of this step is the converted text data.

[0472] Step 3:

[0473] The terminal sends text data to the server. A secure protocol (e.g., HTTPS) is used for transmission, and the data is encrypted to ensure the security of the transmitted content.

[0474] Step 4:

[0475] The server analyzes the received text data using data processing tools. A natural language processing engine is used to extract the intent of the question and keywords, and to identify relevant information. Once this information extraction is complete, data for the next step is generated.

[0476] Step 5:

[0477] The server uses a generative AI model to generate an appropriate response based on the extracted information. The generated response is stored on the server in text format.

[0478] Step 6:

[0479] The server passes the generated text response to the speech synthesis engine, which converts it into speech format. Using the speech synthesis API, it generates natural, human-like speech. Speech format data is then generated.

[0480] Step 7:

[0481] The terminal receives audio data from the server and provides audio output to the user through its speaker. This allows the user to hear the answer to their question.

[0482] Step 8:

[0483] Users can ask additional questions as needed, repeating the process from speech recognition to speech synthesis to obtain the necessary information.

[0484] (Application Example 1)

[0485] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0486] Improving work efficiency and bridging the skills gap among workers with varying levels of experience are key challenges in the manufacturing industry. In particular, there is a lack of intuitive work support and training opportunities for older workers and new recruits; therefore, a flexible system is needed to address these issues.

[0487] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0488] In this invention, the server includes speech recognition means for converting acoustic input into an encoded signal, data processing means for analyzing the encoded signal and obtaining related information, and speech synthesis means for converting the obtained information into an acoustic output and providing it to the user. This enables users to operate and obtain information using their voice, allowing elderly people and newcomers to perform factory work efficiently.

[0489] "Auditory input" refers to sound information obtained through hearing, such as the user's voice or ambient sounds.

[0490] An "encoded signal" is signal data obtained by converting a continuous acoustic input into a digital format.

[0491] "Speech recognition means" refers to a device or program for converting an acoustic input into an encoded signal.

[0492] "Data processing means" refers to a device or program that analyzes an encoded signal and extracts or calculates related information.

[0493] "Speech synthesis means" refers to a device or program that converts acquired information back into an acoustic format and provides it to the user.

[0494] "User interface" is a concept that refers to the means of displaying or inputting information that a user uses to operate a device or system.

[0495] "Elderly people" refers to people of an age group who require support, particularly considering physical or cognitive changes that result from aging.

[0496] Augmented reality is a technological format that overlays virtual information onto the real environment.

[0497] An "information transmission network" is a communication infrastructure used for exchanging data with remote locations.

[0498] "Internet of Things devices" refer to physical devices and systems that are interconnected via the internet.

[0499] "Status data" refers to data about the operating status and environmental conditions collected from Internet of Things devices.

[0500] "Cooperative features" are part of a system that allows users to support each other and share information.

[0501] "Support measures" refer to devices, methods, or tools that enable efficient work.

[0502] In order to implement this invention, it is necessary to use a combination of various hardware and software. Specifically, these include a server, a speech recognition device for encoding acoustic input, a programming library for data analysis, and software for synthesizing sound.

[0503] The server receives an encoded signal from a speech recognition device that receives acoustic input. This is done using a common speech recognition API. For example, the Google Cloud Speech-to-Text API can be used to convert acoustic input into text data. The server then analyzes this encoded signal using data processing tools and collects relevant information. Data analysis can be performed using Python programs or data processing libraries such as NumPy and Pandas.

[0504] The acquired information is converted back into an audio format using speech synthesis technology. This can be done using speech synthesis APIs such as Amazon Polly, allowing the information to be presented in a user-friendly format.

[0505] The user interface should be designed with ease of use for the elderly in mind. It should incorporate an intuitive touch interface and a system that accepts voice commands. For example, if a user says, "Tell me today's work schedule," the system should display the scheduled tasks and provide voice guidance.

[0506] Status data collected via Internet of Things (IoT) devices allows for real-time management of environmental conditions and equipment operating status within the factory, and data analysis enables the proposal of efficient work schedules. This allows for immediate and specific work instructions to be given to qualified workers.

[0507] As a concrete example, let's consider a scenario where a user says, "I want to check the current status of the production line." An example of a prompt for the generating AI model would be, "Please simulate a voice assistant aimed at improving the efficiency of factory operations. This assistant will use speech recognition, data analysis, and speech synthesis to provide the user with real-time production status."

[0508] This invention aims to efficiently and intuitively support work in manufacturing environments, and is designed to be easily operated by all workers, including elderly and new employees.

[0509] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0510] Step 1:

[0511] The device captures the user's audio input and sends it to a speech recognition device. The input is the user's voice, and the output is a digitized encoded signal. This signal is parsed using the Google Cloud Speech-to-Text API, and the audio signal is converted into text.

[0512] Step 2:

[0513] The server receives encoded signals and performs analysis using data processing tools. The input is encoded text data, and the output is relational information. A Python program analyzes the text data and efficiently extracts information using NumPy and Pandas.

[0514] Step 3:

[0515] The server converts the analyzed information back into speech format using speech synthesis technology. The input is the analyzed information, and the output is the synthesized speech. Amazon Polly is used to generate and provide speech in a format that is easy for the user to understand.

[0516] Step 4:

[0517] The device uses speech synthesis output to notify the user of information verbally. Here, the user easily obtains information through the voice interface. The information is delivered at a speed and volume that is easy for elderly people to hear.

[0518] Step 5:

[0519] The user requests status data from an Internet of Things (IoT) device to the server. The input is a voice command from the user, and the server collects this data and performs integrated data analysis.

[0520] Step 6:

[0521] Based on data collected from IoT devices, the server provides users with detailed information about the current environmental conditions and status of the factory. The output is the user's environmental monitoring information, which is used to facilitate efficient operations. An example of its use in a generated AI model is the prompt message: "Simulate a voice assistant aimed at improving the efficiency of factory work. This assistant uses speech recognition, data analysis, and speech synthesis to provide users with real-time production status."

[0522] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0523] This invention is implemented by a system that includes speech recognition means, data processing means, speech synthesis means, interface means, educational support means, environmental data management means, community formation means, and emotion engine means.

[0524] The emotion engine has the function of recognizing emotions from the user's voice input. When the user asks a question or gives an instruction to the voice assistant, the terminal converts the voice into a digital signal. The converted digital signal is sent to a server and analyzed by the emotion engine. The server identifies the user's emotion based on the characteristics of the voice and generates a response optimized for that emotion through a data processing means. The obtained information is output as voice using a speech synthesis means and delivered to the user.

[0525] The interface is designed to enable intuitive operation, making it easy for elderly users to utilize the system. The educational support system provides new experiences based on the user's emotional state. Educational programs for young people are displayed in augmented reality format via the device's camera function, enabling contextually adaptive learning.

[0526] The environmental data management system analyzes agricultural environmental data collected from IoT devices in real time using an information and communication network, improving user work efficiency. Furthermore, the community building system facilitates online information exchange and communication, strengthening collaboration among agricultural workers.

[0527] For example, if a user says, "I'm not feeling very good right now," the device analyzes that emotion, and the server generates appropriate advice to reduce stress. Similarly, in an educational support scenario, if a user feels frustrated with learning, the learning content is adjusted through an emotion engine to make it more engaging. This system makes it possible to provide efficient agricultural support while being mindful of the user's emotions.

[0528] The following describes the processing flow.

[0529] Step 1:

[0530] Users speak to the voice assistant, asking questions, giving instructions, and expressing their current feelings.

[0531] Step 2:

[0532] The device uses a microphone to capture the user's voice and uses speech recognition to convert the voice into a digital signal.

[0533] Step 3:

[0534] The terminal sends the converted digital signal to the server. In addition to the audio content, it also includes user context information.

[0535] Step 4:

[0536] The server utilizes data processing tools to analyze the voice data. Furthermore, an emotion engine identifies the user's emotions from the voice.

[0537] Step 5:

[0538] The server generates a user-optimized response based on the emotions it identifies. For example, if the user is feeling stressed, it might incorporate relaxation advice.

[0539] Step 6:

[0540] The response generated by the server is converted into audio data using speech synthesis technology and sent to the terminal.

[0541] Step 7:

[0542] The device plays the received audio data and provides a response to the user through the speaker.

[0543] Step 8:

[0544] When users access educational content, the device uses educational support tools to deliver the educational program in an augmented reality format. The content is adjusted to the user's emotions to ensure effective learning.

[0545] (Example 2)

[0546] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0547] Systems utilizing speech recognition technology require more user-friendly interfaces for the elderly and flexible responses that can accommodate the emotions of individual users. Furthermore, improving learning efficiency through the provision of agricultural-related educational content is also crucial. The need for means to enhance information sharing and communication with others is also increasing.

[0548] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0549] In this invention, the server includes processing means for converting acoustic data into digital information, computing means for analyzing the converted digital information and extracting relevant information, and acoustic synthesis means for converting the extracted information into acoustic output and providing it to the user. This enables the recognition of emotions from voice input, the provision of responses optimized for individual users, and the promotion of educational content and communication utilizing augmented reality.

[0550] The "processing department" is the department responsible for converting audio data into digital information.

[0551] The "Computation Department" is a department that has the function of analyzing converted digital information and extracting relevant information.

[0552] The "Sound Synthesis Department" is the department responsible for converting extracted information into sound output and providing it to users.

[0553] The "User Interface Department" is a department that simplifies operations through the user interface and provides support for elderly users.

[0554] The "Educational Support Department" is a department that has the function of providing educational materials related to agriculture in augmented reality format.

[0555] The "Emotion Analysis Department" is responsible for recognizing emotions from user voice input and optimizing responses.

[0556] The "Environmental Data Management Department" is a department that has the function of acquiring and analyzing environmental information from sensor devices via an information and communication network.

[0557] The "Collaboration Building Department" is a department that has the function of sharing information with other users and promoting communication.

[0558] This invention is a system that specifically analyzes user voice input, recognizes emotions, and responds appropriately. Specific embodiments for implementing this system are described below.

[0559] The user provides questions or instructions to the system through a voice input device. The terminal receives this voice and, in the first step, converts the voice into a digital signal using speech recognition technology. Speech recognition software is typically used for this purpose. For example, when converting audio data into text data, a speech recognition API can be used.

[0560] The converted digital signal is sent to a server. The server uses an emotion analysis engine to analyze this audio data and identify the user's emotions. Machine learning models and natural language processing are used for the analysis, such as the BERT model. Based on the analysis results, the server understands the content of the voice spoken by the user and generates the optimal response through its data processing functions.

[0561] The generated response is converted back from text to speech using speech synthesis technology. The terminal outputs this speech in a format audible to the user. An example of a prompt is, "Explain how to recognize emotions and generate a response when the user asks about their current feelings." This prompt supports output from a generative AI model that takes user reactions into account.

[0562] For example, if a user says, "I'm feeling down today," the server's emotion analysis engine receives this and generates advice to alleviate stress, outputting it as a voice message from the terminal, such as, "Try taking some deep breaths to relax." This system provides responses that take emotions into account, showing empathy to the user and achieving a more comfortable interaction.

[0563] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0564] Step 1:

[0565] The user speaks into the voice input device, asking questions or giving instructions. The device receives this voice input and uses speech recognition technology to convert the speech into a digital signal. The input is analog audio data, and the output is digital data in text format. This includes the specific operation of converting speech to text using a speech recognition API.

[0566] Step 2:

[0567] The terminal sends the converted digital data to the server. The server sends the received text data to an emotion analysis engine to analyze the user's emotions. It receives text data as input and generates analyzed emotion information as output. Machine learning models are used, and data analysis is performed, particularly using natural language processing techniques.

[0568] Step 3:

[0569] The server generates an appropriate response using data processing functions based on the analysis results. The input is analyzed emotional information, and the output is the text data of the response. The server uses a generative AI model to perform specific actions to construct a natural response that is appropriate to the emotion and scenario.

[0570] Step 4:

[0571] The text data of the response returned from the server to the terminal is converted back into speech by a speech synthesis system. The terminal then plays this speech data to the user. The input is the text data of the response, and the output is the speech data. The specific operation in this step is the conversion from text to speech using a speech synthesis engine.

[0572] (Application Example 2)

[0573] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0574] In modern brick-and-mortar stores, customer service that can immediately respond to the diverse needs and emotions of customers is required. However, conventional systems struggle to provide responsive service that is tailored to the customer's emotions and situation, necessitating further efficiency and personalization. To solve this problem, a system is needed that can recognize customer emotions and provide the most appropriate response.

[0575] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0576] In this invention, the server includes speech recognition means for converting voice input into a digital signal, data processing means for analyzing the converted digital signal and obtaining relevant information, and speech synthesis means for converting the obtained information into voice output and providing it to the user. This makes it possible to analyze the emotions of a customer in real time from their voice and provide an optimal customer service response that corresponds to those emotions.

[0577] "Speech recognition means" refers to a device or process that converts speech into a digital signal.

[0578] "Data processing means" refers to a function or device for analyzing digital signals and obtaining related information.

[0579] "Speech synthesis means" refers to a technology or device for converting acquired information into speech output and providing it to the user.

[0580] "Interface means" refers to a design or device that simplifies operation through a user interface and assists elderly users in using the device.

[0581] "Customer service support tools" refer to functions or systems that support customer service in stores and generate optimal responses according to the customer's emotions.

[0582] "Environmental data management means" refers to a technology or device for acquiring environmental data from IoT devices via an information and communication network and analyzing it using data processing means.

[0583] "Community building means" refers to functions or systems that provide community features, allow users to share information with each other, and facilitate communication.

[0584] The system realizing this invention primarily utilizes means for speech recognition, data processing, speech synthesis, interface, customer service support, environmental data management, and community building. It acquires the user's voice through smart glasses or other voice input devices and converts it into a digital signal in real time. Cloud-based speech recognition services such as Google Cloud Speech-to-Text are used for speech recognition. This digital signal is sent to a server and analyzed by sentiment analysis software such as IBM Watson Natural Language Understanding. Based on the analyzed sentiment data, an appropriate response or customer service method is generated using an OpenAI GPT-based model. This response is then provided to the user via speech synthesis.

[0585] Furthermore, if a user is wearing smart glasses in the store, the display will show customer service recommendations. This allows staff to provide service while considering the customer's emotions. For example, if a customer appears tired when asking about a product, the emotion engine will recognize that the customer is tired and suggest products that promote relaxation or encourage a gentler tone of voice.

[0586] For example, when a user says, "I'm looking for something to help me relax," the server can use the information that "the customer seems to want to relax" to generate recommendations for products with relaxation effects. An example of a prompt to the generating AI model would be, "The customer is asking about products with relaxation effects. What do you recommend?"

[0587] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0588] Step 1:

[0589] The device acquires the user's voice. The input is the user's voice, which is collected using the microphone of smart glasses or a smartphone. The output is real-time audio data. This audio data is temporarily stored within the device.

[0590] Step 2:

[0591] The device converts the audio data into a digital signal. The input is the audio data collected in step 1, which is then converted into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text). The output is text data. This converted digital signal is immediately sent to the server.

[0592] Step 3:

[0593] The server analyzes the text data to identify the user's emotions. The input is the text data generated in step 2, which is analyzed using emotion analysis software (e.g., IBM Watson Natural Language Understanding). The output is data indicating the user's emotions. Further analysis is then performed based on this data.

[0594] Step 4:

[0595] The server generates optimal responses and product recommendations based on the analysis results. The input is the sentiment data obtained in step 3, and a generative AI model (e.g., OpenAI's GPT-based model) is used to generate the response. The output is the content and conversation suggested to the user. This generated response is used in the next step.

[0596] Step 5:

[0597] The device synthesizes the generated response into speech and provides it to the user. The input is the response data generated in step 4, which is converted into speech output using speech synthesis technology. The output is the content provided to the user aurally. As a result, the user can receive a response that aligns with their own emotions.

[0598] Step 6:

[0599] Visual information is displayed on the user's display (such as the screen of smart glasses). The input is the information generated in step 4, and the content is displayed on the display screen at the appropriate time. The output is the visually presented information. This display allows the user to enjoy a deeper customer service experience.

[0600] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0601] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0602] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0603] [Fourth Embodiment]

[0604] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0605] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0606] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0607] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0608] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0609] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0610] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0611] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0612] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0613] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0614] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0615] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0616] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0617] This invention is implemented by a system comprising speech recognition means, data processing means, speech synthesis means, interface means, educational support means, environmental data management means, and community formation means.

[0618] The voice recognition system has the function of converting the user's voice input into a digital signal. When the user asks a question to the voice assistant, the device captures and digitizes the voice. This digital data is sent to a server and analyzed by data processing equipment.

[0619] The server uses data processing equipment to analyze the received digital signals and extract relevant information. This information may include specific agricultural data or data for work support. This information is then converted into a speech format via speech synthesis equipment and provided to the user.

[0620] The interface provides a simple user interface, making it easy for elderly users to operate. This allows users to operate the device intuitively.

[0621] The educational support measures will provide educational programs for young people that utilize augmented reality (AR) technology. The devices will use their camera functions to visually display agricultural equipment and work procedures overlaid on real-world scenery, providing an interactive learning experience.

[0622] The environmental data management system supports the efficiency of agricultural work by collecting agricultural environmental data in real time via IoT devices and analyzing it on a server. This enables timely work instructions.

[0623] The community-building tools provide users with the ability to share information online and communicate with other farmers. This enables knowledge sharing and community revitalization.

[0624] For example, if an elderly user asks the device, "What's the weather like today?", the device recognizes the voice and sends data to the server. The server collects the current weather information and responds to the device verbally. Furthermore, when younger users use an AR educational program on their smartphones, the system can provide visual instruction by overlaying instructions on how to use virtual agricultural machinery onto the real-world environment. In this way, the system bridges the technical gap among agricultural workers and enables more efficient farming.

[0625] The following describes the processing flow.

[0626] Step 1:

[0627] The user speaks aloud to the voice assistant, giving questions or instructions.

[0628] Step 2:

[0629] The device uses a microphone to capture the user's voice and converts the voice into a digital signal using speech recognition technology.

[0630] Step 3:

[0631] The terminal sends the converted digital signal to the server. The transmitted data includes the user's voice content as well as necessary contextual information.

[0632] Step 4:

[0633] The server uses data processing equipment to analyze the received digital signals and identify relevant information. For example, if a user requests weather information, the server will refer to a weather information database to obtain the necessary data.

[0634] Step 5:

[0635] The server converts the acquired information into text data for audio output and sends it to the terminal as a response.

[0636] Step 6:

[0637] The terminal converts the received text data into speech using a speech synthesis system and conveys the information to the user through the speaker.

[0638] Step 7:

[0639] The user takes appropriate action based on the provided voice information. If necessary, they can ask additional questions and restart the process.

[0640] (Example 1)

[0641] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0642] In modern agriculture, older generations find it difficult to utilize digital technology, creating a technological gap with younger generations. Furthermore, insufficient information sharing and utilization of environmental data are leading to decreased efficiency in farm work. To address these challenges, a system is needed that effectively utilizes voice recognition and data analysis while maintaining user-friendly operation for the elderly.

[0643] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0644] In this invention, the server includes a speech recognition means for converting voice input into a digital signal, a data processing means for analyzing the converted digital signal and obtaining relevant information, and a speech synthesis means for converting the obtained information into voice output and providing it to the user. This enables real-time information acquisition and appropriate agricultural work support through a user interface that can be easily operated even by the elderly.

[0645] "Voice recognition means" refers to an element that has the function of converting voice input into a digital signal.

[0646] A "digital signal" is a signal in digital format converted by speech recognition means and analyzed by data processing means.

[0647] A "data processing means" is an element that has the function of analyzing the received digital signal and extracting relevant information.

[0648] A "speech synthesis means" is an element that has the function of converting information acquired by a data processing means into speech output and providing it to the user.

[0649] "Interface means" refers to elements that simplify operation through a user interface and support the use of the device by elderly users.

[0650] An "educational support tool" is an element that has the function of providing agricultural education content in an augmented reality format.

[0651] A "generative model" refers to algorithms and techniques used to generate appropriate responses based on received data.

[0652] "Environmental data management means" refers to elements for acquiring environmental data from sensor devices via an information and communication network and analyzing it using data processing means.

[0653] "Information sharing means" are elements that provide collaborative functions, allow users to share information with each other, and facilitate communication.

[0654] This invention is a comprehensive system for effectively utilizing voice input to provide agricultural information and educational support. The system primarily comprises voice recognition means, data processing means, voice synthesis means, interface means, educational support means, environmental data management means, and information sharing means. This enables both the elderly and young people to receive agricultural support using digital technology.

[0655] The device captures voice input from the user and converts it into a digital signal using speech recognition. This speech recognition uses a general-purpose speech recognition engine (e.g., a speech API) to convert the voice into text data while removing noise.

[0656] The server receives digital data obtained through speech recognition and analyzes it using data processing tools. Here, natural language processing techniques are used to interpret the user's intent, and relevant information is extracted by a generating AI model (e.g., a language model API). The generated information is then converted into natural-sounding speech by speech synthesis tools. For speech synthesis, a speech synthesis engine is used to generate easy-to-understand speech.

[0657] Users can utilize an interface that allows for easy operation through the system. In particular, an intuitive user interface is provided for elderly users to enhance convenience.

[0658] Furthermore, in terms of educational support, augmented reality technology will be used to provide young people with visually and interactively presented educational content about agriculture. This will be achieved by using the camera function of a smartphone to overlay agriculture-related information onto real-world scenery.

[0659] Furthermore, the environmental data management system acquires environmental data in real time from sensor devices via an information and communication network and analyzes it on a server. Based on this information, it becomes possible to provide instructions that support efficient agricultural work.

[0660] Information sharing tools allow users to share information with other users online and facilitate communication. This enables knowledge exchange and collaboration among agricultural workers.

[0661] As a concrete example, when an elderly person speaks to the terminal saying, "Tell me today's weather," the terminal captures the voice and converts it into digital data. The server then analyzes this data to obtain weather information and responds in voice. As an example of a prompt, providing the input "Create a description of the voice recognition system for improving agricultural efficiency" to the generating AI model yields relevant information.

[0662] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0663] Step 1:

[0664] The user speaks into the device, providing voice input. The device captures this voice data with its microphone and filters out background noise to produce clearer audio.

[0665] Step 2:

[0666] The device uses a speech recognition engine to convert the captured audio into a digital signal. Using a speech API, the audio data is converted into text data, making the audio content analyzable. The output of this step is the converted text data.

[0667] Step 3:

[0668] The terminal sends text data to the server. A secure protocol (e.g., HTTPS) is used for transmission, and the data is encrypted to ensure the security of the transmitted content.

[0669] Step 4:

[0670] The server analyzes the received text data using data processing tools. A natural language processing engine is used to extract the intent of the question and keywords, and to identify relevant information. Once this information extraction is complete, data for the next step is generated.

[0671] Step 5:

[0672] The server uses a generative AI model to generate an appropriate response based on the extracted information. The generated response is stored on the server in text format.

[0673] Step 6:

[0674] The server passes the generated text response to the speech synthesis engine, which converts it into speech format. Using the speech synthesis API, it generates natural, human-like speech. Speech format data is then generated.

[0675] Step 7:

[0676] The terminal receives audio data from the server and provides audio output to the user through its speaker. This allows the user to hear the answer to their question.

[0677] Step 8:

[0678] Users can ask additional questions as needed, repeating the process from speech recognition to speech synthesis to obtain the necessary information.

[0679] (Application Example 1)

[0680] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0681] Improving work efficiency and bridging the skills gap among workers with varying levels of experience are key challenges in the manufacturing industry. In particular, there is a lack of intuitive work support and training opportunities for older workers and new recruits; therefore, a flexible system is needed to address these issues.

[0682] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0683] In this invention, the server includes speech recognition means for converting acoustic input into an encoded signal, data processing means for analyzing the encoded signal and obtaining related information, and speech synthesis means for converting the obtained information into an acoustic output and providing it to the user. This enables users to operate and obtain information using their voice, allowing elderly people and newcomers to perform factory work efficiently.

[0684] "Auditory input" refers to sound information obtained through hearing, such as the user's voice or ambient sounds.

[0685] An "encoded signal" is signal data obtained by converting a continuous acoustic input into a digital format.

[0686] "Speech recognition means" refers to a device or program for converting an acoustic input into an encoded signal.

[0687] "Data processing means" refers to a device or program that analyzes an encoded signal and extracts or calculates related information.

[0688] "Speech synthesis means" refers to a device or program that converts acquired information back into an acoustic format and provides it to the user.

[0689] "User interface" is a concept that refers to the means of displaying or inputting information that a user uses to operate a device or system.

[0690] "Elderly people" refers to people of an age group who require support, particularly considering physical or cognitive changes that result from aging.

[0691] Augmented reality is a technological format that overlays virtual information onto the real environment.

[0692] An "information transmission network" is a communication infrastructure used for exchanging data with remote locations.

[0693] "Internet of Things devices" refer to physical devices and systems that are interconnected via the internet.

[0694] "Status data" refers to data about the operating status and environmental conditions collected from Internet of Things devices.

[0695] "Cooperative features" are part of a system that allows users to support each other and share information.

[0696] "Support measures" refer to devices, methods, or tools that enable efficient work.

[0697] In order to implement this invention, it is necessary to use a combination of various hardware and software. Specifically, these include a server, a speech recognition device for encoding acoustic input, a programming library for data analysis, and software for synthesizing sound.

[0698] The server receives an encoded signal from a speech recognition device that receives acoustic input. This is done using a common speech recognition API. For example, the Google Cloud Speech-to-Text API can be used to convert acoustic input into text data. The server then analyzes this encoded signal using data processing tools and collects relevant information. Data analysis can be performed using Python programs or data processing libraries such as NumPy and Pandas.

[0699] The acquired information is converted back into an audio format using speech synthesis technology. This can be done using speech synthesis APIs such as Amazon Polly, allowing the information to be presented in a user-friendly format.

[0700] The user interface should be designed with ease of use for the elderly in mind. It should incorporate an intuitive touch interface and a system that accepts voice commands. For example, if a user says, "Tell me today's work schedule," the system should display the scheduled tasks and provide voice guidance.

[0701] Status data collected via Internet of Things (IoT) devices allows for real-time management of environmental conditions and equipment operating status within the factory, and data analysis enables the proposal of efficient work schedules. This allows for immediate and specific work instructions to be given to qualified workers.

[0702] As a concrete example, let's consider a scenario where a user says, "I want to check the current status of the production line." An example of a prompt for the generating AI model would be, "Please simulate a voice assistant aimed at improving the efficiency of factory operations. This assistant will use speech recognition, data analysis, and speech synthesis to provide the user with real-time production status."

[0703] This invention aims to efficiently and intuitively support work in manufacturing environments, and is designed to be easily operated by all workers, including elderly and new employees.

[0704] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0705] Step 1:

[0706] The device captures the user's audio input and sends it to a speech recognition device. The input is the user's voice, and the output is a digitized encoded signal. This signal is parsed using the Google Cloud Speech-to-Text API, and the audio signal is converted into text.

[0707] Step 2:

[0708] The server receives encoded signals and performs analysis using data processing tools. The input is encoded text data, and the output is relational information. A Python program analyzes the text data and efficiently extracts information using NumPy and Pandas.

[0709] Step 3:

[0710] The server converts the analyzed information back into speech format using speech synthesis technology. The input is the analyzed information, and the output is the synthesized speech. Amazon Polly is used to generate and provide speech in a format that is easy for the user to understand.

[0711] Step 4:

[0712] The device uses speech synthesis output to notify the user of information verbally. Here, the user easily obtains information through the voice interface. The information is delivered at a speed and volume that is easy for elderly people to hear.

[0713] Step 5:

[0714] The user requests status data from an Internet of Things (IoT) device to the server. The input is a voice command from the user, and the server collects this data and performs integrated data analysis.

[0715] Step 6:

[0716] Based on data collected from IoT devices, the server provides users with detailed information about the current environmental conditions and status of the factory. The output is the user's environmental monitoring information, which is used to facilitate efficient operations. An example of its use in a generated AI model is the prompt message: "Simulate a voice assistant aimed at improving the efficiency of factory work. This assistant uses speech recognition, data analysis, and speech synthesis to provide users with real-time production status."

[0717] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0718] This invention is implemented by a system that includes speech recognition means, data processing means, speech synthesis means, interface means, educational support means, environmental data management means, community formation means, and emotion engine means.

[0719] The emotion engine has the function of recognizing emotions from the user's voice input. When the user asks a question or gives an instruction to the voice assistant, the terminal converts the voice into a digital signal. The converted digital signal is sent to a server and analyzed by the emotion engine. The server identifies the user's emotion based on the characteristics of the voice and generates a response optimized for that emotion through a data processing means. The obtained information is output as voice using a speech synthesis means and delivered to the user.

[0720] The interface is designed to enable intuitive operation, making it easy for elderly users to utilize the system. The educational support system provides new experiences based on the user's emotional state. Educational programs for young people are displayed in augmented reality format via the device's camera function, enabling contextually adaptive learning.

[0721] The environmental data management system analyzes agricultural environmental data collected from IoT devices in real time using an information and communication network, improving user work efficiency. Furthermore, the community building system facilitates online information exchange and communication, strengthening collaboration among agricultural workers.

[0722] For example, if a user says, "I'm not feeling very good right now," the device analyzes that emotion, and the server generates appropriate advice to reduce stress. Similarly, in an educational support scenario, if a user feels frustrated with learning, the learning content is adjusted through an emotion engine to make it more engaging. This system makes it possible to provide efficient agricultural support while being mindful of the user's emotions.

[0723] The following describes the processing flow.

[0724] Step 1:

[0725] Users speak to the voice assistant, asking questions, giving instructions, and expressing their current feelings.

[0726] Step 2:

[0727] The device uses a microphone to capture the user's voice and uses speech recognition to convert the voice into a digital signal.

[0728] Step 3:

[0729] The terminal sends the converted digital signal to the server. In addition to the audio content, it also includes user context information.

[0730] Step 4:

[0731] The server utilizes data processing tools to analyze the voice data. Furthermore, an emotion engine identifies the user's emotions from the voice.

[0732] Step 5:

[0733] The server generates a user-optimized response based on the emotions it identifies. For example, if the user is feeling stressed, it might incorporate relaxation advice.

[0734] Step 6:

[0735] The response generated by the server is converted into audio data using speech synthesis technology and sent to the terminal.

[0736] Step 7:

[0737] The device plays the received audio data and provides a response to the user through the speaker.

[0738] Step 8:

[0739] When users access educational content, the device uses educational support tools to deliver the educational program in an augmented reality format. The content is adjusted to the user's emotions to ensure effective learning.

[0740] (Example 2)

[0741] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0742] Systems utilizing speech recognition technology require more user-friendly interfaces for the elderly and flexible responses that can accommodate the emotions of individual users. Furthermore, improving learning efficiency through the provision of agricultural-related educational content is also crucial. The need for means to enhance information sharing and communication with others is also increasing.

[0743] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0744] In this invention, the server includes processing means for converting acoustic data into digital information, computing means for analyzing the converted digital information and extracting relevant information, and acoustic synthesis means for converting the extracted information into acoustic output and providing it to the user. This enables the recognition of emotions from voice input, the provision of responses optimized for individual users, and the promotion of educational content and communication utilizing augmented reality.

[0745] The "processing department" is the department responsible for converting audio data into digital information.

[0746] The "Computation Department" is a department that has the function of analyzing converted digital information and extracting relevant information.

[0747] The "Sound Synthesis Department" is the department responsible for converting extracted information into sound output and providing it to users.

[0748] The "User Interface Department" is a department that simplifies operations through the user interface and provides support for elderly users.

[0749] The "Educational Support Department" is a department that has the function of providing educational materials related to agriculture in augmented reality format.

[0750] The "Emotion Analysis Department" is responsible for recognizing emotions from user voice input and optimizing responses.

[0751] The "Environmental Data Management Department" is a department that has the function of acquiring and analyzing environmental information from sensor devices via an information and communication network.

[0752] The "Collaboration Building Department" is a department that has the function of sharing information with other users and promoting communication.

[0753] This invention is a system that specifically analyzes user voice input, recognizes emotions, and responds appropriately. Specific embodiments for implementing this system are described below.

[0754] The user provides questions or instructions to the system through a voice input device. The terminal receives this voice and, in the first step, converts the voice into a digital signal using speech recognition technology. Speech recognition software is typically used for this purpose. For example, when converting audio data into text data, a speech recognition API can be used.

[0755] The converted digital signal is sent to a server. The server uses an emotion analysis engine to analyze this audio data and identify the user's emotions. Machine learning models and natural language processing are used for the analysis, such as the BERT model. Based on the analysis results, the server understands the content of the voice spoken by the user and generates the optimal response through its data processing functions.

[0756] The generated response is converted back from text to speech using speech synthesis technology. The terminal outputs this speech in a format audible to the user. An example of a prompt is, "Explain how to recognize emotions and generate a response when the user asks about their current feelings." This prompt supports output from a generative AI model that takes user reactions into account.

[0757] For example, if a user says, "I'm feeling down today," the server's emotion analysis engine receives this and generates advice to alleviate stress, outputting it as a voice message from the terminal, such as, "Try taking some deep breaths to relax." This system provides responses that take emotions into account, showing empathy to the user and achieving a more comfortable interaction.

[0758] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0759] Step 1:

[0760] The user speaks into the voice input device, asking questions or giving instructions. The device receives this voice input and uses speech recognition technology to convert the speech into a digital signal. The input is analog audio data, and the output is digital data in text format. This includes the specific operation of converting speech to text using a speech recognition API.

[0761] Step 2:

[0762] The terminal sends the converted digital data to the server. The server sends the received text data to an emotion analysis engine to analyze the user's emotions. It receives text data as input and generates analyzed emotion information as output. Machine learning models are used, and data analysis is performed, particularly using natural language processing techniques.

[0763] Step 3:

[0764] The server generates an appropriate response using data processing functions based on the analysis results. The input is analyzed emotional information, and the output is the text data of the response. The server uses a generative AI model to perform specific actions to construct a natural response that is appropriate to the emotion and scenario.

[0765] Step 4:

[0766] The text data of the response returned from the server to the terminal is converted back into speech by a speech synthesis system. The terminal then plays this speech data to the user. The input is the text data of the response, and the output is the speech data. The specific operation in this step is the conversion from text to speech using a speech synthesis engine.

[0767] (Application Example 2)

[0768] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0769] In modern brick-and-mortar stores, customer service that can immediately respond to the diverse needs and emotions of customers is required. However, conventional systems struggle to provide responsive service that is tailored to the customer's emotions and situation, necessitating further efficiency and personalization. To solve this problem, a system is needed that can recognize customer emotions and provide the most appropriate response.

[0770] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0771] In this invention, the server includes speech recognition means for converting voice input into a digital signal, data processing means for analyzing the converted digital signal and obtaining relevant information, and speech synthesis means for converting the obtained information into voice output and providing it to the user. This makes it possible to analyze the emotions of a customer in real time from their voice and provide an optimal customer service response that corresponds to those emotions.

[0772] "Speech recognition means" refers to a device or process that converts speech into a digital signal.

[0773] "Data processing means" refers to a function or device for analyzing digital signals and obtaining related information.

[0774] "Speech synthesis means" refers to a technology or device for converting acquired information into speech output and providing it to the user.

[0775] "Interface means" refers to a design or device that simplifies operation through a user interface and assists elderly users in using the device.

[0776] "Customer service support tools" refer to functions or systems that support customer service in stores and generate optimal responses according to the customer's emotions.

[0777] "Environmental data management means" refers to a technology or device for acquiring environmental data from IoT devices via an information and communication network and analyzing it using data processing means.

[0778] "Community building means" refers to functions or systems that provide community features, allow users to share information with each other, and facilitate communication.

[0779] The system realizing this invention primarily utilizes means for speech recognition, data processing, speech synthesis, interface, customer service support, environmental data management, and community building. It acquires the user's voice through smart glasses or other voice input devices and converts it into a digital signal in real time. Cloud-based speech recognition services such as Google Cloud Speech-to-Text are used for speech recognition. This digital signal is sent to a server and analyzed by sentiment analysis software such as IBM Watson Natural Language Understanding. Based on the analyzed sentiment data, an appropriate response or customer service method is generated using an OpenAI GPT-based model. This response is then provided to the user via speech synthesis.

[0780] Furthermore, if a user is wearing smart glasses in the store, the display will show customer service recommendations. This allows staff to provide service while considering the customer's emotions. For example, if a customer appears tired when asking about a product, the emotion engine will recognize that the customer is tired and suggest products that promote relaxation or encourage a gentler tone of voice.

[0781] For example, when a user says, "I'm looking for something to help me relax," the server can use the information that "the customer seems to want to relax" to generate recommendations for products with relaxation effects. An example of a prompt to the generating AI model would be, "The customer is asking about products with relaxation effects. What do you recommend?"

[0782] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0783] Step 1:

[0784] The device acquires the user's voice. The input is the user's voice, which is collected using the microphone of smart glasses or a smartphone. The output is real-time audio data. This audio data is temporarily stored within the device.

[0785] Step 2:

[0786] The device converts the audio data into a digital signal. The input is the audio data collected in step 1, which is then converted into text using a speech recognition engine (e.g., Google Cloud Speech-to-Text). The output is text data. This converted digital signal is immediately sent to the server.

[0787] Step 3:

[0788] The server analyzes the text data to identify the user's emotions. The input is the text data generated in step 2, which is analyzed using emotion analysis software (e.g., IBM Watson Natural Language Understanding). The output is data indicating the user's emotions. Further analysis is then performed based on this data.

[0789] Step 4:

[0790] The server generates optimal responses and product recommendations based on the analysis results. The input is the sentiment data obtained in step 3, and a generative AI model (e.g., OpenAI's GPT-based model) is used to generate the response. The output is the content and conversation suggested to the user. This generated response is used in the next step.

[0791] Step 5:

[0792] The device synthesizes the generated response into speech and provides it to the user. The input is the response data generated in step 4, which is converted into speech output using speech synthesis technology. The output is the content provided to the user aurally. As a result, the user can receive a response that aligns with their own emotions.

[0793] Step 6:

[0794] Visual information is displayed on the user's display (such as the screen of smart glasses). The input is the information generated in step 4, and the content is displayed on the display screen at the appropriate time. The output is the visually presented information. This display allows the user to enjoy a deeper customer service experience.

[0795] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0796] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0797] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0798] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0799] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0800] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0801] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0802] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0803] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0804] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0805] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0806] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0807] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0808] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0809] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0810] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0811] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0812] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0813] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0814] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0815] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0816] The following is further disclosed regarding the embodiments described above.

[0817] (Claim 1)

[0818] A voice recognition means that converts voice input into a digital signal,

[0819] A data processing means for analyzing the converted digital signal and obtaining related information,

[0820] A speech synthesis means that converts acquired information into speech output and provides it to the user,

[0821] An interface means that simplifies operation through a user interface and supports use by elderly users,

[0822] Educational support tools for providing agricultural educational content in augmented reality format,

[0823] A system that includes this.

[0824] (Claim 2)

[0825] The system according to claim 1, further comprising environmental data management means for acquiring environmental data from IoT devices via an information and communication network and analyzing it using data processing means.

[0826] (Claim 3)

[0827] The system according to claim 1, further comprising means for community formation to provide community functions and facilitate communication among users by sharing information.

[0828] "Example 1"

[0829] (Claim 1)

[0830] A voice recognition means that converts voice input into a digital signal,

[0831] A data processing means for analyzing the converted digital signal and obtaining related information,

[0832] A speech synthesis means that converts acquired information into speech output and provides it to the user,

[0833] An interface means that simplifies operation through a user interface and supports use by elderly users,

[0834] Educational support tools for providing agricultural education content in augmented reality format,

[0835] Information extraction means that generates an appropriate response based on received data using a generative model,

[0836] A system that includes this.

[0837] (Claim 2)

[0838] The system according to claim 1, further comprising environmental data management means for acquiring environmental data from sensor devices via an information and communication network and analyzing it using data processing means.

[0839] (Claim 3)

[0840] The system according to claim 1, further comprising information sharing means for providing collaborative functions and promoting communication by sharing information with other users.

[0841] "Application Example 1"

[0842] (Claim 1)

[0843] A speech recognition means that converts an acoustic input into an encoded signal,

[0844] A data processing means for analyzing encoded signals and obtaining related information,

[0845] A speech synthesis means that converts acquired information into sound output and provides it to the user,

[0846] An interface means that simplifies operation through the user interface and supports use by the elderly,

[0847] Educational support tools for providing educational information related to manufacturing in augmented reality format,

[0848] A system that includes this.

[0849] (Claim 2)

[0850] The system according to claim 1, further comprising environmental data management means for acquiring state data from Internet of Things devices via an information transmission network and analyzing it using data processing means.

[0851] (Claim 3)

[0852] The system according to claim 1, further comprising means for community formation to provide cooperative functions and facilitate communication by sharing information with other users, and means for supporting efficient work in industrial facilities.

[0853] "Example 2 of combining an emotion engine"

[0854] (Claim 1)

[0855] Processing means for converting acoustic data into digital information,

[0856] Computational means for analyzing converted digital information and extracting relevant information,

[0857] A sound synthesis system that converts extracted information into sound output and provides it to the user,

[0858] A user interface component that simplifies operation through the user interface and supports use by elderly users,

[0859] An educational support department to provide agricultural educational materials in augmented reality format,

[0860] An emotion analysis system that recognizes emotions from the user's voice input and optimizes the response,

[0861] A system that includes this.

[0862] (Claim 2)

[0863] The system according to claim 1, further comprising an environmental data management department that acquires environmental information from sensor devices via an information and communication network and analyzes it using a computing department.

[0864] (Claim 3)

[0865] The system according to claim 1, further comprising a cooperation formation section that provides cooperative functions and facilitates communication by allowing users to share information with each other.

[0866] "Application example 2 when combining with an emotional engine"

[0867] (Claim 1)

[0868] A voice recognition means that converts voice input into a digital signal,

[0869] A data processing means for analyzing the converted digital signal and obtaining related information,

[0870] A speech synthesis means that converts acquired information into speech output and provides it to the user,

[0871] An interface means that simplifies operation through a user interface and supports use by elderly users,

[0872] A customer service support tool that assists with customer service in stores and generates the optimal response according to the customer's emotions,

[0873] A system that includes this.

[0874] (Claim 2)

[0875] The system according to claim 1, which is an environmental data management means that acquires environmental data from IoT devices via an information and communication network and analyzes it using a data processing means.

[0876] (Claim 3)

[0877] The system according to claim 1, which provides community functions and means for community formation to facilitate communication among users by allowing them to share information. [Explanation of Symbols]

[0878] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A voice recognition means that converts voice input into a digital signal, A data processing means for analyzing the converted digital signal and obtaining related information, A speech synthesis means that converts acquired information into speech output and provides it to the user, An interface means that simplifies operation through a user interface and supports use by elderly users, Educational support tools for providing agricultural educational content in augmented reality format, A system that includes this.

2. The system according to claim 1, further comprising environmental data management means for acquiring environmental data from IoT devices via an information and communication network and analyzing it using data processing means.

3. The system according to claim 1, further comprising means for community formation to provide community functions and facilitate communication among users by sharing information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A