system

A system using voice recognition and GPS/Wi-Fi positioning with secure communication aids in efficiently locating lost mobile devices indoors by emitting sound and displaying location on a map, addressing inefficiencies in existing methods.

JP2026073482APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

There is an inconvenience in locating lost mobile devices indoors, particularly for elderly or young individuals living alone or busy household managers, as they often need to use different communication devices to make calls, and existing methods are inefficient and lack visual confirmation.

Method used

A system utilizing voice recognition to emit sound from a portable device, combined with GPS and Wi-Fi positioning, and displaying the device's location on a map, ensuring secure data exchange through encrypted communication.

Benefits of technology

Enables quick and accurate location of mobile devices indoors by integrating acoustic and visual cues, reducing the risk of loss and theft while ensuring secure communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073482000001_ABST
    Figure 2026073482000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Information processing means that generates a command to ring a portable device based on user instructions received by voice recognition means, Means for controlling the acoustic output of a portable device in response to a command generated by the information processing means, A means of acquiring location information from a portable device and visually displaying it in conjunction with map information inside a house, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] There are often situations where a mobile device is lost indoors, and there is an inconvenience that it is necessary to make a phone call again using a different communication device each time. This problem is particularly serious for elderly or young people living alone or busy household managers, and there is a need to quickly and easily locate the lost mobile device.

Means for Solving the Problems

[0005] This invention includes an information processing means that receives user instructions using voice recognition means and generates commands to audibly notify the user of a portable device. The portable device emits a sound based on the received command, thereby making it easier to pinpoint its location. Furthermore, the location information of the portable device is acquired and displayed on a map of the house, making it possible to visually identify the location of the portable device. In addition, the invention provides highly accurate location information by applying Wi-Fi signals and GPS positioning means, and employs technology that ensures secure data exchange using encrypted communication protocols.

[0006] "Voice recognition means" refers to a technological device that analyzes the user's voice input and recognizes the content of the instructions.

[0007] An "information processing means" is a device that generates appropriate commands based on data acquired from a speech recognition means and controls communication terminals and the entire system.

[0008] A "portable device" refers to a portable communication device that receives voice commands and outputs sound.

[0009] "Means for controlling sound output" refers to a mechanism and control method that enables a portable device to play specified sounds or melodies.

[0010] "Location information" refers to data that indicates the current geographical location of a portable device.

[0011] "In-house map information" refers to digital map data that shows the structure and room layout within the building where the user resides.

[0012] "Wi-Fi signal" is a technology that estimates location based on signals provided through a wireless network.

[0013] A "GPS positioning system" is a technology that acquires positional information from satellites and indicates a specific point on Earth.

[0014] An "encrypted communication protocol" is a communication standard used to encrypt data in order to send and receive data securely and prevent unauthorized access from external sources. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention relates to a system that uses voice recognition to make it easier to find a portable device. The system includes voice recognition means, information processing means, and a portable device. The portable device is equipped with a function to control its acoustic output and can acquire its location information and display it on a map of the house.

[0037] System processing flow:

[0038] 1. The user instructs the voice assistant to "make my cell phone ring." The voice assistant then transmits this instruction to the computer.

[0039] 2. The server has a voice recognition system that analyzes voice commands received from the user. Through this analysis, it understands the content of the command and recognizes the user's intention to make the portable device ring.

[0040] 3. The server then uses a predetermined information processing means to generate an appropriate signal based on the voice command. This signal is sent to the portable device to instruct it to perform the necessary action.

[0041] 4. When the terminal (portable device) receives a signal from the server, it plays a pre-registered melody via sound output. This playback allows the user to pinpoint their location by relying on their hearing.

[0042] 5. The device collects its own location information using Wi-Fi signals and GPS positioning and transmits it to the server. The location information is provided to the user as visual location information in conjunction with a map of the house.

[0043] 6. Users can visually determine the location of their mobile device by looking at the provided map information, without relying on sound.

[0044] Specific example:

[0045] For example, if a user is in their living room at home and has forgotten where they put their mobile device, they can say to their voice assistant, "Make my phone ring." Based on this instruction, the mobile device will ring, and the server will display the device's location on a map of the house. The user can then start walking, relying on the sound, while simultaneously checking the map on their smartphone or tablet to find the exact location. In this way, the user can efficiently search for their mobile device.

[0046] The following describes the processing flow.

[0047] Step 1:

[0048] The user instructs the voice assistant to "make my phone ring." This voice message is received by the registered device and sent to the server.

[0049] Step 2:

[0050] The server analyzes the received voice command using speech recognition technology. It recognizes that the command is an instruction to "make the cell phone ring" and determines the corresponding action.

[0051] Step 3:

[0052] The server uses information processing tools to generate a command to play a melody on the mobile device. This command is encrypted and transmitted to the mobile device via the internet.

[0053] Step 4:

[0054] The terminal (portable device) receives a command from the server and plays a pre-specified melody stored on the device. The sound output provides a clue for the user to physically locate the device.

[0055] Step 5:

[0056] The device outputs sound while simultaneously acquiring its own location information. Wi-Fi signals and GPS location measurement methods support this process.

[0057] Step 6:

[0058] The device uploads the acquired location information to the server. The data is transmitted via a secure protocol and associated with map information of the user's home.

[0059] Step 7:

[0060] The server analyzes location information and updates the map for display in the user-accessible interface. The current location of the mobile device is visually displayed, allowing the user to easily pinpoint their location.

[0061] Step 8:

[0062] Users can use sound and map information to quickly locate their portable devices and retrieve them.

[0063] (Example 1)

[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0065] Mobile devices are frequently misplaced in daily life, and locating and searching for them places a significant burden on users. Furthermore, simple acoustic signals alone are insufficient; there is a need for a method that integrates sound and visual information to quickly and accurately pinpoint the device's location.

[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0067] In this invention, the server includes processing means for receiving voice input, interpreting user instructions using natural language processing, and generating a signal to call a mobile terminal; control means for controlling the sound output device of the mobile terminal in response to the signal generated by the processing means; and display means for acquiring location information of the mobile terminal and integrating the location data with spatial information of the house for visual display. This enables the user to efficiently find the mobile terminal by combining acoustic and visual information.

[0068] "Voice input" refers to the process or system of receiving voice data spoken by a user in digital format.

[0069] "Natural language processing" is a technology that interprets the meaning from a user's speech and converts it into instructions or information that a computer can understand.

[0070] A "mobile device" refers to a portable information and communication device, primarily one that has telephone and internet connectivity functions.

[0071] A "signal" is a series of informational data used to transmit communications or commands, and is transmitted electronically or optically.

[0072] An "acoustic output device" is a device or component that receives a signal and reproduces it as sound.

[0073] "Control means" refers to technologies or devices used to guide specific equipment or processes to perform their intended actions.

[0074] "Location information" refers to data that indicates the geographical or spatial location where an object exists.

[0075] "Spatial information" refers to data and mapping information that shows the physical arrangement and relationships in a specific place or environment.

[0076] "Visual display" refers to a method or technique for presenting information to users visually.

[0077] This invention relates to a mobile device search system utilizing voice recognition. Specific embodiments for implementing this system are described below.

[0078] The server is equipped with a speech recognition engine to receive and analyze voice input. For example, it uses the Google® Cloud Speech-to-Text API to perform natural language analysis on the user's voice and generate a signal to call a mobile device. The generated signal is sent to the mobile device via an encrypted communication protocol.

[0079] The device analyzes the received signal to control the audio output device. This analysis is achieved using MediaPlayer on Android® devices and AVAudioPlayer on iOS devices. This allows the device to play a predetermined melody, enabling the user to locate the device using the sound as a clue. Furthermore, the device measures location information using a combination of Wi-Fi and GPS sensors and sends this information back to the server.

[0080] Location information is integrated with spatial information of the house on the server. This spatial information is displayed using APIs such as Google Maps to visually show the mobile device's location on the user's device. This allows users to efficiently locate their mobile device both audibly and visually.

[0081] For example, if a user misplaces their mobile device at home, they can instruct the voice assistant to "make my phone ring," causing the device to emit a sound. They can then quickly locate the device by referring to map information displayed on the server. In this way, the system can support the user's daily life.

[0082] One possible prompt for the generating AI model would be: "To find a lost mobile device in my home, use the voice assistant to make my phone ring and display its location on a map of my house."

[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0084] Step 1:

[0085] The user issues a command to the voice assistant, such as "Ring my cell phone." The voice assistant converts this voice command into digital data and sends it to a server. The input is the user's voice, and the output is voice data. Speech recognition technology is used for this conversion.

[0086] Step 2:

[0087] The server inputs the received audio data into a speech recognition engine, which performs natural language processing to analyze the user's intent. This process generates a command to ring the mobile device. The input is audio data, and the output is a control signal. Specifically, the speech recognition engine converts the audio into text and generates a command based on that.

[0088] Step 3:

[0089] The server encrypts the generated control signals using data protection technology and transmits them to the mobile device. The input is the control signal, and the output is the encrypted signal. Encryption ensures the security of the communication.

[0090] Step 4:

[0091] The terminal receives and decrypts an encrypted signal from the server. At this point, the terminal receives specific commands to control the audio output device. The input is the encrypted signal, and the output is the control command. In terms of specific operation, the terminal's decryption process decodes the signal.

[0092] Step 5:

[0093] The terminal controls the sound output device based on control commands and plays a pre-set melody. This helps the user find the mobile device by relying on sound. The input is the control command, and the output is sound. Specifically, the sound playback software in the terminal operates according to the commands.

[0094] Step 6:

[0095] The device acquires its location information in real time using Wi-Fi and GPS. This location information is transmitted to the server. The input is the location information collection function, and the output is location data. Specifically, the location measurement module operates and collects data.

[0096] Step 7:

[0097] The server integrates the received location data with the room's map information and displays the terminal's location on the user's visual device. The input consists of location data and map information, while the output is visual display data. Specifically, map display software uses the location information for visualization.

[0098] (Application Example 1)

[0099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0100] Traditionally, if users forget where they placed their valuables or personal devices, it has been difficult for them to quickly and accurately locate them. Furthermore, there is a lack of visual confirmation methods, creating a need for technologies that reduce the risk of loss or theft. In addition, secure communication methods are required.

[0101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0102] In this invention, the server includes information processing means for generating a command to ring a portable device based on user instructions received by voice recognition means, means for acquiring location information of the portable device and visually displaying it in conjunction with map information within a building, and means for visually displaying the location of valuables on an augmented reality display device based on voice instructions. This allows users to quickly locate portable devices and valuables using voice and visually confirm their location, thereby reducing the risk of loss or theft. Furthermore, secure information exchange can be achieved through encrypted communication.

[0103] "Voice recognition means" refers to a device or technology that analyzes the voice spoken by a user and converts that voice into appropriate commands or text.

[0104] "Information processing means" refers to a device or system that generates signals based on commands and data converted by speech recognition means and performs processing to execute necessary actions.

[0105] A "portable device" primarily refers to a portable electronic device capable of producing sound output or acquiring location information.

[0106] "Means for controlling sound output" refers to control technologies that manage the type of sound, volume, playback timing, etc., in an audio device to produce the intended sound.

[0107] "Location information" refers to data that indicates the current location of a specific object, and is often represented as geographical coordinate information.

[0108] "Map information" refers to data that visually represents the location within a building or geographical space, as well as information about its surroundings.

[0109] An "augmented reality display device" is a device or technology that overlays digital information and images onto the real world's field of vision.

[0110] An "encrypted communication method" is a technology that ensures the security of communication by encrypting information using a specific algorithm when transmitting data and decrypting it at the receiving end.

[0111] The system for realizing this invention includes a speech recognition means, an information processing means, and a portable device. The server uses a speech recognition library to receive and analyze voice commands spoken by the user through a microphone. The analyzed voice data is converted into text format and passed to the information processing means. There, an appropriate signal for ringing the portable device is generated based on the voice command.

[0112] The device collects current location information using GPS or wireless signals. The acquired location information is transmitted to a server and, in conjunction with map information, is visually displayed on the user's augmented reality display device, such as smart glasses.

[0113] Users can quickly locate their devices by issuing voice commands. Furthermore, they can efficiently find their valuables by referring to the visually displayed map information.

[0114] To ensure security in this process, encrypted communication methods are used for data transmission and reception. For example, a user can use smart glasses in a cafe and ask the voice assistant to "make my phone ring," allowing them to visually locate their phone. This helps prevent loss and allows for secure location tracking.

[0115] An example of a prompt for a generative AI model is, "I've forgotten where I put my wallet, so I'd like to ask my voice assistant to find out where it is. How do I instruct it?"

[0116] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0117] Step 1:

[0118] The user uses a voice assistant and speaks into a voice input device. The user's voice commands are received as input, converted into digital signals by a speech recognition system, and then sent to a server. Here, a speech recognition library is used to analyze the voice data and generate corresponding command text.

[0119] Step 2:

[0120] The server receives the text generated by the speech recognition library and analyzes its content. Based on the analysis, a command to ring the mobile device is generated. In this information processing process, an appropriate signal is created from the input text data and sent to the mobile device as output.

[0121] Step 3:

[0122] The terminal receives a command from the server. Based on this command, the terminal's acoustic device is activated and emits a pre-set sound. This allows the user to locate the terminal by relying on the sound.

[0123] Step 4:

[0124] The device uses its built-in GPS and wireless signals to collect its own location information. The data obtained from these sensors is processed as input to generate the device's current location information. This location information is sent to a server and integrated with map information as output.

[0125] Step 5:

[0126] The server combines the received location information with map information within the user's building and converts it into visual information. This process involves data processing and calculations to make it visually displayable on an augmented reality display device. Finally, the user can confirm the specific location of their device through the presented visual information.

[0127] Step 6:

[0128] Users can refer to visual feedback from augmented reality displays and use both sound and sight to quickly find their desired valuables. This enables decision-making based not only on sound but also on visual cues.

[0129] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0130] This invention provides a system that combines speech recognition and emotion analysis, offering an efficient means of finding a portable device lost inside a house. The system includes speech recognition means, information processing means, an emotion engine, and a portable device. The portable device is equipped with a function to control acoustic output, and by acquiring location information and displaying it on a map of the house, the user can visually understand the device's location.

[0131] System processing flow:

[0132] 1. The user instructs the voice assistant to "make my cell phone ring." The voice assistant collects this voice message and sends it to the server.

[0133] 2. The server analyzes the user's voice received by the speech recognition system. At this time, the emotion engine analyzes the tone, speed, and pitch of the voice to understand the user's emotional state.

[0134] 3. The server generates instructions to ring the portable device based on the voice analysis results. The type and volume of the sound output are adjusted based on the results of the emotion engine.

[0135] 4. The terminal (portable device) receives instructions from the server and outputs the registered melody at an appropriate volume. For example, if the user is emotionally agitated, the volume will increase to attract their attention.

[0136] 5. In parallel with audio output, the terminal collects its own location information using Wi-Fi signals and GPS positioning means and transmits it to the server.

[0137] 6. The server uses location information to update a map of the house and reflects this in an interface that visually displays the location of the mobile device to the user.

[0138] Specific example:

[0139] For example, if a user is working in their home office and forgets where they put their smartphone, they might anxiously ask their voice assistant to "make my phone ring." The voice is analyzed by an emotion engine, and a command is sent to the server to ring the phone at a volume appropriate to the user's emotion. The voice tone indicates that the user is anxious, so the phone plays a melody at a higher volume than usual. Using this sound signal, the user begins searching for their smartphone and simultaneously checks a map, allowing them to find it quickly. In this way, the user experience can be improved by using thoughtful, emotion-sensitive features.

[0140] The following describes the processing flow.

[0141] Step 1:

[0142] The user instructs the voice assistant to "make my phone ring." This voice command is sent to the server via the device.

[0143] Step 2:

[0144] The server analyzes the received audio data using speech recognition to identify the action the user is requesting. Simultaneously, the emotion engine analyzes the tone, speed, and pitch of the voice to determine the user's emotional state (e.g., calm, anxious, angry).

[0145] Step 3:

[0146] The server determines the type and volume of the sound output according to the user's emotional state and prepares a command to be transmitted to the portable device using information processing means. This command includes information on the melody and volume adjusted according to the emotion.

[0147] Step 4:

[0148] The terminal (portable device) receives commands from the server and starts outputting sound. For example, if it is determined that the user is anxious, it will play a melody at a higher volume than usual.

[0149] Step 5:

[0150] As soon as audio output begins, the device acquires location information using Wi-Fi signals and GPS positioning. This location information is then transmitted from the device to the server.

[0151] Step 6:

[0152] The server analyzes the received location information and combines it with indoor map information to pinpoint the location of the mobile device. The updated map information is reflected in the user-accessible interface, visually indicating the location of the mobile device.

[0153] Step 7:

[0154] Users can quickly locate and retrieve their portable devices using acoustic signals and map information, resulting in a less stressful experience.

[0155] (Example 2)

[0156] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0157] There is a need for efficient methods to quickly locate portable devices when they are lost inside a house or when users are searching for them in a state of confusion. In particular, providing appropriate support in emotionally distressed situations has been difficult with conventional technologies. Considering these circumstances, it is necessary to create technologies that can respond flexibly according to the user's psychological state.

[0158] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0159] In this invention, the server includes information processing means for generating a command to ring the portable device based on user instructions received by voice recognition means, means for analyzing the received voice data and evaluating the user's emotional state using an emotion engine, and means for dynamically adjusting the type and volume of the portable device's sound output based on the command and emotion evaluation generated by the information processing means. This enables accurate visual and auditory support for the user to calmly find the portable device, and facilitates efficient device discovery.

[0160] "Voice recognition means" refers to technology that converts voice commands uttered by a user into digital data and analyzes it.

[0161] "Information processing means" refers to a technology that performs predetermined processing on digital data obtained by voice recognition means to generate commands necessary for operating a portable device.

[0162] An "emotion engine" is a technology that analyzes the characteristics of voice to evaluate the user's emotional state and reflects the results in information processing.

[0163] "Means for dynamically adjusting the type and volume of sound output" refers to a technology that flexibly changes the type and volume of sound output from a portable device based on the evaluation results of the emotion engine.

[0164] "Means for acquiring location information of a portable device" refers to technology that uses communication networks and location measurement technology to determine the precise physical location of a portable device.

[0165] "Visual means of display" refers to a technology that visually displays the location information of a portable device on a map of the house so that users can easily understand it.

[0166] This invention provides a method for efficiently locating a portable device lost inside a house using a system that integrates speech recognition and emotion analysis. Specifically, it employs the following techniques and procedures.

[0167] This system allows users to receive support from a voice assistant if they lose their mobile device. By instructing the voice assistant to "make my phone ring," the voice command is sent to the system. This voice is collected via the device's microphone and transmitted to a server.

[0168] The server converts speech to text using speech recognition software (e.g., a common speech recognition API). Simultaneously, an emotion engine analyzes the characteristics of the speech and evaluates the user's emotional state. Based on this evaluation, it generates commands to adjust the acoustic output of the mobile device. In particular, the volume is automatically adjusted if the user is anxious or confused.

[0169] As a concrete example, consider a scenario where a user has lost their smartphone at home and is searching for it. For instance, if the user anxiously requests, "Make my phone ring," the server, based on its sentiment analysis, assigns a high alert level and instructs the mobile device to emit a loud sound. Relying on the sound signal, the user can quickly locate their smartphone. Furthermore, a map is displayed on the user interface, clearly indicating the location of the mobile device, allowing for simultaneous visual confirmation.

[0170] As an example of a prompt, the AI ​​model might be input with a message like, "Read the user's emotions from their voice and generate instructions to play the device at an appropriate volume." In this way, the present invention enables flexible responses tailored to the user's psychological state, making device discovery more efficient.

[0171] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0172] Step 1:

[0173] The user requests the voice assistant to "make my phone ring." The input is the user's voice, which the device collects using its microphone. The collected voice data is converted into a digital signal and stored on the device.

[0174] Step 2:

[0175] The terminal transmits audio data to the server. The input is digitized audio data, which is then transferred to the server. The output is audio data transmitted over the network. Data is transferred quickly and securely using an internet connection.

[0176] Step 3:

[0177] The server uses speech recognition software to convert speech data into text. The input is speech data received from the terminal, which is processed into text format using a speech recognition algorithm. The output is a text-formatted instruction.

[0178] Step 4:

[0179] The server uses an emotion engine to analyze the tone, speed, and pitch of text data and evaluate the user's emotional state. The input is text data, and the emotion analysis algorithm identifies the user's emotions. The output is the evaluation result of the user's emotional state.

[0180] Step 5:

[0181] The server generates a command to ring the portable device based on the speech recognition results and sentiment evaluation. The input is the analyzed command and sentiment evaluation result, which is used to generate a command that determines the appropriate type and volume of sound output. The output is a specific command message.

[0182] Step 6:

[0183] The terminal (portable device) receives commands from the server and executes an audible output according to those commands. The input is a command message from the server, and the device plays music or an alarm at an appropriate volume using its built-in speaker. The output is a physical acoustic signal.

[0184] Step 7:

[0185] The device collects its own location information using Wi-Fi signals and a location measuring device and transmits it to the server. The input is surrounding signal data, which is processed into location information using a location determination algorithm. The output is location information data sent to the server.

[0186] Step 8:

[0187] The server updates a map of the house based on the received location information, visually displaying the location of the mobile device to the user. The input is location data, which is combined with map data and displayed visually on the user interface. The output is an interface showing the precise location of the mobile device.

[0188] (Application Example 2)

[0189] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0190] In recent years, efficiently locating personal electronic devices that are left behind or lost inside vehicles has become a challenge. In particular, it is difficult for passengers to find lost devices themselves in the complex environment of a moving vehicle. Furthermore, the lack of visual means to determine the location of devices necessitates a rapid resolution of the problem.

[0191] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0192] In this invention, the server includes data processing means that generate a command to sound a personal electronic device based on user instructions and emotion analysis received by voice recognition means; means that control the sound output of the personal electronic device in response to the command generated by the data processing means; and means that acquire location data of the personal electronic device and display it visually in conjunction with map information within the mobile vehicle. This makes it possible to efficiently find a personal electronic device that has been lost within the mobile vehicle.

[0193] "Voice recognition means" refers to a device or program that receives voice instructions from a user and processes them as digital data.

[0194] "Emotional analysis" is a technology that analyzes and evaluates a person's emotional state through voice, facial expressions, and text.

[0195] A "personal electronic device" is a portable electronic device such as a smartphone or tablet, which is a digital terminal used by an individual.

[0196] "Data processing means" refers to a device or program that has the function of receiving information, analyzing it, and generating commands or responses.

[0197] "Means for controlling sound output" refers to a device or program that has the function of setting and adjusting alarm sounds or melodies to be played at an appropriate volume and type.

[0198] "Location data" refers to information that indicates the geographical or spatial location of a specific object.

[0199] "Map information" refers to data that visually represents the geographical elements of a specific space.

[0200] The term "mobile object" refers to a vehicle or its components that can move by carrying people or objects.

[0201] The system that realizes this invention is a technology for quickly locating personal electronic devices such as smartphones and tablets when they are lost inside a vehicle. This system consists of voice recognition means, emotion analysis means, data processing means, sound output control means, location data acquisition means, and map information display means.

[0202] The server first uses a speech recognition API (e.g., Google Speech-to-Text) to acquire and analyze the user's voice commands, and converts the command content into digital data. It also uses an emotion analysis engine (e.g., Affectiva, IBM Watson® Tone Analyzer) to read the user's emotional state from their voice and incorporates this into the data processing. Based on this emotional information, the data processing system generates a command for the optimal acoustic output and transmits it to the personal electronic device.

[0203] Personal electronic devices (smartphones and tablets) receive commands from the server, adjust the volume and sound type, and then sound an alarm. During this process, they collect location data using Bluetooth, Wi-Fi signal strength, and in some cases, the Global Positioning System (GPS), and send this information back to the server. The server then updates the in-vehicle map information based on this location data, providing the user with an interface that intuitively shows the device's location.

[0204] For example, if a user loses sight of their smartphone that they placed on the passenger seat while driving, they can instruct the voice assistant inside the car to "find my smartphone in the car." If the emotion analysis engine detects that the user is in a hurry, the smartphone will emit a loud sound and its location will be displayed on the in-car monitor, allowing the user to find their smartphone immediately.

[0205] An example of a prompt for a generative AI model is: "I want to develop a system to efficiently find a smart device that has been lost inside a car. I want it to use speech recognition and sentiment analysis to visually indicate the device's location."

[0206] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0207] Step 1:

[0208] The user gives instructions to the in-car voice recognition system. The voice command, "Find my smart device," is taken as input, converted into digital data by the voice recognition API, and that data is sent to the server.

[0209] Step 2:

[0210] The server analyzes the received audio data. Using the speech recognition results as input, the emotion analysis engine evaluates the emotional state of the speech (e.g., anxiety, calmness). This process generates two outputs: an audio instruction and an emotional state.

[0211] Step 3:

[0212] Based on the analyzed voice commands and emotional state, the server uses data processing to generate a command for the appropriate acoustic output. This command includes the type and volume of sound for the device. The generated command is sent to the terminal as output data.

[0213] Step 4:

[0214] The terminal (personal electronic device) receives an audio output command from the server. Based on this input, the device sounds an alarm at the specified volume and type. The execution of the audio output results in a physical sound being emitted.

[0215] Step 5:

[0216] Simultaneously, the device collects its own location data. It obtains location information using Wi-Fi signals, Bluetooth, and in some cases, the Global Positioning System (GPS). The acquired location data is then transferred to a server as output.

[0217] Step 6:

[0218] The server updates the in-vehicle map information based on the transmitted location data. Using the location data as input, it generates a user interface that shows the device's location. This interface is displayed on the in-vehicle monitor, providing a visual output.

[0219] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0220] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0221] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0222] [Second Embodiment]

[0223] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0224] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0225] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0226] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0227] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0228] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0229] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0230] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0231] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0232] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0233] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0234] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0235] This invention relates to a system that uses voice recognition to make it easier to find a portable device. The system includes voice recognition means, information processing means, and a portable device. The portable device is equipped with a function to control its acoustic output and can acquire its location information and display it on a map of the house.

[0236] System processing flow:

[0237] 1. The user instructs the voice assistant to "make my cell phone ring." The voice assistant then transmits this instruction to the computer.

[0238] 2. The server has a voice recognition system that analyzes voice commands received from the user. Through this analysis, it understands the content of the command and recognizes the user's intention to make the portable device ring.

[0239] 3. The server then uses a predetermined information processing means to generate an appropriate signal based on the voice command. This signal is sent to the portable device to instruct it to perform the necessary action.

[0240] 4. When the terminal (portable device) receives a signal from the server, it plays a pre-registered melody via sound output. This playback allows the user to pinpoint their location by relying on their hearing.

[0241] 5. The device collects its own location information using Wi-Fi signals and GPS positioning and transmits it to the server. The location information is provided to the user as visual location information in conjunction with a map of the house.

[0242] 6. Users can visually determine the location of their mobile device by looking at the provided map information, without relying on sound.

[0243] Specific example:

[0244] For example, if a user is in their living room at home and has forgotten where they put their mobile device, they can say to their voice assistant, "Make my phone ring." Based on this instruction, the mobile device will ring, and the server will display the device's location on a map of the house. The user can then start walking, relying on the sound, while simultaneously checking the map on their smartphone or tablet to find the exact location. In this way, the user can efficiently search for their mobile device.

[0245] The following describes the processing flow.

[0246] Step 1:

[0247] The user instructs the voice assistant to "make my phone ring." This voice message is received by the registered device and sent to the server.

[0248] Step 2:

[0249] The server analyzes the received voice command using speech recognition technology. It recognizes that the command is an instruction to "make the cell phone ring" and determines the corresponding action.

[0250] Step 3:

[0251] The server uses information processing tools to generate a command to play a melody on the mobile device. This command is encrypted and transmitted to the mobile device via the internet.

[0252] Step 4:

[0253] The terminal (portable device) receives a command from the server and plays a pre-specified melody stored on the device. The sound output provides a clue for the user to physically locate the device.

[0254] Step 5:

[0255] The device outputs sound while simultaneously acquiring its own location information. Wi-Fi signals and GPS location measurement methods support this process.

[0256] Step 6:

[0257] The device uploads the acquired location information to the server. The data is transmitted via a secure protocol and associated with map information of the user's home.

[0258] Step 7:

[0259] The server analyzes location information and updates the map for display in the user-accessible interface. The current location of the mobile device is visually displayed, allowing the user to easily pinpoint their location.

[0260] Step 8:

[0261] Users can use sound and map information to quickly locate their portable devices and retrieve them.

[0262] (Example 1)

[0263] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0264] Mobile devices are frequently misplaced in daily life, and locating and searching for them places a significant burden on users. Furthermore, simple acoustic signals alone are insufficient; there is a need for a method that integrates sound and visual information to quickly and accurately pinpoint the device's location.

[0265] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0266] In this invention, the server includes processing means for receiving voice input, interpreting user instructions using natural language processing, and generating a signal to call a mobile terminal; control means for controlling the sound output device of the mobile terminal in response to the signal generated by the processing means; and display means for acquiring location information of the mobile terminal and integrating the location data with spatial information of the house for visual display. This enables the user to efficiently find the mobile terminal by combining acoustic and visual information.

[0267] "Voice input" refers to the process or system of receiving voice data spoken by a user in digital format.

[0268] "Natural language processing" is a technology that interprets the meaning from a user's speech and converts it into instructions or information that a computer can understand.

[0269] A "mobile device" refers to a portable information and communication device, primarily one that has telephone and internet connectivity functions.

[0270] A "signal" is a series of informational data used to transmit communications or commands, and is transmitted electronically or optically.

[0271] An "acoustic output device" is a device or component that receives a signal and reproduces it as sound.

[0272] "Control means" refers to technologies or devices used to guide specific equipment or processes to perform their intended actions.

[0273] "Location information" refers to data that indicates the geographical or spatial location where an object exists.

[0274] "Spatial information" refers to data and mapping information that shows the physical arrangement and relationships in a specific place or environment.

[0275] "Visual display" refers to a method or technique for presenting information to users visually.

[0276] This invention relates to a mobile device search system utilizing voice recognition. Specific embodiments for implementing this system are described below.

[0277] The server is equipped with a speech recognition engine to receive and analyze voice input. For example, using the Google Cloud Speech-to-Text API, it performs natural language processing on the user's voice and generates a signal to call a mobile device. The generated signal is then sent to the mobile device via an encrypted communication protocol.

[0278] The terminal analyzes the received signal to control the audio output device. This analysis is achieved by using MediaPlayer for Android terminals and AVAudioPlayer for iOS terminals. As a result, the terminal plays a predetermined melody so that the user can identify the position of the terminal based on the sound. Furthermore, the terminal combines Wi-Fi and the GPS sensor to measure the location information and sends this information to the server.

[0279] The location information is integrated with the indoor space information on the server. To display this space information, the Google Maps API or the like is used to visually present the location of the mobile terminal on the user's device. This enables the user to efficiently find the mobile terminal both acoustically and visually.

[0280] As a specific example, when the user forgets to place the mobile terminal inside the house, by instructing the voice assistant to "ring the mobile phone", a sound is emitted from the terminal, and while referring to the map information displayed on the server, the mobile terminal can be quickly found. In this way, the user's life can be supported.

[0281] As an example of a prompt sentence for the generative AI model, an instruction such as "To find a lost mobile device inside the house, use the voice assistant to ring the mobile and display its position on the map of the house." can be given.

[0282] The flow of the specific process in Example 1 will be described using FIG. 11

[0283] Step 1:

[0284] The user issues an instruction such as "ring the mobile phone" to the voice assistant. The voice assistant converts this voice command into digital data and sends it to the server. The input is the user's voice, and the output is voice data. Voice recognition technology is used for this conversion.

[0285] Step 2:

[0286] The server inputs the received voice data into a voice recognition engine, performs natural language processing to analyze the user's intention, and generates a command to ring the mobile terminal through this process. The input is voice data, and the output is a control signal. As a specific operation, the voice recognition engine converts the voice into text and generates a command based on it.

[0287] Step 3:

[0288] The server encrypts the generated control signal using data protection technology and transmits it to the mobile terminal. The input is the control signal, and the output is the encrypted signal. Encryption ensures the security of communication.

[0289] Step 4:

[0290] The terminal receives the encrypted signal from the server and decrypts it. At this point, the terminal receives a specific command for controlling the acoustic output device. The input is the encrypted signal, and the output is the control command. As a specific operation, the decryption process of the terminal decodes the signal.

[0291] Step 5:

[0292] The terminal controls the acoustic output device based on the control command and plays the pre-set melody. This helps the user find the mobile terminal relying on the sound. The input is the control command, and the output is sound. Specifically, in the terminal, the acoustic playback software operates according to the command.

[0293] ]>[[ID=]31]Step 6:

[0294] The terminal utilizes Wi-Fi and GPS in real time to obtain its own location information, which is transmitted to the server. The input is the location information collection function, and the output is location data. As a specific operation, the location measurement module operates to collect data.

[0295] Step

[0296] The server integrates the received location data with the room's map information and displays the terminal's location on the user's visual device. The input consists of location data and map information, while the output is visual display data. Specifically, map display software uses the location information for visualization.

[0297] (Application Example 1)

[0298] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0299] Traditionally, if users forget where they placed their valuables or personal devices, it has been difficult for them to quickly and accurately locate them. Furthermore, there is a lack of visual confirmation methods, creating a need for technologies that reduce the risk of loss or theft. In addition, secure communication methods are required.

[0300] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0301] In this invention, the server includes information processing means for generating a command to ring a portable device based on user instructions received by voice recognition means, means for acquiring location information of the portable device and visually displaying it in conjunction with map information within a building, and means for visually displaying the location of valuables on an augmented reality display device based on voice instructions. This allows users to quickly locate portable devices and valuables using voice and visually confirm their location, thereby reducing the risk of loss or theft. Furthermore, secure information exchange can be achieved through encrypted communication.

[0302] "Voice recognition means" refers to a device or technology that analyzes the voice spoken by a user and converts that voice into appropriate commands or text.

[0303] The "information processing means" is a device or system that generates a signal based on commands and data converted by the voice recognition means and performs processing for executing necessary operations.

[0304] The "portable device" mainly refers to a portable electronic device capable of acoustic output and acquisition of position information.

[0305] The "means for controlling acoustic output" is a control technology for managing the type of sound, volume, playback timing, etc. in an acoustic device and outputting the intended sound.

[0306] "Position information" is data indicating where a specific object is currently located, and is often represented as geographical coordinate information in particular.

[0307] "Map information" is data that visually represents the position and surrounding information in a building or geographical space.

[0308] The "augmented reality display device" is a device or technology for superimposing digital information and images on the real field of view for display.

[0309] The "encrypted communication method" is a technology for ensuring the security of communication by encrypting information with a specific algorithm during data transmission and decrypting it on the receiving side.

[0310] The system for realizing this invention includes voice recognition means, information processing means, and a portable device. The server uses a voice recognition library to receive and analyze voice instructions uttered by the user through the microphone. The analyzed voice data is converted into text format and passed to the information processing means. Here, an appropriate signal for sounding the portable device is generated based on the voice command.

[0311] On the terminal side, the current position information is collected using GPS or wireless signals. The acquired position information is transmitted to the server and visually shown on an augmented reality display device such as the user's smart glasses in cooperation with map information.

[0312] Users can quickly locate their devices by issuing voice commands. Furthermore, they can efficiently find their valuables by referring to the visually displayed map information.

[0313] To ensure security in this process, encrypted communication methods are used for data transmission and reception. For example, a user can use smart glasses in a cafe and ask the voice assistant to "make my phone ring," allowing them to visually locate their phone. This helps prevent loss and allows for secure location tracking.

[0314] An example of a prompt for a generative AI model is, "I've forgotten where I put my wallet, so I'd like to ask my voice assistant to find out where it is. How do I instruct it?"

[0315] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0316] Step 1:

[0317] The user uses a voice assistant and speaks into a voice input device. The user's voice commands are received as input, converted into digital signals by a speech recognition system, and then sent to a server. Here, a speech recognition library is used to analyze the voice data and generate corresponding command text.

[0318] Step 2:

[0319] The server receives the text generated by the speech recognition library and analyzes its content. Based on the analysis, a command to ring the mobile device is generated. In this information processing process, an appropriate signal is created from the input text data and sent to the mobile device as output.

[0320] Step 3:

[0321] The terminal receives a command from the server. Based on this command, the terminal's acoustic device is activated and emits a pre-set sound. This allows the user to locate the terminal by relying on the sound.

[0322] Step 4:

[0323] The device uses its built-in GPS and wireless signals to collect its own location information. The data obtained from these sensors is processed as input to generate the device's current location information. This location information is sent to a server and integrated with map information as output.

[0324] Step 5:

[0325] The server combines the received location information with map information within the user's building and converts it into visual information. This process involves data processing and calculations to make it visually displayable on an augmented reality display device. Finally, the user can confirm the specific location of their device through the presented visual information.

[0326] Step 6:

[0327] Users can refer to visual feedback from augmented reality displays, using both sound and sight to quickly find their valuables. This allows for decision-making based not only on sound but also on visual cues.

[0328] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0329] This invention provides a system that combines speech recognition and emotion analysis, offering an efficient means of finding a portable device lost inside a house. The system includes speech recognition means, information processing means, an emotion engine, and a portable device. The portable device is equipped with a function to control acoustic output, and by acquiring location information and displaying it on a map of the house, the user can visually understand the device's location.

[0330] System processing flow:

[0331] 1. The user instructs the voice assistant to "make my cell phone ring." The voice assistant collects this voice message and sends it to the server.

[0332] 2. The server analyzes the user's voice received by the speech recognition system. At this time, the emotion engine analyzes the tone, speed, and pitch of the voice to understand the user's emotional state.

[0333] 3. The server generates instructions to ring the portable device based on the voice analysis results. The type and volume of the sound output are adjusted based on the results of the emotion engine.

[0334] 4. The terminal (portable device) receives instructions from the server and outputs the registered melody at an appropriate volume. For example, if the user is emotionally agitated, the volume will increase to attract their attention.

[0335] 5. In parallel with audio output, the terminal collects its own location information using Wi-Fi signals and GPS positioning means and transmits it to the server.

[0336] 6. The server uses location information to update a map of the house and reflects this in an interface that visually displays the location of the mobile device to the user.

[0337] Specific example:

[0338] For example, if a user is working in their home office and forgets where they put their smartphone, they might anxiously ask their voice assistant to "make my phone ring." The voice is analyzed by an emotion engine, and a command is sent to the server to make the phone ring at a volume appropriate to the user's emotion. The voice tone indicates that the user is anxious, so the phone plays a melody at a higher volume than usual. Using this sound signal, the user begins searching for their smartphone and simultaneously checks a map, allowing them to find it quickly. In this way, the user experience can be improved by using thoughtful, emotion-sensitive features.

[0339] The following describes the processing flow.

[0340] Step 1:

[0341] The user instructs the voice assistant to "make my phone ring." This voice command is sent to the server via the device.

[0342] Step 2:

[0343] The server analyzes the received audio data using speech recognition to identify the action the user is requesting. Simultaneously, the emotion engine analyzes the tone, speed, and pitch of the voice to determine the user's emotional state (e.g., calm, anxious, angry).

[0344] Step 3:

[0345] The server determines the type and volume of the sound output according to the user's emotional state and prepares a command to be transmitted to the portable device using information processing means. This command includes information on the melody and volume adjusted according to the emotion.

[0346] Step 4:

[0347] The terminal (portable device) receives commands from the server and starts outputting sound. For example, if it is determined that the user is anxious, it will play a melody at a higher volume than usual.

[0348] Step 5:

[0349] As soon as audio output begins, the device acquires location information using Wi-Fi signals and GPS positioning. This location information is then transmitted from the device to the server.

[0350] Step 6:

[0351] The server analyzes the received location information and combines it with indoor map information to pinpoint the location of the mobile device. The updated map information is reflected in the user-accessible interface, visually indicating the location of the mobile device.

[0352] Step 7:

[0353] Users can quickly locate and retrieve their portable devices using acoustic signals and map information, resulting in a less stressful experience.

[0354] (Example 2)

[0355] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0356] There is a need for efficient methods to quickly locate portable devices when they are lost inside a house or when users are searching for them in a state of confusion. In particular, providing appropriate support in emotionally distressed situations has been difficult with conventional technologies. Considering these circumstances, it is necessary to create technologies that can respond flexibly according to the user's psychological state.

[0357] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0358] In this invention, the server includes information processing means for generating a command to ring the portable device based on user instructions received by voice recognition means, means for analyzing the received voice data and evaluating the user's emotional state using an emotion engine, and means for dynamically adjusting the type and volume of the portable device's sound output based on the command and emotion evaluation generated by the information processing means. This enables accurate visual and auditory support for the user to calmly find the portable device, and facilitates efficient device discovery.

[0359] "Voice recognition means" refers to technology that converts voice commands uttered by a user into digital data and analyzes it.

[0360] "Information processing means" refers to a technology that performs predetermined processing on digital data obtained by voice recognition means to generate commands necessary for operating a portable device.

[0361] An "emotion engine" is a technology that analyzes the characteristics of voice to evaluate the user's emotional state and reflects the results in information processing.

[0362] "Means for dynamically adjusting the type and volume of sound output" refers to a technology that flexibly changes the type and volume of sound output from a portable device based on the evaluation results of the emotion engine.

[0363] "Means for acquiring location information of a portable device" refers to technologies that use communication networks and location measurement technologies to determine the precise physical location of a portable device.

[0364] "Visual means of display" refers to a technology that visually displays the location information of a portable device on a map of the house so that users can easily understand it.

[0365] This invention provides a method for efficiently locating a portable device lost inside a house using a system that integrates voice recognition and emotion analysis. Specifically, it employs the following techniques and procedures.

[0366] This system allows users to receive support from a voice assistant if they lose their mobile device. By instructing the voice assistant to "make my phone ring," the voice command is sent to the system. This voice is collected via the device's microphone and transmitted to a server.

[0367] The server converts speech to text using speech recognition software (e.g., a common speech recognition API). Simultaneously, an emotion engine analyzes the characteristics of the speech and evaluates the user's emotional state. Based on this evaluation, it generates commands to adjust the acoustic output of the mobile device. In particular, the volume is automatically adjusted if the user is anxious or confused.

[0368] As a concrete example, consider a scenario where a user has lost their smartphone at home and is searching for it. For instance, if the user anxiously requests, "Make my phone ring," the server, based on its sentiment analysis, assigns a high alert level and instructs the mobile device to emit a loud sound. Relying on the sound signal, the user can quickly locate their smartphone. Furthermore, a map is displayed on the user interface, clearly indicating the location of the mobile device, allowing for simultaneous visual confirmation.

[0369] As an example of a prompt, the AI ​​model might be input with a message like, "Read the user's emotions from their voice and generate instructions to play the device at an appropriate volume." In this way, the present invention enables flexible responses tailored to the user's psychological state, making device discovery more efficient.

[0370] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0371] Step 1:

[0372] The user requests the voice assistant to "make my phone ring." The input is the user's voice, which the device collects using its microphone. The collected voice data is converted into a digital signal and stored on the device.

[0373] Step 2:

[0374] The terminal sends audio data to the server. The input is digitized audio data, which is then transferred to the server. The output is audio data transmitted over the network. Data is transferred quickly and securely using an internet connection.

[0375] Step 3:

[0376] The server uses speech recognition software to convert speech data into text. The input is speech data received from the terminal, which is processed into text format using a speech recognition algorithm. The output is a text-formatted instruction.

[0377] Step 4:

[0378] The server uses an emotion engine to analyze the tone, speed, and pitch of text data and evaluate the user's emotional state. The input is text data, and the emotion analysis algorithm identifies the user's emotions. The output is the evaluation result of the user's emotional state.

[0379] Step 5:

[0380] The server generates a command to ring the portable device based on the speech recognition results and sentiment evaluation. The input is the analyzed command and sentiment evaluation results, which are used to generate a command that determines the appropriate type and volume of sound output. The output is a specific command message.

[0381] Step 6:

[0382] The terminal (portable device) receives commands from the server and executes an audible output according to those commands. The input is a command message from the server, and the device plays music or an alarm at an appropriate volume using its built-in speaker. The output is a physical acoustic signal.

[0383] Step 7:

[0384] The device collects its own location information using Wi-Fi signals and a location measuring device and transmits it to the server. The input is surrounding signal data, which is processed into location information using a location determination algorithm. The output is location information data sent to the server.

[0385] Step 8:

[0386] The server updates a map of the house based on the received location information and visually displays the location of the mobile device to the user. The input is location data, which is combined with map data and displayed visually on the user interface. The output is an interface showing the precise location of the mobile device.

[0387] (Application Example 2)

[0388] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0389] In recent years, efficiently locating personal electronic devices that are left behind or lost inside vehicles has become a challenge. In particular, it is difficult for passengers to find lost devices themselves in the complex environment of a moving vehicle. Furthermore, the lack of visual means to determine the location of devices necessitates a rapid resolution of the problem.

[0390] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0391] In this invention, the server includes data processing means that generate a command to sound a personal electronic device based on user instructions and emotion analysis received by voice recognition means; means that control the sound output of the personal electronic device in response to the command generated by the data processing means; and means that acquire location data of the personal electronic device and display it visually in conjunction with map information within the mobile vehicle. This makes it possible to efficiently find a personal electronic device that has been lost within the mobile vehicle.

[0392] "Voice recognition means" refers to a device or program that receives voice instructions from a user and processes them as digital data.

[0393] "Emotional analysis" is a technology that analyzes and evaluates a person's emotional state through voice, facial expressions, and text.

[0394] A "personal electronic device" is a portable electronic device such as a smartphone or tablet, which is a digital terminal used by an individual.

[0395] "Data processing means" refers to a device or program that has the function of receiving information, analyzing it, and generating commands or responses.

[0396] "Means for controlling sound output" refers to a device or program that has the function of setting and adjusting alarm sounds or melodies to be played at an appropriate volume and type.

[0397] "Location data" refers to information that indicates the geographical or spatial location of a specific object.

[0398] "Map information" refers to data that visually represents the geographical elements of a specific space.

[0399] The term "mobile object" refers to a vehicle or its components that can move by carrying people or objects.

[0400] The system that realizes this invention is a technology for quickly locating personal electronic devices such as smartphones and tablets when they are lost inside a vehicle. This system consists of voice recognition means, emotion analysis means, data processing means, sound output control means, location data acquisition means, and map information display means.

[0401] The server first uses a speech recognition API (e.g., Google Speech-to-Text) to acquire and analyze the user's voice commands, and converts the command content into digital data. It also uses an emotion analysis engine (e.g., Affectiva, IBM Watson Tone Analyzer) to read the user's emotional state from their voice and incorporates this into the data processing. Based on this emotional information, the data processing system generates a command for the optimal acoustic output and transmits it to the personal electronic device.

[0402] Personal electronic devices (smartphones and tablets) receive commands from the server, adjust the volume and sound type, and then sound an alarm. During this process, they collect location data using Bluetooth, Wi-Fi signal strength, and in some cases, the Global Positioning System (GPS), and send this information back to the server. The server then updates the in-vehicle map information based on this location data, providing the user with an interface that intuitively shows the device's location.

[0403] For example, if a user loses sight of their smartphone that they placed on the passenger seat while driving, they can instruct the voice assistant inside the car to "find my smartphone in the car." If the emotion analysis engine detects that the user is in a hurry, the smartphone will emit a loud sound and its location will be displayed on the in-car monitor, allowing the user to find their smartphone immediately.

[0404] An example of a prompt for a generative AI model is: "I want to develop a system to efficiently find a smart device that has been lost inside a car. I want it to use speech recognition and sentiment analysis to visually indicate the device's location."

[0405] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0406] Step 1:

[0407] The user gives instructions to the in-car voice recognition system. The voice command, "Find my smart device," is taken as input, converted into digital data by the voice recognition API, and that data is sent to the server.

[0408] Step 2:

[0409] The server analyzes the received audio data. Using the speech recognition results as input, the emotion analysis engine evaluates the emotional state of the voice (e.g., anxiety, calmness). This process generates two outputs: an audio instruction and an emotional state.

[0410] Step 3:

[0411] Based on the analyzed voice commands and emotional state, the server uses data processing to generate a command for the appropriate acoustic output. This command includes the type and volume of sound for the device. The generated command is then sent to the terminal as output data.

[0412] Step 4:

[0413] The terminal (personal electronic device) receives an audio output command from the server. Based on this input, the device sounds an alarm at the specified volume and type. The execution of the audio output results in a physical sound being emitted.

[0414] Step 5:

[0415] Simultaneously, the device collects its own location data. It obtains location information using Wi-Fi signals, Bluetooth, and in some cases, the Global Positioning System (GPS). The acquired location data is then transferred to a server as output.

[0416] Step 6:

[0417] The server updates the in-vehicle map information based on the transmitted location data. Using the location data as input, it generates a user interface that shows the device's location. This interface is displayed on the in-vehicle monitor, providing a visual output.

[0418] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0419] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0420] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0421] [Third Embodiment]

[0422] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0423] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0424] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0425] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0426] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0427] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0428] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0429] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0430] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0431] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0432] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0433] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0434] This invention relates to a system that uses voice recognition to make it easier to find a portable device. The system includes voice recognition means, information processing means, and a portable device. The portable device is equipped with a function to control its acoustic output and can acquire its location information and display it on a map of the house.

[0435] System processing flow:

[0436] 1. The user instructs the voice assistant to "make my cell phone ring." The voice assistant then transmits this instruction to the computer.

[0437] 2. The server has a voice recognition system that analyzes voice commands received from the user. Through this analysis, it understands the content of the command and recognizes the user's intention to make the portable device ring.

[0438] 3. The server then uses a predetermined information processing means to generate an appropriate signal based on the voice command. This signal is sent to the portable device to instruct it to perform the necessary action.

[0439] 4. When the terminal (portable device) receives a signal from the server, it plays a pre-registered melody via sound output. This playback allows the user to pinpoint their location by relying on their hearing.

[0440] 5. The device collects its own location information using Wi-Fi signals and GPS positioning and transmits it to the server. The location information is provided to the user as visual location information in conjunction with a map of the house.

[0441] 6. Users can visually determine the location of their mobile device by looking at the provided map information, without relying on sound.

[0442] Specific example:

[0443] For example, if a user is in their living room at home and has forgotten where they put their mobile device, they can say to their voice assistant, "Make my phone ring." Based on this instruction, the mobile device will ring, and the server will display the device's location on a map of the house. The user can then start walking, relying on the sound, while simultaneously checking the map on their smartphone or tablet to find the exact location. In this way, the user can efficiently search for their mobile device.

[0444] The following describes the processing flow.

[0445] Step 1:

[0446] The user instructs the voice assistant to "make my phone ring." This voice message is received by the registered device and sent to the server.

[0447] Step 2:

[0448] The server analyzes the received voice command using speech recognition technology. It recognizes that the command is an instruction to "make the cell phone ring" and determines the corresponding action.

[0449] Step 3:

[0450] The server uses information processing tools to generate a command to play a melody on the mobile device. This command is encrypted and transmitted to the mobile device via the internet.

[0451] Step 4:

[0452] The terminal (portable device) receives a command from the server and plays a pre-specified melody stored on the device. The sound output provides a clue for the user to physically locate the device.

[0453] Step 5:

[0454] The device outputs sound while simultaneously acquiring its own location information. Wi-Fi signals and GPS location measurement methods support this process.

[0455] Step 6:

[0456] The device uploads the acquired location information to the server. The data is transmitted via a secure protocol and associated with map information of the user's home.

[0457] Step 7:

[0458] The server analyzes location information and updates the map for display in the user-accessible interface. The current location of the mobile device is visually displayed, allowing the user to easily pinpoint their location.

[0459] Step 8:

[0460] Users can use sound and map information to quickly locate their portable devices and retrieve them.

[0461] (Example 1)

[0462] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0463] Mobile devices are frequently misplaced in daily life, and locating and searching for them places a significant burden on users. Furthermore, simple acoustic signals alone are insufficient; there is a need for a method that integrates sound and visual information to quickly and accurately pinpoint the device's location.

[0464] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0465] In this invention, the server includes processing means for receiving voice input, interpreting user instructions using natural language processing, and generating a signal to call a mobile terminal; control means for controlling the sound output device of the mobile terminal in response to the signal generated by the processing means; and display means for acquiring location information of the mobile terminal and integrating the location data with spatial information of the house for visual display. This enables the user to efficiently find the mobile terminal by combining acoustic and visual information.

[0466] "Voice input" refers to the process or system of receiving voice data spoken by a user in digital format.

[0467] "Natural language processing" is a technology that interprets the meaning from a user's speech and converts it into instructions or information that a computer can understand.

[0468] A "mobile device" refers to a portable information and communication device, primarily one that has telephone and internet connectivity functions.

[0469] A "signal" is a series of informational data used to transmit communications or commands, and is transmitted electronically or optically.

[0470] An "acoustic output device" is a device or component that receives a signal and reproduces it as sound.

[0471] "Control means" refers to technologies or devices used to guide specific equipment or processes to perform their intended actions.

[0472] "Location information" refers to data that indicates the geographical or spatial location where an object exists.

[0473] "Spatial information" refers to data and mapping information that shows the physical arrangement and relationships in a specific place or environment.

[0474] "Visual display" refers to a method or technique for presenting information to users visually.

[0475] This invention relates to a mobile device search system utilizing voice recognition. Specific embodiments for implementing this system are described below.

[0476] The server is equipped with a speech recognition engine to receive and analyze voice input. For example, using the Google Cloud Speech-to-Text API, it performs natural language processing on the user's voice and generates a signal to call a mobile device. The generated signal is then sent to the mobile device via an encrypted communication protocol.

[0477] The device analyzes the received signal to control the audio output device. This analysis is performed using MediaPlayer on Android devices and AVAudioPlayer on iOS devices. This allows the device to play a predetermined melody, enabling the user to locate the device using the sound as a clue. Furthermore, the device measures location information using a combination of Wi-Fi and GPS sensors and sends this information back to the server.

[0478] Location information is integrated with spatial information of the house on the server. This spatial information is displayed using APIs such as Google Maps to visually show the mobile device's location on the user's device. This allows users to efficiently locate their mobile device both audibly and visually.

[0479] For example, if a user misplaces their mobile device at home, they can instruct the voice assistant to "make my phone ring," causing the device to emit a sound. They can then quickly locate the device by referring to map information displayed on the server. In this way, the system can support the user's daily life.

[0480] One possible prompt for the generating AI model would be: "To find a lost mobile device in my home, use the voice assistant to make my phone ring and display its location on a map of my house."

[0481] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0482] Step 1:

[0483] The user issues a command to the voice assistant, such as "Ring my cell phone." The voice assistant converts this voice command into digital data and sends it to a server. The input is the user's voice, and the output is voice data. Speech recognition technology is used for this conversion.

[0484] Step 2:

[0485] The server inputs the received audio data into a speech recognition engine, which performs natural language processing to analyze the user's intent. This process generates a command to ring the mobile device. The input is audio data, and the output is a control signal. Specifically, the speech recognition engine converts the audio into text and generates a command based on that.

[0486] Step 3:

[0487] The server encrypts the generated control signals using data protection technology and transmits them to the mobile device. The input is the control signal, and the output is the encrypted signal. Encryption ensures the security of the communication.

[0488] Step 4:

[0489] The terminal receives and decrypts an encrypted signal from the server. At this point, the terminal receives specific commands to control the audio output device. The input is the encrypted signal, and the output is the control command. In terms of specific operation, the terminal's decryption process decodes the signal.

[0490] Step 5:

[0491] The terminal controls the sound output device based on control commands and plays a pre-set melody. This helps the user find the mobile device by relying on sound. The input is the control command, and the output is sound. Specifically, the sound playback software in the terminal operates according to the commands.

[0492] Step 6:

[0493] The device acquires its location information in real time using Wi-Fi and GPS. This location information is transmitted to the server. The input is the location information collection function, and the output is location data. Specifically, the location measurement module operates and collects data.

[0494] Step 7:

[0495] The server integrates the received location data with the room's map information and displays the terminal's location on the user's visual device. The input consists of location data and map information, while the output is visual display data. Specifically, map display software uses the location information for visualization.

[0496] (Application Example 1)

[0497] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0498] Traditionally, if users forget where they placed their valuables or personal devices, it has been difficult for them to quickly and accurately locate them. Furthermore, there is a lack of visual confirmation methods, creating a need for technologies that reduce the risk of loss or theft. In addition, secure communication methods are required.

[0499] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0500] In this invention, the server includes information processing means for generating a command to ring a portable device based on user instructions received by voice recognition means, means for acquiring location information of the portable device and visually displaying it in conjunction with map information within a building, and means for visually displaying the location of valuables on an augmented reality display device based on voice instructions. This allows users to quickly locate portable devices and valuables using voice and visually confirm their location, thereby reducing the risk of loss or theft. Furthermore, secure information exchange can be achieved through encrypted communication.

[0501] "Voice recognition means" refers to a device or technology that analyzes the voice spoken by a user and converts that voice into appropriate commands or text.

[0502] "Information processing means" refers to a device or system that generates signals based on commands and data converted by speech recognition means and performs processing to execute necessary actions.

[0503] A "portable device" primarily refers to a portable electronic device capable of producing sound output or acquiring location information.

[0504] "Means for controlling sound output" refers to control technologies that manage the type of sound, volume, playback timing, etc., in an audio device to produce the intended sound.

[0505] "Location information" refers to data that indicates the current location of a specific object, and is often represented as geographical coordinate information.

[0506] "Map information" refers to data that visually represents the location within a building or geographical space, as well as information about its surroundings.

[0507] An "augmented reality display device" is a device or technology that overlays digital information and images onto the real world's field of vision.

[0508] An "encrypted communication method" is a technology that ensures the security of communication by encrypting information using a specific algorithm when transmitting data and decrypting it at the receiving end.

[0509] The system for realizing this invention includes a speech recognition means, an information processing means, and a portable device. The server uses a speech recognition library to receive and analyze voice commands spoken by the user through a microphone. The analyzed voice data is converted into text format and passed to the information processing means. There, an appropriate signal for ringing the portable device is generated based on the voice command.

[0510] The device collects current location information using GPS or wireless signals. The acquired location information is transmitted to a server and, in conjunction with map information, is visually displayed on the user's augmented reality display device, such as smart glasses.

[0511] Users can quickly locate their devices by issuing voice commands. Furthermore, they can efficiently find their valuables by referring to the visually displayed map information.

[0512] To ensure security in this process, encrypted communication methods are used for data transmission and reception. For example, a user can use smart glasses in a cafe and ask the voice assistant to "make my phone ring," allowing them to visually locate their phone. This helps prevent loss and allows for secure location tracking.

[0513] An example of a prompt for a generative AI model is, "I've forgotten where I put my wallet, so I'd like to ask my voice assistant to find out where it is. How do I instruct it?"

[0514] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0515] Step 1:

[0516] The user uses a voice assistant and speaks into a voice input device. The user's voice commands are received as input, converted into digital signals by a speech recognition system, and then sent to a server. Here, a speech recognition library is used to analyze the voice data and generate corresponding command text.

[0517] Step 2:

[0518] The server receives the text generated by the speech recognition library and analyzes its content. Based on the analysis, a command to ring the mobile device is generated. In this information processing process, an appropriate signal is created from the input text data and sent to the mobile device as output.

[0519] Step 3:

[0520] The terminal receives a command from the server. Based on this command, the terminal's acoustic device is activated and emits a pre-set sound. This allows the user to locate the terminal by relying on the sound.

[0521] Step 4:

[0522] The device uses its built-in GPS and wireless signals to collect its own location information. The data obtained from these sensors is processed as input to generate the device's current location information. This location information is sent to a server and integrated with map information as output.

[0523] Step 5:

[0524] The server combines the received location information with map information within the user's building and converts it into visual information. This process involves data processing and calculations to make it visually displayable on an augmented reality display device. Finally, the user can confirm the specific location of their device through the presented visual information.

[0525] Step 6:

[0526] Users can refer to visual feedback from augmented reality displays, using both sound and sight to quickly find their valuables. This allows for decision-making based not only on sound but also on visual cues.

[0527] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0528] This invention provides a system that combines speech recognition and emotion analysis, offering an efficient means of finding a portable device lost inside a house. The system includes speech recognition means, information processing means, an emotion engine, and a portable device. The portable device is equipped with a function to control acoustic output, and by acquiring location information and displaying it on a map of the house, the user can visually understand the device's location.

[0529] System processing flow:

[0530] 1. The user instructs the voice assistant to "make my cell phone ring." The voice assistant collects this voice message and sends it to the server.

[0531] 2. The server analyzes the user's voice received by the speech recognition system. At this time, the emotion engine analyzes the tone, speed, and pitch of the voice to understand the user's emotional state.

[0532] 3. The server generates instructions to ring the portable device based on the voice analysis results. The type and volume of the sound output are adjusted based on the results of the emotion engine.

[0533] 4. The terminal (portable device) receives instructions from the server and outputs the registered melody at an appropriate volume. For example, if the user is emotionally agitated, the volume will increase to attract their attention.

[0534] 5. In parallel with audio output, the terminal collects its own location information using Wi-Fi signals and GPS positioning means and transmits it to the server.

[0535] 6. The server uses location information to update a map of the house and reflects this in an interface that visually displays the location of the mobile device to the user.

[0536] Specific example:

[0537] For example, if a user is working in their home office and forgets where they put their smartphone, they might anxiously ask their voice assistant to "make my phone ring." The voice is analyzed by an emotion engine, and a command is sent to the server to make the phone ring at a volume appropriate to the user's emotion. The voice tone indicates that the user is anxious, so the phone plays a melody at a higher volume than usual. Using this sound signal, the user begins searching for their smartphone and simultaneously checks a map, allowing them to find it quickly. In this way, the user experience can be improved by using thoughtful, emotion-sensitive features.

[0538] The following describes the processing flow.

[0539] Step 1:

[0540] The user instructs the voice assistant to "make my phone ring." This voice command is sent to the server via the device.

[0541] Step 2:

[0542] The server analyzes the received audio data using speech recognition to identify the action the user is requesting. Simultaneously, the emotion engine analyzes the tone, speed, and pitch of the voice to determine the user's emotional state (e.g., calm, anxious, angry).

[0543] Step 3:

[0544] The server determines the type and volume of the sound output according to the user's emotional state and prepares a command to be transmitted to the portable device using information processing means. This command includes information on the melody and volume adjusted according to the emotion.

[0545] Step 4:

[0546] The terminal (portable device) receives commands from the server and starts outputting sound. For example, if it is determined that the user is anxious, it will play a melody at a higher volume than usual.

[0547] Step 5:

[0548] As soon as audio output begins, the device acquires location information using Wi-Fi signals and GPS positioning. This location information is then transmitted from the device to the server.

[0549] Step 6:

[0550] The server analyzes the received location information and combines it with indoor map information to pinpoint the location of the mobile device. The updated map information is reflected in the user-accessible interface, visually indicating the location of the mobile device.

[0551] Step 7:

[0552] Users can quickly locate and retrieve their portable devices using acoustic signals and map information, resulting in a less stressful experience.

[0553] (Example 2)

[0554] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0555] There is a need for efficient methods to quickly locate portable devices when they are lost inside a house or when users are searching for them in a state of confusion. In particular, providing appropriate support in emotionally distressed situations has been difficult with conventional technologies. Considering these circumstances, it is necessary to create technologies that can respond flexibly according to the user's psychological state.

[0556] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0557] In this invention, the server includes information processing means for generating a command to ring the portable device based on user instructions received by voice recognition means, means for analyzing the received voice data and evaluating the user's emotional state using an emotion engine, and means for dynamically adjusting the type and volume of the portable device's sound output based on the command and emotion evaluation generated by the information processing means. This enables accurate visual and auditory support for the user to calmly find the portable device, and facilitates efficient device discovery.

[0558] "Voice recognition means" refers to technology that converts voice commands uttered by a user into digital data and analyzes it.

[0559] "Information processing means" refers to a technology that performs predetermined processing on digital data obtained by voice recognition means to generate commands necessary for operating a portable device.

[0560] An "emotion engine" is a technology that analyzes the characteristics of voice to evaluate the user's emotional state and reflects the results in information processing.

[0561] "Means for dynamically adjusting the type and volume of sound output" refers to a technology that flexibly changes the type and volume of sound output from a portable device based on the evaluation results of the emotion engine.

[0562] "Means for acquiring location information of a portable device" refers to technology that uses communication networks and location measurement technology to determine the precise physical location of a portable device.

[0563] "Visual means of display" refers to a technology that visually displays the location information of a portable device on a map of the house so that users can easily understand it.

[0564] This invention provides a method for efficiently locating a portable device lost inside a house using a system that integrates speech recognition and emotion analysis. Specifically, it employs the following techniques and procedures.

[0565] This system allows users to receive support from a voice assistant if they lose their mobile device. By instructing the voice assistant to "make my phone ring," the voice command is sent to the system. This voice is collected via the device's microphone and transmitted to a server.

[0566] The server converts speech to text using speech recognition software (e.g., a common speech recognition API). Simultaneously, an emotion engine analyzes the characteristics of the speech and evaluates the user's emotional state. Based on this evaluation, it generates commands to adjust the acoustic output of the mobile device. In particular, the volume is automatically adjusted if the user is anxious or confused.

[0567] As a concrete example, consider a scenario where a user has lost their smartphone at home and is searching for it. For instance, if the user anxiously requests, "Make my phone ring," the server, based on its sentiment analysis, assigns a high alert level and instructs the mobile device to emit a loud sound. Relying on the sound signal, the user can quickly locate their smartphone. Furthermore, a map is displayed on the user interface, clearly indicating the location of the mobile device, allowing for simultaneous visual confirmation.

[0568] As an example of a prompt, the AI ​​model might be input with a message like, "Read the user's emotions from their voice and generate instructions to play the device at an appropriate volume." In this way, the present invention enables flexible responses tailored to the user's psychological state, making device discovery more efficient.

[0569] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0570] Step 1:

[0571] The user requests the voice assistant to "make my phone ring." The input is the user's voice, which the device collects using its microphone. The collected voice data is converted into a digital signal and stored on the device.

[0572] Step 2:

[0573] The terminal transmits audio data to the server. The input is digitized audio data, which is then transferred to the server. The output is audio data transmitted over the network. Data is transferred quickly and securely using an internet connection.

[0574] Step 3:

[0575] The server uses speech recognition software to convert speech data into text. The input is speech data received from the terminal, which is processed into text format using a speech recognition algorithm. The output is a text-formatted instruction.

[0576] Step 4:

[0577] The server uses an emotion engine to analyze the tone, speed, and pitch of text data and evaluate the user's emotional state. The input is text data, and the emotion analysis algorithm identifies the user's emotions. The output is the evaluation result of the user's emotional state.

[0578] Step 5:

[0579] The server generates a command to ring the portable device based on the speech recognition results and sentiment evaluation. The input is the analyzed command and sentiment evaluation result, which is used to generate a command that determines the appropriate type and volume of sound output. The output is a specific command message.

[0580] Step 6:

[0581] The terminal (portable device) receives commands from the server and executes an audible output according to those commands. The input is a command message from the server, and the device plays music or an alarm at an appropriate volume using its built-in speaker. The output is a physical acoustic signal.

[0582] Step 7:

[0583] The device collects its own location information using Wi-Fi signals and a location measuring device and transmits it to the server. The input is surrounding signal data, which is processed into location information using a location determination algorithm. The output is location information data sent to the server.

[0584] Step 8:

[0585] The server updates a map of the house based on the received location information, visually displaying the location of the mobile device to the user. The input is location data, which is combined with map data and displayed visually on the user interface. The output is an interface showing the precise location of the mobile device.

[0586] (Application Example 2)

[0587] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0588] In recent years, efficiently locating personal electronic devices that are left behind or lost inside vehicles has become a challenge. In particular, it is difficult for passengers to find lost devices themselves in the complex environment of a moving vehicle. Furthermore, the lack of visual means to determine the location of devices necessitates a rapid resolution of the problem.

[0589] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0590] In this invention, the server includes data processing means that generate a command to sound a personal electronic device based on user instructions and emotion analysis received by voice recognition means; means that control the sound output of the personal electronic device in response to the command generated by the data processing means; and means that acquire location data of the personal electronic device and display it visually in conjunction with map information within the mobile vehicle. This makes it possible to efficiently find a personal electronic device that has been lost within the mobile vehicle.

[0591] "Voice recognition means" refers to a device or program that receives voice instructions from a user and processes them as digital data.

[0592] "Emotional analysis" is a technology that analyzes and evaluates a person's emotional state through voice, facial expressions, and text.

[0593] A "personal electronic device" is a portable electronic device such as a smartphone or tablet, which is a digital terminal used by an individual.

[0594] "Data processing means" refers to a device or program that has the function of receiving information, analyzing it, and generating commands or responses.

[0595] "Means for controlling sound output" refers to a device or program that has the function of setting and adjusting alarm sounds or melodies to be played at an appropriate volume and type.

[0596] "Location data" refers to information that indicates the geographical or spatial location of a specific object.

[0597] "Map information" refers to data that visually represents the geographical elements of a specific space.

[0598] The term "mobile object" refers to a vehicle or its components that can move by carrying people or objects.

[0599] The system that realizes this invention is a technology for quickly locating personal electronic devices such as smartphones and tablets when they are lost inside a vehicle. This system consists of voice recognition means, emotion analysis means, data processing means, sound output control means, location data acquisition means, and map information display means.

[0600] The server first uses a speech recognition API (e.g., Google Speech-to-Text) to acquire and analyze the user's voice commands, and converts the command content into digital data. It also uses an emotion analysis engine (e.g., Affectiva, IBM Watson Tone Analyzer) to read the user's emotional state from their voice and incorporates this into the data processing. Based on this emotional information, the data processing system generates a command for the optimal acoustic output and transmits it to the personal electronic device.

[0601] Personal electronic devices (smartphones and tablets) receive commands from the server, adjust the volume and sound type, and then sound an alarm. During this process, they collect location data using Bluetooth, Wi-Fi signal strength, and in some cases, the Global Positioning System (GPS), and send this information back to the server. The server then updates the in-vehicle map information based on this location data, providing the user with an interface that intuitively shows the device's location.

[0602] For example, if a user loses sight of their smartphone that they placed on the passenger seat while driving, they can instruct the voice assistant inside the car to "find my smartphone in the car." If the emotion analysis engine detects that the user is in a hurry, the smartphone will emit a loud sound and its location will be displayed on the in-car monitor, allowing the user to find their smartphone immediately.

[0603] An example of a prompt for a generative AI model is: "I want to develop a system to efficiently find a smart device that has been lost inside a car. I want it to use speech recognition and sentiment analysis to visually indicate the device's location."

[0604] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0605] Step 1:

[0606] The user gives instructions to the in-car voice recognition system. The voice command, "Find my smart device," is taken as input, converted into digital data by the voice recognition API, and that data is sent to the server.

[0607] Step 2:

[0608] The server analyzes the received audio data. Using the speech recognition results as input, the emotion analysis engine evaluates the emotional state of the speech (e.g., anxiety, calmness). This process generates two outputs: an audio instruction and an emotional state.

[0609] Step 3:

[0610] Based on the analyzed voice commands and emotional state, the server uses data processing to generate a command for the appropriate acoustic output. This command includes the type and volume of sound for the device. The generated command is sent to the terminal as output data.

[0611] Step 4:

[0612] The terminal (personal electronic device) receives an audio output command from the server. Based on this input, the device sounds an alarm at the specified volume and type. The execution of the audio output results in a physical sound being emitted.

[0613] Step 5:

[0614] Simultaneously, the device collects its own location data. It obtains location information using Wi-Fi signals, Bluetooth, and in some cases, the Global Positioning System (GPS). The acquired location data is then transferred to a server as output.

[0615] Step 6:

[0616] The server updates the in-vehicle map information based on the transmitted location data. Using the location data as input, it generates a user interface that shows the device's location. This interface is displayed on the in-vehicle monitor, providing a visual output.

[0617] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0618] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0619] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0620] [Fourth Embodiment]

[0621] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0622] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0623] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0624] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0625] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0626] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0627] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0628] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0629] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0630] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0631] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0632] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0633] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0634] This invention is a system that uses voice recognition to make it easier to find a portable device. The system includes voice recognition means, information processing means, and a portable device. The portable device is equipped with a function to control its sound output and can acquire its location information and display it on a map of the house.

[0635] System processing flow:

[0636] 1. The user instructs the voice assistant to "make my cell phone ring." The voice assistant then transmits this instruction to the computer.

[0637] 2. The server has a voice recognition system that analyzes voice commands received from the user. Through this analysis, it understands the content of the command and recognizes the user's intention to make the portable device ring.

[0638] 3. The server then uses a predetermined information processing means to generate an appropriate signal based on the voice command. This signal is sent to the portable device to instruct it to perform the necessary action.

[0639] 4. When the terminal (portable device) receives a signal from the server, it plays a pre-registered melody via sound output. This playback allows the user to pinpoint their location by relying on their hearing.

[0640] 5. The device collects its own location information using Wi-Fi signals and GPS positioning and transmits it to the server. The location information is provided to the user as visual location information in conjunction with a map of the house.

[0641] 6. Users can visually determine the location of their mobile device by looking at the provided map information, without relying on sound.

[0642] Specific example:

[0643] For example, if a user is in their living room at home and has forgotten where they put their mobile device, they can say to their voice assistant, "Make my phone ring." Based on this instruction, the mobile device will ring, and the server will display the device's location on a map of the house. The user can then start walking, relying on the sound, while simultaneously checking the map on their smartphone or tablet to find the exact location. In this way, the user can efficiently search for their mobile device.

[0644] The following describes the processing flow.

[0645] Step 1:

[0646] The user instructs the voice assistant to "make my cell phone ring." This voice message is received by the registered device and sent to the server.

[0647] Step 2:

[0648] The server analyzes the received voice command using speech recognition technology. It recognizes that the command is an instruction to "make the cell phone ring" and determines the corresponding action.

[0649] Step 3:

[0650] The server uses information processing tools to generate a command to play a melody on the mobile device. This command is encrypted and transmitted to the mobile device via the internet.

[0651] Step 4:

[0652] The terminal (portable device) receives a command from the server and plays a pre-specified melody stored on the device. The sound output provides a clue for the user to physically locate the device.

[0653] Step 5:

[0654] The device outputs sound while simultaneously acquiring its own location information. Wi-Fi signals and GPS location measurement methods support this process.

[0655] Step 6:

[0656] The device uploads the acquired location information to the server. The data is transmitted via a secure protocol and associated with map information of the user's home.

[0657] Step 7:

[0658] The server analyzes location information and updates the map for display in the user-accessible interface. The current location of the mobile device is visually displayed, allowing the user to easily pinpoint their location.

[0659] Step 8:

[0660] Users can use sound and map information to quickly locate their portable devices and retrieve them.

[0661] (Example 1)

[0662] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0663] Mobile devices are frequently misplaced in daily life, and locating and searching for them places a significant burden on users. Furthermore, simple acoustic signals alone are insufficient; there is a need for a method that integrates sound and visual information to quickly and accurately pinpoint the device's location.

[0664] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0665] In this invention, the server includes processing means for receiving voice input, interpreting user instructions using natural language processing, and generating a signal to call a mobile terminal; control means for controlling the sound output device of the mobile terminal in response to the signal generated by the processing means; and display means for acquiring location information of the mobile terminal and integrating the location data with spatial information of the house for visual display. This enables the user to efficiently find the mobile terminal by combining acoustic and visual information.

[0666] "Voice input" refers to the process or system of receiving voice data spoken by a user in digital format.

[0667] "Natural language processing" is a technology that interprets the meaning from a user's speech and converts it into instructions or information that a computer can understand.

[0668] A "mobile device" refers to a portable information and communication device, primarily one that has telephone and internet connectivity functions.

[0669] A "signal" is a series of informational data used to transmit communications or commands, and is transmitted electronically or optically.

[0670] An "acoustic output device" is a device or component that receives a signal and reproduces it as sound.

[0671] "Control means" refers to technologies or devices used to guide specific equipment or processes to perform their intended actions.

[0672] "Location information" refers to data that indicates the geographical or spatial location where an object exists.

[0673] "Spatial information" refers to data and mapping information that shows the physical arrangement and relationships in a specific place or environment.

[0674] "Visual display" refers to a method or technique for presenting information to users visually.

[0675] This invention relates to a mobile device search system utilizing voice recognition. Specific embodiments for implementing this system are described below.

[0676] The server is equipped with a speech recognition engine to receive and analyze voice input. For example, using the Google Cloud Speech-to-Text API, it performs natural language processing on the user's voice and generates a signal to call a mobile device. The generated signal is then sent to the mobile device via an encrypted communication protocol.

[0677] The device analyzes the received signal to control the audio output device. This analysis is performed using MediaPlayer on Android devices and AVAudioPlayer on iOS devices. This allows the device to play a predetermined melody, enabling the user to locate the device using the sound as a clue. Furthermore, the device measures location information using a combination of Wi-Fi and GPS sensors and sends this information back to the server.

[0678] Location information is integrated with spatial information of the house on the server. This spatial information is displayed using APIs such as Google Maps to visually show the mobile device's location on the user's device. This allows users to efficiently locate their mobile device both audibly and visually.

[0679] For example, if a user misplaces their mobile device at home, they can instruct the voice assistant to "make my phone ring," causing the device to emit a sound. They can then quickly locate the device by referring to map information displayed on the server. In this way, the system can support the user's daily life.

[0680] One possible prompt for the generating AI model would be: "To find a lost mobile device in my home, use the voice assistant to make my phone ring and display its location on a map of my house."

[0681] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0682] Step 1:

[0683] The user issues a command to the voice assistant, such as "Ring my cell phone." The voice assistant converts this voice command into digital data and sends it to a server. The input is the user's voice, and the output is voice data. Speech recognition technology is used for this conversion.

[0684] Step 2:

[0685] The server inputs the received audio data into a speech recognition engine, which performs natural language processing to analyze the user's intent. This process generates a command to ring the mobile device. The input is audio data, and the output is a control signal. Specifically, the speech recognition engine converts the audio into text and generates a command based on that.

[0686] Step 3:

[0687] The server encrypts the generated control signals using data protection technology and transmits them to the mobile device. The input is the control signal, and the output is the encrypted signal. Encryption ensures the security of the communication.

[0688] Step 4:

[0689] The terminal receives and decrypts an encrypted signal from the server. At this point, the terminal receives specific commands to control the audio output device. The input is the encrypted signal, and the output is the control command. In terms of operation, the terminal's decryption process decodes the signal.

[0690] Step 5:

[0691] The terminal controls the sound output device based on control commands and plays a pre-set melody. This helps the user find the mobile device by relying on sound. The input is the control command, and the output is sound. Specifically, the sound playback software in the terminal operates according to the commands.

[0692] Step 6:

[0693] The device acquires its location information in real time using Wi-Fi and GPS. This location information is transmitted to the server. The input is the location information collection function, and the output is location data. Specifically, the location measurement module operates and collects data.

[0694] Step 7:

[0695] The server integrates the received location data with the room's map information and displays the terminal's location on the user's visual device. The input consists of location data and map information, while the output is visual display data. Specifically, map display software uses the location information for visualization.

[0696] (Application Example 1)

[0697] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0698] Traditionally, if users forget where they placed their valuables or personal devices, it has been difficult for them to quickly and accurately locate them. Furthermore, there is a lack of visual confirmation methods, creating a need for technologies that reduce the risk of loss or theft. In addition, secure communication methods are required.

[0699] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0700] In this invention, the server includes information processing means for generating a command to ring a portable device based on user instructions received by voice recognition means, means for acquiring location information of the portable device and visually displaying it in conjunction with map information within a building, and means for visually displaying the location of valuables on an augmented reality display device based on voice instructions. This allows users to quickly locate portable devices and valuables using voice and visually confirm their location, thereby reducing the risk of loss or theft. Furthermore, secure information exchange can be achieved through encrypted communication.

[0701] "Voice recognition means" refers to a device or technology that analyzes the voice spoken by a user and converts that voice into appropriate commands or text.

[0702] "Information processing means" refers to a device or system that generates signals based on commands and data converted by speech recognition means and performs processing to execute necessary actions.

[0703] A "portable device" primarily refers to a portable electronic device capable of producing sound output or acquiring location information.

[0704] "Means for controlling sound output" refers to control technologies that manage the type of sound, volume, playback timing, etc., in an audio device to produce the intended sound.

[0705] "Location information" refers to data that indicates the current location of a specific object, and is often represented as geographical coordinate information.

[0706] "Map information" refers to data that visually represents the location within a building or geographical space, as well as information about its surroundings.

[0707] An "augmented reality display device" is a device or technology that overlays digital information and images onto the real world's field of vision.

[0708] An "encrypted communication method" is a technology that ensures the security of communication by encrypting information using a specific algorithm when transmitting data and decrypting it at the receiving end.

[0709] The system for realizing this invention includes a speech recognition means, an information processing means, and a portable device. The server uses a speech recognition library to receive and analyze voice commands spoken by the user through a microphone. The analyzed voice data is converted into text format and passed to the information processing means. There, an appropriate signal for ringing the portable device is generated based on the voice command.

[0710] The device collects current location information using GPS or wireless signals. The acquired location information is transmitted to a server and, in conjunction with map information, is visually displayed on the user's augmented reality display device, such as smart glasses.

[0711] Users can quickly locate their devices by issuing voice commands. Furthermore, they can efficiently find their valuables by referring to the visually displayed map information.

[0712] To ensure security in this process, encrypted communication methods are used for data transmission and reception. For example, a user can use smart glasses in a cafe and ask the voice assistant to "make my phone ring," allowing them to visually locate their phone. This helps prevent loss and allows for secure location tracking.

[0713] An example of a prompt for a generative AI model is, "I've forgotten where I put my wallet, so I'd like to ask my voice assistant to find out where it is. How do I instruct it?"

[0714] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0715] Step 1:

[0716] The user uses a voice assistant and speaks into a voice input device. The user's voice commands are received as input, converted into digital signals by a speech recognition system, and then sent to a server. Here, a speech recognition library is used to analyze the voice data and generate corresponding command text.

[0717] Step 2:

[0718] The server receives the text generated by the speech recognition library and analyzes its content. Based on the analysis, a command to ring the mobile device is generated. In this information processing process, an appropriate signal is created from the input text data and sent to the mobile device as output.

[0719] Step 3:

[0720] The terminal receives a command from the server. Based on this command, the terminal's acoustic device is activated and emits a pre-set sound. This allows the user to locate the terminal by relying on the sound.

[0721] Step 4:

[0722] The device uses its built-in GPS and wireless signals to collect its own location information. The data obtained from these sensors is processed as input to generate the device's current location information. This location information is sent to a server and integrated with map information as output.

[0723] Step 5:

[0724] The server combines the received location information with map information within the user's building and converts it into visual information. This process involves data processing and calculations to make it visually displayable on an augmented reality display device. Finally, the user can confirm the specific location of their device through the presented visual information.

[0725] Step 6:

[0726] Users can refer to visual feedback from augmented reality displays, using both sound and sight to quickly find their valuables. This allows for decision-making based not only on sound but also on visual cues.

[0727] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0728] This invention provides a system that combines speech recognition and emotion analysis, offering an efficient means of finding a portable device lost inside a house. The system includes speech recognition means, information processing means, an emotion engine, and a portable device. The portable device is equipped with a function to control acoustic output, and by acquiring location information and displaying it on a map of the house, the user can visually understand the device's location.

[0729] System processing flow:

[0730] 1. The user instructs the voice assistant to "make my cell phone ring." The voice assistant collects this voice message and sends it to the server.

[0731] 2. The server analyzes the user's voice received by the speech recognition system. At this time, the emotion engine analyzes the tone, speed, and pitch of the voice to understand the user's emotional state.

[0732] 3. The server generates instructions to ring the portable device based on the voice analysis results. The type and volume of the sound output are adjusted based on the results of the emotion engine.

[0733] 4. The terminal (portable device) receives instructions from the server and outputs the registered melody at an appropriate volume. For example, if the user is emotionally agitated, the volume will increase to attract their attention.

[0734] 5. In parallel with audio output, the terminal collects its own location information using Wi-Fi signals and GPS positioning means and transmits it to the server.

[0735] 6. The server uses location information to update a map of the house and reflects this in an interface that visually displays the location of the mobile device to the user.

[0736] Specific example:

[0737] For example, if a user is working in their home office and forgets where they put their smartphone, they might anxiously ask their voice assistant to "make my phone ring." The voice is analyzed by an emotion engine, and a command is sent to the server to make the phone ring at a volume appropriate to the user's emotion. The voice tone indicates that the user is anxious, so the phone plays a melody at a higher volume than usual. Using this sound signal, the user begins searching for their smartphone and simultaneously checks a map, allowing them to find it quickly. In this way, the user experience can be improved by using thoughtful, emotion-sensitive features.

[0738] The following describes the processing flow.

[0739] Step 1:

[0740] The user instructs the voice assistant to "make my phone ring." This voice command is sent to the server via the device.

[0741] Step 2:

[0742] The server analyzes the received audio data using speech recognition to identify the action the user is requesting. Simultaneously, the emotion engine analyzes the tone, speed, and pitch of the voice to determine the user's emotional state (e.g., calm, anxious, angry).

[0743] Step 3:

[0744] The server determines the type and volume of the sound output according to the user's emotional state and prepares a command to be transmitted to the portable device using information processing means. This command includes information on the melody and volume adjusted according to the emotion.

[0745] Step 4:

[0746] The terminal (portable device) receives commands from the server and starts outputting sound. For example, if it is determined that the user is anxious, it will play a melody at a higher volume than usual.

[0747] Step 5:

[0748] As soon as audio output begins, the device acquires location information using Wi-Fi signals and GPS positioning. This location information is then transmitted from the device to the server.

[0749] Step 6:

[0750] The server analyzes the received location information and combines it with indoor map information to pinpoint the location of the mobile device. The updated map information is reflected in the user-accessible interface, visually indicating the location of the mobile device.

[0751] Step 7:

[0752] Users can quickly locate and retrieve their portable devices using acoustic signals and map information, resulting in a less stressful experience.

[0753] (Example 2)

[0754] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0755] There is a need for efficient methods to quickly locate portable devices when they are lost inside a house or when users are searching for them in a state of confusion. In particular, providing appropriate support in emotionally distressed situations has been difficult with conventional technologies. Considering these circumstances, it is necessary to create technologies that can respond flexibly according to the user's psychological state.

[0756] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0757] In this invention, the server includes information processing means for generating a command to ring the portable device based on user instructions received by voice recognition means, means for analyzing the received voice data and evaluating the user's emotional state using an emotion engine, and means for dynamically adjusting the type and volume of the portable device's sound output based on the command and emotion evaluation generated by the information processing means. This enables accurate visual and auditory support for the user to calmly find the portable device, and facilitates efficient device discovery.

[0758] "Voice recognition means" refers to technology that converts voice commands uttered by a user into digital data and analyzes it.

[0759] "Information processing means" refers to a technology that performs predetermined processing on digital data obtained by voice recognition means to generate commands necessary for operating a portable device.

[0760] An "emotion engine" is a technology that analyzes the characteristics of voice to evaluate the user's emotional state and reflects the results in information processing.

[0761] "Means for dynamically adjusting the type and volume of sound output" refers to a technology that flexibly changes the type and volume of sound output from a portable device based on the evaluation results of the emotion engine.

[0762] "Means for acquiring location information of a portable device" refers to technology that uses communication networks and location measurement technology to determine the precise physical location of a portable device.

[0763] "Visual means of display" refers to a technology that visually displays the location information of a portable device on a map of the house so that users can easily understand it.

[0764] This invention provides a method for efficiently locating a portable device lost inside a house using a system that integrates speech recognition and emotion analysis. Specifically, it employs the following techniques and procedures.

[0765] This system allows users to receive support from a voice assistant if they lose their mobile device. By instructing the voice assistant to "make my phone ring," the voice command is sent to the system. This voice is collected via the device's microphone and transmitted to a server.

[0766] The server converts speech to text using speech recognition software (e.g., a common speech recognition API). Simultaneously, an emotion engine analyzes the characteristics of the speech and evaluates the user's emotional state. Based on this evaluation, it generates commands to adjust the acoustic output of the mobile device. In particular, the volume is automatically adjusted if the user is anxious or confused.

[0767] As a concrete example, consider a scenario where a user has lost their smartphone at home and is searching for it. For instance, if the user anxiously requests, "Make my phone ring," the server, based on its sentiment analysis, assigns a high alert level and instructs the mobile device to emit a loud sound. Relying on the sound signal, the user can quickly locate their smartphone. Furthermore, a map is displayed on the user interface, clearly indicating the location of the mobile device, allowing for simultaneous visual confirmation.

[0768] As an example of a prompt, the AI ​​model might be input with a message like, "Read the user's emotions from their voice and generate instructions to play the device at an appropriate volume." In this way, the present invention enables flexible responses tailored to the user's psychological state, making device discovery more efficient.

[0769] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0770] Step 1:

[0771] The user requests the voice assistant to "make my phone ring." The input is the user's voice, which the device collects using its microphone. The collected voice data is converted into a digital signal and stored on the device.

[0772] Step 2:

[0773] The terminal transmits audio data to the server. The input is digitized audio data, which is then transferred to the server. The output is audio data transmitted over the network. Data is transferred quickly and securely using an internet connection.

[0774] Step 3:

[0775] The server uses speech recognition software to convert speech data into text. The input is speech data received from the terminal, which is processed into text format using a speech recognition algorithm. The output is a text-formatted instruction.

[0776] Step 4:

[0777] The server uses an emotion engine to analyze the tone, speed, and pitch of text data and evaluate the user's emotional state. The input is text data, and the emotion analysis algorithm identifies the user's emotions. The output is the evaluation result of the user's emotional state.

[0778] Step 5:

[0779] The server generates a command to ring the portable device based on the speech recognition results and sentiment evaluation. The input is the analyzed command and sentiment evaluation result, which is used to generate a command that determines the appropriate type and volume of sound output. The output is a specific command message.

[0780] Step 6:

[0781] The terminal (portable device) receives commands from the server and executes an audible output according to those commands. The input is a command message from the server, and the device plays music or an alarm at an appropriate volume using its built-in speaker. The output is a physical acoustic signal.

[0782] Step 7:

[0783] The device collects its own location information using Wi-Fi signals and a location measuring device and transmits it to the server. The input is surrounding signal data, which is processed into location information using a location determination algorithm. The output is location information data sent to the server.

[0784] Step 8:

[0785] The server updates a map of the house based on the received location information, visually displaying the location of the mobile device to the user. The input is location data, which is combined with map data and displayed visually on the user interface. The output is an interface showing the precise location of the mobile device.

[0786] (Application Example 2)

[0787] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0788] In recent years, efficiently locating personal electronic devices that are left behind or lost inside vehicles has become a challenge. In particular, it is difficult for passengers to find lost devices themselves in the complex environment of a moving vehicle. Furthermore, the lack of visual means to determine the location of devices necessitates a rapid resolution of the problem.

[0789] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0790] In this invention, the server includes data processing means that generate a command to sound a personal electronic device based on user instructions and emotion analysis received by voice recognition means; means that control the sound output of the personal electronic device in response to the command generated by the data processing means; and means that acquire location data of the personal electronic device and display it visually in conjunction with map information within the mobile vehicle. This makes it possible to efficiently find a personal electronic device that has been lost within the mobile vehicle.

[0791] "Voice recognition means" refers to a device or program that receives voice instructions from a user and processes them as digital data.

[0792] "Emotional analysis" is a technology that analyzes and evaluates a person's emotional state through voice, facial expressions, and text.

[0793] A "personal electronic device" is a portable electronic device such as a smartphone or tablet, which is a digital terminal used by an individual.

[0794] "Data processing means" refers to a device or program that has the function of receiving information, analyzing it, and generating commands or responses.

[0795] "Means for controlling sound output" refers to a device or program that has the function of setting and adjusting alarm sounds or melodies to be played at an appropriate volume and type.

[0796] "Location data" refers to information that indicates the geographical or spatial location of a specific object.

[0797] "Map information" refers to data that visually represents the geographical elements of a specific space.

[0798] The term "mobile object" refers to a vehicle or its components that can move by carrying people or objects.

[0799] The system that realizes this invention is a technology for quickly locating personal electronic devices such as smartphones and tablets when they are lost inside a vehicle. This system consists of voice recognition means, emotion analysis means, data processing means, sound output control means, location data acquisition means, and map information display means.

[0800] The server first uses a speech recognition API (e.g., Google Speech-to-Text) to acquire and analyze the user's voice commands, and converts the command content into digital data. It also uses an emotion analysis engine (e.g., Affectiva, IBM Watson Tone Analyzer) to read the user's emotional state from their voice and incorporates this into the data processing. Based on this emotional information, the data processing system generates a command for the optimal acoustic output and transmits it to the personal electronic device.

[0801] Personal electronic devices (smartphones and tablets) receive commands from the server, adjust the volume and sound type, and then sound an alarm. During this process, they collect location data using Bluetooth, Wi-Fi signal strength, and in some cases, the Global Positioning System (GPS), and send this information back to the server. The server then updates the in-vehicle map information based on this location data, providing the user with an interface that intuitively shows the device's location.

[0802] For example, if a user loses sight of their smartphone that they placed on the passenger seat while driving, they can instruct the voice assistant inside the car to "find my smartphone in the car." If the emotion analysis engine detects that the user is in a hurry, the smartphone will emit a loud sound and its location will be displayed on the in-car monitor, allowing the user to find their smartphone immediately.

[0803] An example of a prompt for a generative AI model is: "I want to develop a system to efficiently find a smart device that has been lost inside a car. I want it to use speech recognition and sentiment analysis to visually indicate the device's location."

[0804] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0805] Step 1:

[0806] The user gives instructions to the in-car voice recognition system. The voice command, "Find my smart device," is taken as input, converted into digital data by the voice recognition API, and that data is sent to the server.

[0807] Step 2:

[0808] The server analyzes the received audio data. Using the speech recognition results as input, the emotion analysis engine evaluates the emotional state of the speech (e.g., anxiety, calmness). This process generates two outputs: an audio instruction and an emotional state.

[0809] Step 3:

[0810] Based on the analyzed voice commands and emotional state, the server uses data processing to generate a command for the appropriate acoustic output. This command includes the type and volume of sound for the device. The generated command is sent to the terminal as output data.

[0811] Step 4:

[0812] The terminal (personal electronic device) receives an audio output command from the server. Based on this input, the device sounds an alarm at the specified volume and type. The execution of the audio output results in a physical sound being emitted.

[0813] Step 5:

[0814] Simultaneously, the device collects its own location data. It obtains location information using Wi-Fi signals, Bluetooth, and in some cases, the Global Positioning System (GPS). The acquired location data is then transferred to a server as output.

[0815] Step 6:

[0816] The server updates the in-vehicle map information based on the transmitted location data. Using the location data as input, it generates a user interface that shows the device's location. This interface is displayed on the in-vehicle monitor, providing a visual output.

[0817] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0818] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0819] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0820] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0821] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0822] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0823] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0824] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0825] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0826] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0827] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0828] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0829] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0830] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0831] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0832] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0833] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0834] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0835] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0836] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0837] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0838] The following is further disclosed regarding the embodiments described above.

[0839] (Claim 1)

[0840] Information processing means that generates a command to ring a portable device based on user instructions received by voice recognition means,

[0841] Means for controlling the acoustic output of a portable device in response to a command generated by the information processing means,

[0842] A means of acquiring location information from a portable device and visually displaying it in conjunction with map information inside a house,

[0843] A system that includes this.

[0844] (Claim 2)

[0845] The system according to claim 1, further comprising means for acquiring the aforementioned location information by combining a Wi-Fi signal and a GPS position measurement means.

[0846] (Claim 3)

[0847] The system according to claim 1, wherein the information processing means includes means for communicating with a portable device using an encrypted communication protocol.

[0848] "Example 1"

[0849] (Claim 1)

[0850] A processing means that receives voice input, interprets user instructions using natural language processing, and generates a signal to call a mobile device,

[0851] Control means for controlling the acoustic output device of a mobile terminal in response to the signal generated by the processing means,

[0852] A display means that acquires location information from a mobile device, integrates the location data with spatial information inside the house, and displays it visually.

[0853] A means of notifying the user to visually and audibly locate the device by measuring its position when sound output is initiated,

[0854] A system that includes this.

[0855] (Claim 2)

[0856] The system according to claim 1, further comprising means for acquiring the aforementioned location information using a wireless signal and a global positioning system position measurement means.

[0857] (Claim 3)

[0858] The system according to claim 1, wherein the processing means includes means for performing secure communication with a mobile terminal using data protection technology.

[0859] "Application Example 1"

[0860] (Claim 1)

[0861] Information processing means that generates a command to ring a portable device based on user instructions received by voice recognition means,

[0862] Means for controlling the acoustic output of a portable device in response to a command generated by the information processing means,

[0863] A means of acquiring location information from a portable device and visually displaying it in conjunction with map information within a building,

[0864] A means for visually displaying the location of valuables on an augmented reality display device based on voice commands,

[0865] A system that includes this.

[0866] (Claim 2)

[0867] The system according to claim 1, further comprising means for acquiring the aforementioned location information by combining a wireless signal and a location measurement means.

[0868] (Claim 3)

[0869] The system according to claim 1, wherein the information processing means comprises means for communicating with a portable device using an encrypted communication method.

[0870] "Example 2 of combining an emotion engine"

[0871] (Claim 1)

[0872] Information processing means that generates a command to ring a portable device based on user instructions received by voice recognition means,

[0873] A means of analyzing received audio data and evaluating the user's emotional state using an emotion engine,

[0874] A means for dynamically adjusting the type and volume of the sound output of a portable device based on the commands and emotional evaluations generated by the aforementioned information processing means,

[0875] A means of acquiring location information from a portable device and visually displaying it in conjunction with map information inside a house,

[0876] A system that includes this.

[0877] (Claim 2)

[0878] The system according to claim 1, further comprising means for acquiring the aforementioned location information by combining a wireless signal and a location measurement means.

[0879] (Claim 3)

[0880] The system according to claim 1, wherein the information processing means includes means for communicating with a portable device using an encrypted communication protocol.

[0881] "Application example 2 when combining with an emotional engine"

[0882] (Claim 1)

[0883] A data processing means that generates a command to sound a personal electronic device based on the user's instructions and emotion analysis received by a voice recognition means,

[0884] Means for controlling the sound output of a personal electronic device in response to commands generated by the data processing means,

[0885] A means of acquiring location data of personal electronic devices and visually displaying it in conjunction with map information within a mobile device,

[0886] A system that includes this.

[0887] (Claim 2)

[0888] The system according to claim 1, further comprising means for acquiring the aforementioned position data by combining a wireless signal and a position measurement means of a global positioning system.

[0889] (Claim 3)

[0890] The system according to claim 1, wherein the data processing means includes means for communicating with a personal electronic device using a security communication protocol. [Explanation of symbols]

[0891] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Information processing means that generates a command to ring a portable device based on user instructions received by voice recognition means, Means for controlling the acoustic output of a portable device in response to a command generated by the information processing means, A means of acquiring location information from a portable device and visually displaying it in conjunction with map information inside a house, A system that includes this.

2. The system according to claim 1, further comprising means for acquiring the aforementioned location information by combining a Wi-Fi signal and a GPS position measurement means.

3. The system according to claim 1, wherein the information processing means includes means for communicating with a portable device using an encrypted communication protocol.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A