system

The system addresses the challenge of manually setting call tones by using a selection, detection, and adjustment unit to automatically adapt the ringtone based on user preferences and ambient sounds, providing a personalized and context-aware experience.

JP2026066711APending Publication Date: 2026-04-17SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-07
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Conventional systems lack the ability to automatically select and adjust an incoming call tone based on user preferences and situational context, requiring manual setting.

Method used

A system comprising a selection unit, detection unit, and adjustment unit that selects and adjusts the incoming call tone based on user preferences and ambient sound detection, using AI to learn and adapt to user feedback.

Benefits of technology

Automatically selects and adjusts the optimal ringtone according to user preferences and circumstances, ensuring a personalized and context-aware experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026066711000001_ABST
    Figure 2026066711000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to automatically select and adjust the optimal ringtone according to the user's preferences and circumstances. [Solution] The system according to the embodiment comprises a selection unit, a detection unit, and an adjustment unit. The selection unit selects a ringtone based on the user's preference or the user's current situation. The detection unit detects ambient sounds around the user. The adjustment unit adjusts the volume or type of the selected ringtone based on the ambient sounds detected by the detection unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional technology, there is a problem that the incoming call tone cannot be automatically selected and adjusted according to the user's preference and situation, and it is necessary to set it manually.

[0005] The system according to the embodiment aims to automatically select and adjust an optimal incoming call tone according to the user's preference and situation.

Means for Solving the Problems

[0006] The system according to the embodiment includes a selection unit, a detection unit, and an adjustment unit. The selection unit selects an incoming call tone based on the user's preference or the user's current situation. The detection unit detects the ambient sound around the user. The adjustment unit adjusts the volume or type of the selected incoming call tone based on the ambient sound detected by the detection unit. [Effects of the Invention]

[0007] The system according to this embodiment can automatically select and adjust the optimal ringtone according to the user's preferences and circumstances. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the reception device 38, the output device 40, and the camera 42 are connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) An AI assistant system according to an embodiment of the present invention is a system that automatically selects the optimal ringtone according to the user's preferences and circumstances. The AI ​​assistant system selects a ringtone based on the user's preferences and the situation in which the user is located. Next, the AI ​​assistant system detects ambient sounds around the user and adjusts the volume or type of the selected ringtone. This mechanism ensures that the user always receives the optimal ringtone. For example, if the user is in a quiet place, the AI ​​assistant system will select quiet music, and if they are in a noisy place, it will select louder music. This selection is learned by the AI ​​based on user feedback to select the optimal ringtone. Next, the AI ​​assistant system detects ambient sounds around the user. For example, it collects ambient sounds using a microphone and the AI ​​analyzes those sounds. This allows the system to understand the situation in which the user is located. Furthermore, the AI ​​assistant system adjusts the volume or type of the selected ringtone based on the detected ambient sounds. For example, it lowers the volume in quiet places and increases the volume in noisy places. It can also prioritize the selection of a specific genre of music. For example, if the user prefers rock music, it will prioritize the selection of rock music. It can also select a ringtone suitable for a specific time of day. For example, it can select quiet music at night and upbeat music during the day. In this way, it can provide ringtones that match the user's lifestyle. Furthermore, the AI ​​assistant system can estimate the user's emotions and select ringtones based on those emotions. For example, if the user is relaxed, it will select relaxing music, and if they are stressed, it will select refreshing music. Finally, the AI ​​assistant system can analyze the user's past selection history and select ringtones appropriate for specific events or situations. For example, if the user selected a particular piece of music for a particular event, it will select the same music when attending that event again. In this way, it can provide optimal ringtones tailored to the user's preferences and circumstances.

[0029] The AI ​​assistant system according to this embodiment comprises a selection unit, a detection unit, and an adjustment unit. The selection unit selects a ringtone based on the user's preference or the user's current situation. For example, if the user is in a quiet place, the selection unit will select quiet music. The selection unit may also select louder music if the user is in a noisy place. Furthermore, the selection unit can learn from the user's feedback and select the optimal ringtone. For example, the selection unit will analyze the history of ringtones the user has previously selected and select the same ringtone in similar situations. The detection unit detects ambient sounds around the user. For example, the detection unit will collect ambient sounds using a microphone. The detection unit can analyze the collected ambient sounds to understand the situation the user is in. For example, the detection unit will detect that there are few ambient sounds in a quiet place and many ambient sounds in a noisy place. The adjustment unit adjusts the volume or type of the selected ringtone based on the ambient sounds detected by the detection unit. For example, the adjustment unit will lower the volume in a quiet place and increase the volume in a noisy place. Furthermore, the adjustment unit can prioritize selecting music of a specific genre. For example, if the user prefers rock music, the adjustment unit will prioritize selecting rock music. In addition, the adjustment unit can select a ringtone suitable for a specific time of day. For example, the adjustment unit will select quiet music at night and lively music during the day. This allows the AI ​​assistant system according to the embodiment to provide the optimal ringtone according to the user's preferences and circumstances.

[0030] The selection unit chooses a ringtone based on the user's preferences or current situation. For example, if the user is in a quiet place, the selection unit will choose quiet music. Conversely, if the user is in a noisy place, the selection unit can also choose louder music. Specifically, the selection unit determines the user's current environment based on location information and environmental sensor data obtained from the user's smartphone or wearable device. For example, it analyzes GPS data and Wi-Fi signal strength to determine whether the user is in a quiet place such as a library or cafe, or a noisy place such as a train station or shopping mall. The selection unit also learns the user's past behavior and selection history to understand their preferences. For example, it analyzes ringtone patterns the user has selected in specific situations in the past and suggests the same ringtone in similar situations. Furthermore, the selection unit can learn from user feedback and select the optimal ringtone. For example, if the user provides feedback such as "like" or "dislike" regarding a particular ringtone, the selection unit learns this information and reflects it in future selections. This allows the selection unit to provide the optimal ringtone according to the user's preferences and circumstances.

[0031] The detection unit detects ambient sounds around the user. The detection unit collects ambient sounds, for example, using a microphone. Specifically, it uses a microphone built into a smartphone or wearable device to collect ambient sounds in real time. The detection unit analyzes the collected ambient sounds to understand the situation the user is in. For example, the detection unit analyzes ambient sounds using speech recognition technology to detect that there are few ambient sounds in quiet places and many ambient sounds in noisy places. Furthermore, the detection unit can also identify specific sound patterns. For example, it can identify car engine sounds, people talking, types of music, etc., and based on that, it can understand the user's situation in more detail. As a result, the detection unit can accurately understand the environment the user is currently in and provide appropriate information to the selection and adjustment units.

[0032] The adjustment unit adjusts the volume or type of the selected ringtone based on the ambient sound detected by the detection unit. For example, the adjustment unit lowers the volume in quiet places and increases it in noisy places. Specifically, the adjustment unit applies a volume adjustment algorithm to set the optimal volume based on the ambient sound data provided by the detection unit. The adjustment unit can also prioritize the selection of a specific genre of music. For example, if the user prefers rock music, the adjustment unit will prioritize rock music. Furthermore, the adjustment unit can select a ringtone suitable for a specific time of day. For example, the adjustment unit will select quiet music at night and lively music during the day. This allows the adjustment unit to provide the optimal ringtone according to the user's preferences and circumstances. In addition, the adjustment unit can fine-tune the volume and type of music based on user feedback. For example, if the user provides feedback such as "too loud" or "too quiet" for a particular volume or type of music, the adjustment unit learns this information and reflects it in subsequent adjustments. This allows the adjustment unit to provide the optimal ringtone according to the user's preferences and circumstances.

[0033] The selection unit can learn which ringtones to select based on user feedback. For example, it can analyze the user's past ringtone selection history and select the same ringtone in similar situations. It can also prioritize music of a specific genre if the user prefers it. Furthermore, it can select ringtones suitable for specific times of day. For example, it might select quiet music at night and lively music during the day. This allows it to select more appropriate ringtones by learning from user feedback. Some or all of the above processing in the selection unit may be performed using AI, or not. For example, the selection unit can input user feedback data into a generating AI and have the generating AI select the optimal ringtone.

[0034] The detection unit can collect ambient sounds using a microphone. For example, the detection unit can collect ambient sounds using a built-in microphone. The detection unit can also collect ambient sounds using an external microphone. Furthermore, the detection unit can collect ambient sounds using multiple microphones. For example, the detection unit can arrange multiple microphones to collect ambient sounds in three dimensions. This allows for an accurate understanding of the surrounding situation by collecting ambient sounds using microphones. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input the collected ambient sound data into a generating AI and have the generating AI perform an analysis of the ambient sounds.

[0035] The selection unit can prioritize selecting a specific music genre. For example, if the user prefers classical music, the selection unit will prioritize classical music. It can also prioritize pop music if the user prefers pop music. Furthermore, if the user prefers rock music, the selection unit can prioritize rock music. For example, the selection unit can learn the user's music genre preferences and prioritize music of that genre. This allows the system to provide ringtones that match the user's preferences by prioritizing specific genres. Some or all of the above processing in the selection unit may be performed using AI, or not. For example, the selection unit can input the user's music genre preference data into a generating AI and have the generating AI select the optimal music genre.

[0036] The selection unit can select a ringtone suitable for a specific time of day. For example, the selection unit can select upbeat music in the morning. It can also select lively music during the day. Furthermore, it can select quiet music at night. For example, the selection unit can learn music suitable for a specific time of day and select music appropriate for that time. This allows the selection unit to provide a ringtone that matches the user's daily rhythm by selecting a ringtone suitable for a specific time of day. Some or all of the above processing in the selection unit may be performed using AI, for example, or without AI. For example, the selection unit can input music data suitable for a specific time of day into a generating AI and have the generating AI select the music that is best suited to that time of day.

[0037] The selection unit can analyze the user's past selection history and select a ringtone appropriate for a specific event or situation. For example, if the user selected a particular piece of music for a specific event, the selection unit will select the same music when the user attends that event again. Similarly, if the user selected a particular piece of music for a specific situation, the selection unit can select the same music again when the user encounters that situation again. Furthermore, the selection unit can learn from the user's past selection history, identify specific patterns, and select the optimal ringtone based on those patterns. For example, the selection unit can analyze the user's past music selection history and select the same music in similar situations. This allows the selection unit to provide ringtones appropriate for specific events or situations by analyzing past selection history. Some or all of the above processing in the selection unit may be performed using AI, or not. For example, the selection unit can input the user's past selection history data into a generating AI and have the generating AI select the optimal ringtone.

[0038] The selection unit can analyze the user's past selection history and select a ringtone appropriate for a specific season or weather. For example, the selection unit might select warm music in winter, relaxing music on rainy days, and refreshing music in summer. By analyzing the user's past selection history and selecting the optimal ringtone for a specific season or weather, the selection unit can provide a ringtone that suits the user's situation. Some or all of the above processing in the selection unit may be performed using AI, for example, or without AI. For example, the selection unit can input the user's past selection history data into a generating AI and have the generating AI select the optimal ringtone for the season or weather.

[0039] The selection unit can select an appropriate ringtone based on the user's current activity. For example, if the user is exercising, the selection unit may select energetic music. It may also select music to enhance concentration if the user is working. Furthermore, it may select relaxing music if the user is taking a break. For instance, the selection unit can detect the user's current activity and select the optimal ringtone based on that activity. This allows the system to provide a ringtone that suits the user's situation by offering the most suitable ringtone based on the user's current activity. Some or all of the above processing in the selection unit may be performed using AI, for example, or without AI. For example, the selection unit can input the user's current activity data into a generating AI and have the generating AI select the optimal ringtone.

[0040] The selection unit can select a ringtone appropriate to the culture and customs of the region based on the user's geographical location. For example, if the user is in Japan, the selection unit can select Japanese-style music. It can also select country music if the user is in the United States. Furthermore, if the user is in India, it can select Bollywood music. In short, the selection unit selects the optimal ringtone appropriate to the culture and customs of the region based on the user's geographical location. This allows for the provision of ringtones appropriate to the local culture and customs by selecting ringtones based on geographical location. Some or all of the above processing in the selection unit may be performed using AI, for example, or without AI. For example, the selection unit can input the user's geographical location data into a generating AI and have the generating AI select the optimal ringtone appropriate to the culture and customs of the region.

[0041] The selection unit can analyze the user's social media activity and select a ringtone based on trends. For example, the selection unit can select a ringtone based on music the user has recently shared. It can also select a ringtone based on music preferred by the user's followers. Furthermore, it can select a ringtone based on music related to events the user is participating in. For example, the selection unit analyzes the user's social media activity and selects the optimal ringtone based on trends. This allows the selection of ringtones based on social media activity to provide ringtones that are in line with current trends. Some or all of the above processing in the selection unit may be performed using AI, for example, or not. For example, the selection unit can input the user's social media activity data into a generating AI and have the generating AI select the optimal trend-based ringtone.

[0042] The detection unit can analyze collected ambient sounds in real time and respond immediately to sudden noises. For example, the detection unit can detect a sudden car horn sound and temporarily increase the ringtone volume. It can also detect a sudden dog bark and change the type of ringtone. Furthermore, it can detect a sudden human scream and temporarily decrease the ringtone volume. For example, the detection unit analyzes collected ambient sounds in real time and responds immediately to sudden noises. This allows the user to receive an appropriate ringtone by responding immediately to sudden noises. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input collected ambient sound data into a generating AI and have the generating AI perform the detection and response to sudden noises.

[0043] The detection unit can classify the collected ambient sounds and perform special processing on specific sounds. For example, the detection unit can detect a car horn and increase the volume of the ringtone. It can also detect a dog barking and change the type of ringtone. Furthermore, it can detect human speech and decrease the volume of the ringtone. For example, the detection unit classifies the collected ambient sounds and performs special processing on specific sounds. This allows the user to receive an appropriate ringtone by performing special processing on specific sounds. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input the collected ambient sound data into a generating AI and have the generating AI perform the classification and processing of specific sounds.

[0044] The detection unit can adjust the collected ambient sound based on the battery level of the user's device. For example, if the battery level is low, the detection unit can set the ambient sound detection frequency to a low level. Conversely, if the battery level is sufficient, the detection unit can also set the ambient sound detection frequency to a high level. Furthermore, if the battery level is moderate, the detection unit can also set the ambient sound detection frequency to a moderate level. For example, the detection unit adjusts the collected ambient sound based on the battery level of the user's device. This allows for efficient use of the device's battery by adjusting the ambient sound detection frequency based on the battery level. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input the collected ambient sound data into a generating AI and have the generating AI perform adjustments based on the battery level.

[0045] The detection unit can combine collected ambient sounds with the user's schedule information to perform detections tailored to specific time periods. For example, the detection unit can set a higher detection frequency for ambient sounds when the user is in a meeting. It can also set a lower detection frequency when the user is on a break. Furthermore, it can set a moderate detection frequency when the user is exercising. For example, the detection unit combines collected ambient sounds with the user's schedule information to perform detections tailored to specific time periods. By adjusting the detection frequency of ambient sounds based on the schedule information, more appropriate detection of ambient sounds becomes possible. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input collected ambient sound data into a generating AI and have the generating AI perform detections based on the schedule information.

[0046] The adjustment unit can optimize the sound quality of the selected ringtone based on the speaker characteristics of the user's device. For example, if a high-quality speaker is being used, the adjustment unit will set the sound quality to a high level. It can also set the sound quality to a low level if a low-quality speaker is being used. Furthermore, if a medium-quality speaker is being used, the adjustment unit can set the sound quality to a medium level. For example, the adjustment unit optimizes the sound quality of the selected ringtone based on the speaker characteristics of the user's device. This allows for the provision of better-sounding ringtones by optimizing the sound quality based on the speaker characteristics. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or without AI. For example, the adjustment unit can input the speaker characteristic data of the user's device into a generating AI and have the generating AI perform the optimal sound quality setting.

[0047] The adjustment unit can adjust the duration of the selected ringtone based on the user's current situation. For example, the adjustment unit may set the ringtone duration shorter if the user is in a meeting. It may also set the ringtone duration longer if the user is driving. Furthermore, it may set the ringtone duration to a medium level if the user is taking a break. For example, the adjustment unit adjusts the duration of the selected ringtone based on the user's current situation. This allows for the provision of a more appropriate ringtone by adjusting the ringtone duration based on the user's current situation. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or without AI. For example, the adjustment unit can input the user's current situation data into a generating AI and have the generating AI set the optimal duration.

[0048] The adjustment unit can optimize the volume of the selected ringtone based on the battery level of the user's device. For example, if the battery level is low, the adjustment unit can set the ringtone volume low. Conversely, if the battery level is sufficient, the adjustment unit can also set the ringtone volume high. Furthermore, if the battery level is moderate, the adjustment unit can also set the ringtone volume to a moderate level. For example, the adjustment unit optimizes the volume of the selected ringtone based on the battery level of the user's device. This allows for efficient use of the device's battery by optimizing the volume based on the battery level. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or without AI. For example, the adjustment unit can input the battery level data of the user's device into a generating AI and have the generating AI perform the optimal volume setting.

[0049] The adjustment unit can adjust the selected ringtone type in conjunction with the user's schedule information to optimize it for specific events. For example, the adjustment unit can select a quiet ringtone during a meeting. It can also select an energetic ringtone during exercise. Furthermore, it can select a relaxing ringtone during a break. For example, the adjustment unit adjusts the selected ringtone type in conjunction with the user's schedule information to optimize it for specific events. This allows the adjustment unit to provide a ringtone suitable for specific events by adjusting the ringtone type based on the schedule information. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or without AI. For example, the adjustment unit can input the user's schedule information data into a generating AI and have the generating AI set the optimal ringtone type.

[0050] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0051] The selection unit can choose a ringtone based on the user's preferences, circumstances, and activity history. For example, if the user previously selected a specific song at a particular location, the unit will select the same song when the user revisits that location. Similarly, if the user selected a specific song during a particular time period, the unit can select the same song again during that time period. Furthermore, if the user selected a specific song at a particular event, the unit can select the same song again when the user attends that event again. This allows the system to provide the most suitable ringtone based on the user's activity history.

[0052] The detection unit can detect the temperature and humidity around the user and adjust the ringtone accordingly. For example, if the ambient temperature is high, the detection unit can select cool music. It can also select refreshing music if the ambient humidity is high. Furthermore, if the ambient temperature is low, it can select warm music. This allows the system to provide the optimal ringtone based on ambient temperature and humidity.

[0053] The selection function can choose a ringtone based on the user's preferences and circumstances, as well as the battery level of the user's device. For example, if the battery level is low, the selection function will select a short ringtone. If the battery level is sufficient, it can also select a long ringtone. Furthermore, if the battery level is moderate, it can select a ringtone of medium length. This allows the system to provide the optimal ringtone based on the device's battery level.

[0054] The selection function can choose a ringtone based on the user's preferences, circumstances, and schedule information. For example, if the user is in a meeting, the selection function will select a quiet ringtone. It can also select an energetic ringtone if the user is exercising. Furthermore, it can select a relaxing ringtone if the user is taking a break. This allows the system to provide the optimal ringtone based on the user's schedule information.

[0055] The following briefly describes the processing flow for example form 1.

[0056] Step 1: The selection unit chooses a ringtone based on the user's preferences or current situation. For example, if the user is in a quiet place, it will choose quiet music; if they are in a noisy place, it will choose louder music. The selection unit can also learn from the user's feedback and select the optimal ringtone. For example, it will analyze the user's past ringtone selection history and select the same ringtone in similar situations. Step 2: The detection unit detects ambient noise around the user. For example, it uses a microphone to collect ambient noise and analyzes the collected noise to understand the situation the user is in. It detects that there is little ambient noise in quiet places and a lot of ambient noise in noisy places. Step 3: The adjustment unit adjusts the volume or type of the selected ringtone based on the ambient sound detected by the detection unit. For example, it lowers the volume in quiet places and increases it in noisy places. It can also prioritize the selection of a specific genre of music. Furthermore, it can select a ringtone suitable for a specific time of day. For example, it can select quiet music at night and lively music during the day.

[0057] (Example of form 2) An AI assistant system according to an embodiment of the present invention is a system that automatically selects the optimal ringtone according to the user's preferences and circumstances. The AI ​​assistant system selects a ringtone based on the user's preferences and the situation in which the user is located. Next, the AI ​​assistant system detects ambient sounds around the user and adjusts the volume or type of the selected ringtone. This mechanism ensures that the user always receives the optimal ringtone. For example, if the user is in a quiet place, the AI ​​assistant system will select quiet music, and if they are in a noisy place, it will select louder music. This selection is learned by the AI ​​based on user feedback to select the optimal ringtone. Next, the AI ​​assistant system detects ambient sounds around the user. For example, it collects ambient sounds using a microphone and the AI ​​analyzes those sounds. This allows the system to understand the situation in which the user is located. Furthermore, the AI ​​assistant system adjusts the volume or type of the selected ringtone based on the detected ambient sounds. For example, it lowers the volume in quiet places and increases the volume in noisy places. It can also prioritize the selection of a specific genre of music. For example, if the user prefers rock music, it will prioritize the selection of rock music. It can also select a ringtone suitable for a specific time of day. For example, it can select quiet music at night and upbeat music during the day. In this way, it can provide ringtones that match the user's lifestyle. Furthermore, the AI ​​assistant system can estimate the user's emotions and select ringtones based on those emotions. For example, if the user is relaxed, it will select relaxing music, and if they are stressed, it will select refreshing music. Finally, the AI ​​assistant system can analyze the user's past selection history and select ringtones appropriate for specific events or situations. For example, if the user selected a particular piece of music for a particular event, it will select the same music when attending that event again. In this way, it can provide optimal ringtones tailored to the user's preferences and circumstances.

[0058] The AI ​​assistant system according to this embodiment comprises a selection unit, a detection unit, and an adjustment unit. The selection unit selects a ringtone based on the user's preference or the user's current situation. For example, if the user is in a quiet place, the selection unit will select quiet music. The selection unit may also select louder music if the user is in a noisy place. Furthermore, the selection unit can learn from the user's feedback and select the optimal ringtone. For example, the selection unit will analyze the history of ringtones the user has previously selected and select the same ringtone in similar situations. The detection unit detects ambient sounds around the user. For example, the detection unit will collect ambient sounds using a microphone. The detection unit can analyze the collected ambient sounds to understand the situation the user is in. For example, the detection unit will detect that there are few ambient sounds in a quiet place and many ambient sounds in a noisy place. The adjustment unit adjusts the volume or type of the selected ringtone based on the ambient sounds detected by the detection unit. For example, the adjustment unit will lower the volume in a quiet place and increase the volume in a noisy place. Furthermore, the adjustment unit can prioritize selecting music of a specific genre. For example, if the user prefers rock music, the adjustment unit will prioritize selecting rock music. In addition, the adjustment unit can select a ringtone suitable for a specific time of day. For example, the adjustment unit will select quiet music at night and lively music during the day. This allows the AI ​​assistant system according to the embodiment to provide the optimal ringtone according to the user's preferences and circumstances.

[0059] The selection unit chooses a ringtone based on the user's preferences or current situation. For example, if the user is in a quiet place, the selection unit will choose quiet music. Conversely, if the user is in a noisy place, the selection unit can also choose louder music. Specifically, the selection unit determines the user's current environment based on location information and environmental sensor data obtained from the user's smartphone or wearable device. For example, it analyzes GPS data and Wi-Fi signal strength to determine whether the user is in a quiet place such as a library or cafe, or a noisy place such as a train station or shopping mall. The selection unit also learns the user's past behavior and selection history to understand their preferences. For example, it analyzes ringtone patterns the user has selected in specific situations in the past and suggests the same ringtone in similar situations. Furthermore, the selection unit can learn from user feedback and select the optimal ringtone. For example, if the user provides feedback such as "like" or "dislike" regarding a particular ringtone, the selection unit learns this information and reflects it in future selections. This allows the selection unit to provide the optimal ringtone according to the user's preferences and circumstances.

[0060] The detection unit detects ambient sounds around the user. The detection unit collects ambient sounds, for example, using a microphone. Specifically, it uses a microphone built into a smartphone or wearable device to collect ambient sounds in real time. The detection unit analyzes the collected ambient sounds to understand the situation the user is in. For example, the detection unit analyzes ambient sounds using speech recognition technology to detect that there are few ambient sounds in quiet places and many ambient sounds in noisy places. Furthermore, the detection unit can also identify specific sound patterns. For example, it can identify car engine sounds, people talking, types of music, etc., and based on that, it can understand the user's situation in more detail. As a result, the detection unit can accurately understand the environment the user is currently in and provide appropriate information to the selection and adjustment units.

[0061] The adjustment unit adjusts the volume or type of the selected ringtone based on the ambient sound detected by the detection unit. For example, the adjustment unit lowers the volume in quiet places and increases it in noisy places. Specifically, the adjustment unit applies a volume adjustment algorithm to set the optimal volume based on the ambient sound data provided by the detection unit. The adjustment unit can also prioritize the selection of a specific genre of music. For example, if the user prefers rock music, the adjustment unit will prioritize rock music. Furthermore, the adjustment unit can select a ringtone suitable for a specific time of day. For example, the adjustment unit will select quiet music at night and lively music during the day. This allows the adjustment unit to provide the optimal ringtone according to the user's preferences and circumstances. In addition, the adjustment unit can fine-tune the volume and type of music based on user feedback. For example, if the user provides feedback such as "too loud" or "too quiet" for a particular volume or type of music, the adjustment unit learns this information and reflects it in subsequent adjustments. This allows the adjustment unit to provide the optimal ringtone according to the user's preferences and circumstances.

[0062] The selection unit can learn which ringtones to select based on user feedback. For example, it can analyze the user's past ringtone selection history and select the same ringtone in similar situations. It can also prioritize music of a specific genre if the user prefers it. Furthermore, it can select ringtones suitable for specific times of day. For example, it might select quiet music at night and lively music during the day. This allows it to select more appropriate ringtones by learning from user feedback. Some or all of the above processing in the selection unit may be performed using AI, or not. For example, the selection unit can input user feedback data into a generating AI and have the generating AI select the optimal ringtone.

[0063] The detection unit can collect ambient sounds using a microphone. For example, the detection unit can collect ambient sounds using a built-in microphone. The detection unit can also collect ambient sounds using an external microphone. Furthermore, the detection unit can collect ambient sounds using multiple microphones. For example, the detection unit can arrange multiple microphones to collect ambient sounds in three dimensions. This allows for an accurate understanding of the surrounding situation by collecting ambient sounds using microphones. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input the collected ambient sound data into a generating AI and have the generating AI perform an analysis of the ambient sounds.

[0064] The selection unit can prioritize selecting a specific music genre. For example, if the user prefers classical music, the selection unit will prioritize classical music. It can also prioritize pop music if the user prefers pop music. Furthermore, if the user prefers rock music, the selection unit can prioritize rock music. For example, the selection unit can learn the user's music genre preferences and prioritize music of that genre. This allows the system to provide ringtones that match the user's preferences by prioritizing specific genres. Some or all of the above processing in the selection unit may be performed using AI, or not. For example, the selection unit can input the user's music genre preference data into a generating AI and have the generating AI select the optimal music genre.

[0065] The selection unit can select a ringtone suitable for a specific time of day. For example, the selection unit can select upbeat music in the morning. It can also select lively music during the day. Furthermore, it can select quiet music at night. For example, the selection unit can learn music suitable for a specific time of day and select music appropriate for that time. This allows the selection unit to provide a ringtone that matches the user's daily rhythm by selecting a ringtone suitable for a specific time of day. Some or all of the above processing in the selection unit may be performed using AI, for example, or without AI. For example, the selection unit can input music data suitable for a specific time of day into a generating AI and have the generating AI select the music that is best suited to that time of day.

[0066] The selection unit can estimate the user's emotions and select a ringtone based on those emotions. For example, if the user is relaxed, the selection unit can select relaxing music. If the user is stressed, the selection unit can also select refreshing music. Furthermore, if the user is concentrating, the selection unit can select music that helps maintain concentration. For example, the selection unit can estimate the user's emotions using technologies such as facial recognition or voice analysis, and select the optimal ringtone based on those emotions. This allows for the provision of more appropriate ringtones by selecting ringtones based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the selection unit may be performed using AI, or not using AI. For example, the selection unit can input the user's emotion data into a generative AI and have the generative AI select the optimal ringtone.

[0067] The selection unit can analyze the user's past selection history and select a ringtone appropriate for a specific event or situation. For example, if the user selected a particular piece of music for a specific event, the selection unit will select the same music when the user attends that event again. Similarly, if the user selected a particular piece of music for a specific situation, the selection unit can select the same music again when the user encounters that situation again. Furthermore, the selection unit can learn from the user's past selection history, identify specific patterns, and select the optimal ringtone based on those patterns. For example, the selection unit can analyze the user's past music selection history and select the same music in similar situations. This allows the selection unit to provide ringtones appropriate for specific events or situations by analyzing past selection history. Some or all of the above processing in the selection unit may be performed using AI, or not. For example, the selection unit can input the user's past selection history data into a generating AI and have the generating AI select the optimal ringtone.

[0068] The selection unit can estimate the user's emotions and adjust the tempo of the ringtone based on the estimated emotions. For example, if the user is relaxed, the selection unit can select a slow-tempo ringtone. If the user is stressed, the selection unit can also select a fast-tempo ringtone to refresh them. Furthermore, if the user is focused, the selection unit can select a medium-tempo ringtone to help them maintain their concentration. For example, the selection unit can estimate the user's emotions using technologies such as facial recognition or voice analysis and select a ringtone with the optimal tempo based on those emotions. This allows for the provision of more appropriate ringtones by adjusting the tempo of the ringtone based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the selection unit may be performed using AI, for example, or without AI. For example, the selection unit can input the user's emotional data into a generating AI and have the AI ​​select a ringtone with the optimal tempo.

[0069] The selection unit can analyze the user's past selection history and select a ringtone appropriate for a specific season or weather. For example, the selection unit might select warm music in winter, relaxing music on rainy days, and refreshing music in summer. By analyzing the user's past selection history and selecting the optimal ringtone for a specific season or weather, the selection unit can provide a ringtone that suits the user's situation. Some or all of the above processing in the selection unit may be performed using AI, for example, or without AI. For example, the selection unit can input the user's past selection history data into a generating AI and have the generating AI select the optimal ringtone for the season or weather.

[0070] The selection unit can select an appropriate ringtone based on the user's current activity. For example, if the user is exercising, the selection unit may select energetic music. It may also select music to enhance concentration if the user is working. Furthermore, it may select relaxing music if the user is taking a break. For instance, the selection unit can detect the user's current activity and select the optimal ringtone based on that activity. This allows the system to provide a ringtone that suits the user's situation by offering the most suitable ringtone based on the user's current activity. Some or all of the above processing in the selection unit may be performed using AI, for example, or without AI. For example, the selection unit can input the user's current activity data into a generating AI and have the generating AI select the optimal ringtone.

[0071] The selection unit can estimate the user's emotions and change the genre of the ringtone based on the estimated emotions. For example, if the user is relaxed, the selection unit may select classical music. If the user is stressed, the selection unit may select pop music. Furthermore, if the user is excited, the selection unit may select rock music. For example, the selection unit estimates the user's emotions using technologies such as facial recognition or voice analysis, and selects the most appropriate genre of ringtone based on those emotions. This allows for the provision of more appropriate ringtones by changing the genre of ringtone based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the selection unit may be performed using AI, or not using AI. For example, the selection unit can input the user's emotion data into the generative AI and have the generative AI select the most appropriate genre of ringtone.

[0072] The selection unit can select a ringtone appropriate to the culture and customs of the region based on the user's geographical location. For example, if the user is in Japan, the selection unit can select Japanese-style music. It can also select country music if the user is in the United States. Furthermore, if the user is in India, it can select Bollywood music. In short, the selection unit selects the optimal ringtone appropriate to the culture and customs of the region based on the user's geographical location. This allows for the provision of ringtones appropriate to the local culture and customs by selecting ringtones based on geographical location. Some or all of the above processing in the selection unit may be performed using AI, for example, or without AI. For example, the selection unit can input the user's geographical location data into a generating AI and have the generating AI select the optimal ringtone appropriate to the culture and customs of the region.

[0073] The selection unit can analyze the user's social media activity and select a ringtone based on trends. For example, the selection unit can select a ringtone based on music the user has recently shared. It can also select a ringtone based on music preferred by the user's followers. Furthermore, it can select a ringtone based on music related to events the user is participating in. For example, the selection unit analyzes the user's social media activity and selects the optimal ringtone based on trends. This allows the selection of ringtones based on social media activity to provide ringtones that are in line with current trends. Some or all of the above processing in the selection unit may be performed using AI, for example, or not. For example, the selection unit can input the user's social media activity data into a generating AI and have the generating AI select the optimal trend-based ringtone.

[0074] The detection unit can estimate the user's emotions and adjust the detection sensitivity of ambient sounds based on the estimated emotions. For example, if the user is relaxed, the detection unit can set the ambient sound detection sensitivity low. Conversely, if the user is stressed, the detection unit can also set the ambient sound detection sensitivity high. Furthermore, if the user is concentrating, the detection unit can also set the ambient sound detection sensitivity to a medium level. For example, the detection unit can estimate the user's emotions using technologies such as facial recognition or voice analysis, and set the optimal detection sensitivity based on those emotions. This allows for more appropriate detection of ambient sounds by adjusting the detection sensitivity based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input the user's emotional data into a generating AI, which can then execute the setting of the optimal detection sensitivity.

[0075] The detection unit can analyze collected ambient sounds in real time and respond immediately to sudden noises. For example, the detection unit can detect a sudden car horn sound and temporarily increase the ringtone volume. It can also detect a sudden dog bark and change the type of ringtone. Furthermore, it can detect a sudden human scream and temporarily decrease the ringtone volume. For example, the detection unit analyzes collected ambient sounds in real time and responds immediately to sudden noises. This allows the user to receive an appropriate ringtone by responding immediately to sudden noises. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input collected ambient sound data into a generating AI and have the generating AI perform the detection and response to sudden noises.

[0076] The detection unit can classify the collected ambient sounds and perform special processing on specific sounds. For example, the detection unit can detect a car horn and increase the volume of the ringtone. It can also detect a dog barking and change the type of ringtone. Furthermore, it can detect human speech and decrease the volume of the ringtone. For example, the detection unit classifies the collected ambient sounds and performs special processing on specific sounds. This allows the user to receive an appropriate ringtone by performing special processing on specific sounds. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input the collected ambient sound data into a generating AI and have the generating AI perform the classification and processing of specific sounds.

[0077] The detection unit can estimate the user's emotions and adjust the detection frequency of ambient sounds based on the estimated emotions. For example, if the user is relaxed, the detection unit can set the detection frequency of ambient sounds low. Conversely, if the user is stressed, the detection unit can set the detection frequency of ambient sounds high. Furthermore, if the user is concentrating, the detection unit can set the detection frequency of ambient sounds to a medium level. For example, the detection unit estimates the user's emotions using technologies such as facial recognition or voice analysis, and sets the optimal detection frequency based on those emotions. This allows for more appropriate detection of ambient sounds by adjusting the detection frequency based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input the user's emotional data into a generating AI, which can then execute the setting of the optimal detection frequency.

[0078] The detection unit can adjust the collected ambient sound based on the battery level of the user's device. For example, if the battery level is low, the detection unit can set the ambient sound detection frequency to a low level. Conversely, if the battery level is sufficient, the detection unit can also set the ambient sound detection frequency to a high level. Furthermore, if the battery level is moderate, the detection unit can also set the ambient sound detection frequency to a moderate level. For example, the detection unit adjusts the collected ambient sound based on the battery level of the user's device. This allows for efficient use of the device's battery by adjusting the ambient sound detection frequency based on the battery level. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input the collected ambient sound data into a generating AI and have the generating AI perform adjustments based on the battery level.

[0079] The detection unit can combine collected ambient sounds with the user's schedule information to perform detections tailored to specific time periods. For example, the detection unit can set a higher detection frequency for ambient sounds when the user is in a meeting. It can also set a lower detection frequency when the user is on a break. Furthermore, it can set a moderate detection frequency when the user is exercising. For example, the detection unit combines collected ambient sounds with the user's schedule information to perform detections tailored to specific time periods. By adjusting the detection frequency of ambient sounds based on the schedule information, more appropriate detection of ambient sounds becomes possible. Some or all of the above processing in the detection unit may be performed using AI, for example, or without AI. For example, the detection unit can input collected ambient sound data into a generating AI and have the generating AI perform detections based on the schedule information.

[0080] The adjustment unit can estimate the user's emotions and fine-tune the ringtone volume based on the estimated emotions. For example, if the user is relaxed, the adjustment unit can set the ringtone volume low. If the user is stressed, the adjustment unit can also set the ringtone volume high. Furthermore, if the user is focused, the adjustment unit can also set the ringtone volume to a medium level. For example, the adjustment unit can estimate the user's emotions using technologies such as facial recognition or voice analysis and set the optimal volume based on those emotions. This allows for a more appropriate ringtone by fine-tuning the ringtone volume based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the adjustment unit may be performed using AI, or not using AI. For example, the adjustment unit can input the user's emotion data into the generative AI and have the generative AI set the optimal volume.

[0081] The adjustment unit can optimize the sound quality of the selected ringtone based on the speaker characteristics of the user's device. For example, if a high-quality speaker is being used, the adjustment unit will set the sound quality to a high level. It can also set the sound quality to a low level if a low-quality speaker is being used. Furthermore, if a medium-quality speaker is being used, the adjustment unit can set the sound quality to a medium level. For example, the adjustment unit optimizes the sound quality of the selected ringtone based on the speaker characteristics of the user's device. This allows for the provision of better-sounding ringtones by optimizing the sound quality based on the speaker characteristics. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or without AI. For example, the adjustment unit can input the speaker characteristic data of the user's device into a generating AI and have the generating AI perform the optimal sound quality setting.

[0082] The adjustment unit can adjust the duration of the selected ringtone based on the user's current situation. For example, the adjustment unit may set the ringtone duration shorter if the user is in a meeting. It may also set the ringtone duration longer if the user is driving. Furthermore, it may set the ringtone duration to a medium level if the user is taking a break. For example, the adjustment unit adjusts the duration of the selected ringtone based on the user's current situation. This allows for the provision of a more appropriate ringtone by adjusting the ringtone duration based on the user's current situation. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or without AI. For example, the adjustment unit can input the user's current situation data into a generating AI and have the generating AI set the optimal duration.

[0083] The adjustment unit can estimate the user's emotions and adjust the ringtone effects based on the estimated emotions. For example, if the user is relaxed, the adjustment unit may set a strong echo effect. It may also set a strong reverb effect if the user is stressed. Furthermore, if the user is focused, the adjustment unit may set the effect to a moderate level. For example, the adjustment unit estimates the user's emotions using technologies such as facial recognition or voice analysis and sets the optimal effect based on those emotions. This allows for the provision of more appropriate ringtones by adjusting the effects based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the adjustment unit may be performed using AI, or not. For example, the adjustment unit can input the user's emotion data into the generative AI and have the generative AI set the optimal effect.

[0084] The adjustment unit can optimize the volume of the selected ringtone based on the battery level of the user's device. For example, if the battery level is low, the adjustment unit can set the ringtone volume low. Conversely, if the battery level is sufficient, the adjustment unit can also set the ringtone volume high. Furthermore, if the battery level is moderate, the adjustment unit can also set the ringtone volume to a moderate level. For example, the adjustment unit optimizes the volume of the selected ringtone based on the battery level of the user's device. This allows for efficient use of the device's battery by optimizing the volume based on the battery level. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or without AI. For example, the adjustment unit can input the battery level data of the user's device into a generating AI and have the generating AI perform the optimal volume setting.

[0085] The adjustment unit can adjust the selected ringtone type in conjunction with the user's schedule information to optimize it for specific events. For example, the adjustment unit can select a quiet ringtone during a meeting. It can also select an energetic ringtone during exercise. Furthermore, it can select a relaxing ringtone during a break. For example, the adjustment unit adjusts the selected ringtone type in conjunction with the user's schedule information to optimize it for specific events. This allows the adjustment unit to provide a ringtone suitable for specific events by adjusting the ringtone type based on the schedule information. Some or all of the above processing in the adjustment unit may be performed using AI, for example, or without AI. For example, the adjustment unit can input the user's schedule information data into a generating AI and have the generating AI set the optimal ringtone type.

[0086] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0087] The selection function can choose a ringtone based on the user's preferences, circumstances, and health condition. For example, if the user is exercising and has a high heart rate, the selection function will select relaxing music. It can also select an alarm sound if the user is sleeping and has a low heart rate. Furthermore, if the user is feeling stressed, the selection function can select music with stress-reducing effects. This allows the system to provide the optimal ringtone based on the user's health condition.

[0088] The selection unit can choose a ringtone based on the user's preferences, circumstances, and activity history. For example, if the user previously selected a specific song at a particular location, the unit will select the same song when the user revisits that location. Similarly, if the user selected a specific song during a particular time period, the unit can select the same song again during that time period. Furthermore, if the user selected a specific song at a particular event, the unit can select the same song again when the user attends that event again. This allows the system to provide the most suitable ringtone based on the user's activity history.

[0089] The detection unit can detect the temperature and humidity around the user and adjust the ringtone accordingly. For example, if the ambient temperature is high, the detection unit can select cool music. It can also select refreshing music if the ambient humidity is high. Furthermore, if the ambient temperature is low, it can select warm music. This allows the system to provide the optimal ringtone based on ambient temperature and humidity.

[0090] The selection function can choose a ringtone based on the user's preferences and circumstances, as well as the battery level of the user's device. For example, if the battery level is low, the selection function will select a short ringtone. If the battery level is sufficient, it can also select a long ringtone. Furthermore, if the battery level is moderate, it can select a ringtone of medium length. This allows the system to provide the optimal ringtone based on the device's battery level.

[0091] The selection function can choose a ringtone based on the user's preferences, circumstances, and schedule information. For example, if the user is in a meeting, the selection function will select a quiet ringtone. It can also select an energetic ringtone if the user is exercising. Furthermore, it can select a relaxing ringtone if the user is taking a break. This allows the system to provide the optimal ringtone based on the user's schedule information.

[0092] The selection unit can estimate the user's emotions and adjust the ringtone tempo based on those emotions. For example, if the user is relaxed, it can select a slow-tempo ringtone. If the user is stressed, it can select a fast-tempo ringtone to help them refresh. Furthermore, if the user is concentrating, it can select a medium-tempo ringtone to help them maintain their focus. By adjusting the ringtone tempo based on the user's emotions, it can provide a more appropriate ringtone.

[0093] The selection function can estimate the user's emotions and change the ringtone genre based on those emotions. For example, if the user is relaxed, it can select classical music. If the user is stressed, it can select pop music. Furthermore, if the user is excited, it can select rock music. By changing the ringtone genre based on the user's emotions, it can provide a more appropriate ringtone.

[0094] The detection unit can estimate the user's emotions and adjust the sensitivity of ambient sound detection based on the estimated emotions. For example, if the user is relaxed, the sensitivity of ambient sound detection can be set low. Conversely, if the user is stressed, the sensitivity of ambient sound detection can be set high. Furthermore, if the user is concentrating, the sensitivity of ambient sound detection can be set to a medium level. By adjusting the sensitivity of ambient sound detection based on the user's emotions, more appropriate detection of ambient sounds becomes possible.

[0095] The adjustment unit can estimate the user's emotions and fine-tune the ringtone volume based on those emotions. For example, if the user is relaxed, the ringtone volume can be set low. If the user is stressed, the ringtone volume can be set high. Furthermore, if the user is concentrating, the ringtone volume can be set to a medium level. This allows for fine-tuning the ringtone volume based on the user's emotions, providing a more appropriate ringtone.

[0096] The adjustment unit can estimate the user's emotions and adjust the ringtone effects based on those emotions. For example, if the user is relaxed, the echo effect can be set to a stronger level. If the user is stressed, the reverb effect can also be set to a stronger level. Furthermore, if the user is focused, the effect can be set to a moderate level. This allows for the provision of more appropriate ringtones by adjusting the effects based on the user's emotions.

[0097] The following briefly describes the processing flow for example form 2.

[0098] Step 1: The selection unit chooses a ringtone based on the user's preferences or current situation. For example, if the user is in a quiet place, it will choose quiet music; if they are in a noisy place, it will choose louder music. The selection unit can also learn from the user's feedback and select the optimal ringtone. For example, it will analyze the user's past ringtone selection history and select the same ringtone in similar situations. Step 2: The detection unit detects ambient noise around the user. For example, it uses a microphone to collect ambient noise and analyzes the collected noise to understand the situation the user is in. It detects that there is little ambient noise in quiet places and a lot of ambient noise in noisy places. Step 3: The adjustment unit adjusts the volume or type of the selected ringtone based on the ambient sound detected by the detection unit. For example, it lowers the volume in quiet places and increases it in noisy places. It can also prioritize the selection of a specific genre of music. Furthermore, it can select a ringtone suitable for a specific time of day. For example, it can select quiet music at night and lively music during the day.

[0099] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0100] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0101] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0102] For example, each of the multiple elements, including the selection unit, detection unit, and adjustment unit, is implemented in at least one of the smart device 14 and the data processing device 12. For example, the selection unit is implemented by the control unit 46A of the smart device 14 and selects a ringtone based on the user's preference and circumstances. The detection unit collects ambient sound using the microphone 38B of the smart device 14 and analyzes it with the control unit 46A. The adjustment unit is implemented by the control unit 46A of the smart device 14 and adjusts the volume and type of the ringtone based on the detected ambient sound. The selection unit, detection unit, and adjustment unit may also be implemented by the specific processing unit 290 of the data processing device 12. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0103] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0104] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0105] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0106] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0107] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0108] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0109] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0110] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0111] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0112] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0113] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0114] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0115] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0116] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0117] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0118] For example, each of the multiple elements, including the selection unit, detection unit, and adjustment unit, is implemented in at least one of the smart glasses 214 and the data processing device 12. For example, the selection unit is implemented by the control unit 46A of the smart glasses 214 and selects a ringtone based on the user's preference and circumstances. The detection unit collects ambient sounds using the microphone 238 of the smart glasses 214 and analyzes them with the control unit 46A. The adjustment unit is implemented by the control unit 46A of the smart glasses 214 and adjusts the volume and type of the ringtone based on the detected ambient sounds. The selection unit, detection unit, and adjustment unit may also be implemented by the specific processing unit 290 of the data processing device 12. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0119] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0120] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0121] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0122] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0123] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0124] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0125] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0126] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0127] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0128] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0129] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0130] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0131] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0132] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0133] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0134] For example, each of the multiple elements, including the selection unit, detection unit, and adjustment unit, is implemented in at least one of the headset terminal 314 and the data processing device 12. For example, the selection unit is implemented by the control unit 46A of the headset terminal 314 and selects a ringtone based on the user's preference and circumstances. The detection unit collects ambient sounds using the microphone 238 of the headset terminal 314 and analyzes them with the control unit 46A. The adjustment unit is implemented by the control unit 46A of the headset terminal 314 and adjusts the volume and type of the ringtone based on the detected ambient sounds. The selection unit, detection unit, and adjustment unit may also be implemented by the specific processing unit 290 of the data processing device 12. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0135] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0136] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0137] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0138] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0139] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0140] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0141] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0142] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0143] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0144] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0145] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0146] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0147] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0148] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0149] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0150] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0151] For example, each of the multiple elements, including the selection unit, detection unit, and adjustment unit, is implemented by at least one of the robot 414 and the data processing unit 12. For example, the selection unit is implemented by the control unit 46A of the robot 414 and selects a ringtone based on the user's preference and circumstances. The detection unit collects ambient sounds using the microphone 238 of the robot 414 and analyzes them by the control unit 46A. The adjustment unit is implemented by the control unit 46A of the robot 414 and adjusts the volume and type of the ringtone based on the detected ambient sounds. The selection unit, detection unit, and adjustment unit may also be implemented by the specific processing unit 290 of the data processing unit 12. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0152] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0153] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0154] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0155] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0156] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0157] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0158] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0159] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0160] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0161] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0162] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0163] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0164] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0165] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0166] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0167] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0168] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0169] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0170] (Note 1) A selection unit that selects a ringtone based on the user's preference or the user's current situation, The aforementioned detection unit detects ambient sounds around the user, The system includes an adjustment unit that adjusts the volume or type of selected ringtone based on ambient sound detected by the aforementioned detection unit. A system characterized by the following features. (Note 2) The aforementioned selection unit is Learns the ringtones to choose based on user feedback. The system described in Appendix 1, characterized by the features described herein. (Note 3) The detection unit, Use a microphone to collect ambient sounds. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned selection unit is Prioritize selecting a specific music genre. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned selection unit is Select a ringtone suitable for a specific time of day (e.g., morning, noon, evening). The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned selection unit is It estimates the user's emotions and selects a ringtone based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned selection unit is It analyzes the user's past selection history and selects ringtones appropriate for specific events or situations. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned selection unit is It estimates the user's emotions and adjusts the ringtone tempo based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned selection unit is It analyzes the user's past selection history and selects ringtones appropriate for specific seasons and weather conditions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned selection unit is Selects an appropriate ringtone based on the user's current activity. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned selection unit is It estimates the user's emotions and changes the ringtone genre based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned selection unit is Based on the user's geographical location, the system selects a ringtone appropriate to the culture and customs of that region. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned selection unit is Analyzes the user's social media activity and selects ringtones based on trends. The system described in Appendix 1, characterized by the features described herein. (Note 14) The detection unit, The system estimates the user's emotions and adjusts the sensitivity of ambient sound detection based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 15) The detection unit, The system analyzes collected ambient sounds in real time and responds immediately to sudden noises. The system described in Appendix 1, characterized by the features described herein. (Note 16) The detection unit, The collected ambient sounds are classified, and special processing is applied to specific sounds. The system described in Appendix 1, characterized by the features described herein. (Note 17) The detection unit, The system estimates the user's emotions and adjusts the frequency of ambient sound detection based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 18) The detection unit, The collected ambient sounds are adjusted based on the battery level of the user's device. The system described in Appendix 1, characterized by the features described herein. (Note 19) The detection unit, The collected ambient sounds are linked with the user's schedule information to perform detection specifically tailored to certain time periods. The system described in Appendix 1, characterized by the features described herein. (Note 20) The adjustment unit is, It estimates the user's emotions and fine-tunes the ringtone volume based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 21) The adjustment unit is, The sound quality of the selected ringtone is optimized based on the speaker characteristics of the user's device. The system described in Appendix 1, characterized by the features described herein. (Note 22) The adjustment unit is, Adjust the duration of the selected ringtone based on the user's current situation. The system described in Appendix 1, characterized by the features described herein. (Note 23) The adjustment unit is, It estimates the user's emotions and adjusts the ringtone effects based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 24) The adjustment unit is, The volume of the selected ringtone is optimized based on the battery level of the user's device. The system described in Appendix 1, characterized by the features described herein. (Note 25) The adjustment unit is, The selected ringtone type is linked to the user's schedule information and customized for specific events. The system described in Appendix 1, characterized by the features described herein. [Explanation of symbols]

[0171] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A selection unit that selects a ringtone based on the user's preference or the user's current situation, The aforementioned detection unit detects ambient sounds around the user, The system includes an adjustment unit that adjusts the volume or type of selected ringtone based on ambient sound detected by the aforementioned detection unit. A system characterized by the following features.

2. The aforementioned selection unit is Based on the user's feedback, the system learns the ringtones to be selected. The system according to feature 1.

3. The detection unit, Use a microphone to collect ambient sounds. The system according to feature 1.

4. The aforementioned selection unit is Prioritize selecting a specific music genre. The system according to feature 1.

5. The aforementioned selection unit is Select a ringtone suitable for a specific time of day. The system according to feature 1.

6. The aforementioned selection unit is The system estimates the user's emotions and selects a ringtone based on those estimated emotions. The system according to feature 1.

7. The aforementioned selection unit is The system analyzes the user's past selection history and selects a ringtone appropriate for a specific event or situation. The system according to feature 1.

8. The aforementioned selection unit is The system estimates the user's emotions and adjusts the tempo of the ringtone based on the estimated emotions of the user. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A