system

The system efficiently identifies and provides wild bird call information using generative AI, enhancing birdwatching experiences and promoting nature conservation.

JP2026073607APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing systems face difficulties in efficiently identifying the calls of wild birds and providing this information to users.

Method used

A system comprising a collection unit, analysis unit, and display unit that uses generative AI to analyze ambient sounds, identify specific wild bird calls, and provide information to users through a smartphone app, including direction estimation on a map.

Benefits of technology

Enables efficient identification and provision of wild bird call information, making birdwatching easier for users and raising awareness of nature conservation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073607000001_ABST
    Figure 2026073607000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to efficiently identify and provide bird calls to the user. [Solution] The system according to the embodiment comprises a collection unit, an analysis unit, an identification unit, and a display unit. The collection unit collects ambient sounds. The analysis unit analyzes the sounds collected by the collection unit. The identification unit identifies the calls of specific wild birds from the sounds analyzed by the analysis unit. The display unit provides the user with the information identified by the identification unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document No. 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, there is a problem that it is difficult to efficiently identify the calls of wild birds and provide them to users.

[0005] The system according to the embodiment aims to efficiently identify the calls of wild birds and provide them to users.

Means for Solving the Problems

[0006] The system according to the embodiment includes a collection unit, an analysis unit, an identification unit, and a display unit. The collection unit collects ambient sound. The analysis unit analyzes the sound collected by the collection unit. The identification unit identifies the calls of specific wild birds from the sound analyzed by the analysis unit. The display unit provides the information identified by the identification unit to the user.

Advantages of the Invention

[0007] The system according to this embodiment can efficiently identify and provide information about wild bird calls to the user. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the reception device 38, the output device 40, and the camera 42 are connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The Bird Research System according to an embodiment of the present invention is a system for supporting birdwatching. The Bird Research System uses earphones, a microphone, and a smartphone as devices. For beginners in particular, distinguishing between the calls of multiple wild birds in nature and finding a specific bird is difficult, and acquiring the necessary skills takes time and experience. The Bird Research System uses generative AI to identify only the calls of the target wild bird from among the noisy calls of wild birds, indicates its direction, and provides functions such as listing the types and numbers of wild birds in the surrounding area. First, the user connects the earphones and microphone to their smartphone and launches the Bird Research app. Next, the generative AI analyzes the surrounding sounds in real time and identifies the calls of wild birds. For example, if the user registers the calls of a specific wild bird in advance, the app will notify the user when that bird calls. The generative AI also accurately identifies the calls of multiple wild birds and identifies their types and numbers. Furthermore, the Bird Research System estimates the direction of the calls through acoustic analysis and displays it on a map, allowing for a more accurate understanding of the location of wild birds. For example, if a user is pointing their smartphone north, the direction of the bird call will also be displayed as north. This allows the user to receive navigation tailored to the direction they are facing. This system can broaden the base of birdwatching, from beginners to intermediate users. Users can intuitively identify and observe wild birds' calls without complex operations. Furthermore, with the support of the generating AI, anyone can easily enjoy birdwatching, and it is expected that this will raise awareness of nature conservation. In this way, the Triresearch system makes it easier for users to enjoy birdwatching and raises awareness of nature conservation.

[0029] The research system according to this embodiment comprises a collection unit, an analysis unit, an identification unit, and a display unit. The collection unit collects ambient sounds. The collection unit can collect ambient sounds using, for example, a microphone array connected to a smartphone or earphones. The collection unit can collect, for example, natural sounds, artificial sounds, sounds in a specific frequency band, etc. The analysis unit analyzes the sounds collected by the collection unit. The analysis unit can analyze the sounds using, for example, methods such as speech recognition algorithms, frequency analysis, and time-domain analysis. The analysis unit can analyze the collected sounds using generative AI. The identification unit identifies the calls of specific wild birds from the sounds analyzed by the analysis unit. The identification unit can identify the calls of specific wild birds using, for example, methods such as machine learning algorithms and pattern matching. The identification unit can identify the calls of specific wild birds using generative AI. The display unit provides the user with the information identified by the identification unit. The display unit can provide information using, for example, text display, voice notification, and visual display. The display unit can list the types and numbers of identified wild birds. The display unit can also estimate the direction of bird calls through acoustic analysis and display it on a map. As a result, the birdwatching research system according to this embodiment can make birdwatching easier for users and raise awareness of nature conservation.

[0030] The sound collection unit collects ambient sounds. For example, the collection unit can collect ambient sounds using a microphone array connected to a smartphone or earphones. Specifically, it uses a microphone built into the smartphone or an externally connected high-sensitivity microphone array to collect ambient sounds with high accuracy. A microphone array can arrange multiple microphones to estimate the direction and distance of a sound source. This allows the collection unit to emphasize sounds from a specific direction and reduce noise. The collection unit can collect, for example, natural sounds, artificial sounds, and sounds in specific frequency bands. Natural sounds include wind, rain, babbling brooks, and bird songs, while artificial sounds include car engine noises, human speech, and machine operation sounds. To collect sounds in a specific frequency band, the collection unit can use filtering technology to extract only the desired sounds. For example, when collecting bird songs, a filter is set to a specific frequency band, emphasizing and collecting sounds in that band. This allows the collection unit to collect ambient environmental sounds with high accuracy and provide them to the analysis unit. Furthermore, the data collection unit can process the collected sound data in real time and transmit it to the analysis unit. This allows the data collection unit to collect sound data efficiently and effectively, improving the overall performance of the system.

[0031] The analysis unit analyzes the sound collected by the collection unit. The analysis unit can analyze the sound using methods such as speech recognition algorithms, frequency analysis, and time-domain analysis. Specifically, it uses speech recognition algorithms to identify speech and specific sounds from the collected sound data. Frequency analysis analyzes the frequency components of sound and extracts sounds characteristic of specific frequency bands. Time-domain analysis analyzes the temporal changes in sound and identifies sound patterns and rhythms. The analysis unit can also analyze the collected sound using generative AI. Generative AI uses deep learning models to analyze sound data and extract sound features with high accuracy. For example, the generative AI takes collected sound data as input, generates a sound spectrogram, and analyzes that spectrogram to extract sound features. The generative AI learns from past sound data and has the ability to identify specific sound patterns and features. This allows the analysis unit to quickly and accurately analyze the collected sound data and provide it to the identification unit. Furthermore, the analysis unit can analyze sound data in real time and instantly grasp the surrounding situation. This allows the analysis unit to analyze sound data efficiently and effectively, improving the overall performance of the system.

[0032] The identification unit identifies the calls of specific wild birds from the sounds analyzed by the analysis unit. The identification unit can identify the calls of specific wild birds using methods such as machine learning algorithms and pattern matching. Specifically, it uses machine learning algorithms to extract the characteristics of wild bird calls from the analyzed sound data and identifies the calls of specific wild birds based on those characteristics. In pattern matching, the analyzed sound data is compared with a database of known wild bird calls and matching patterns are identified. The identification unit can also identify the calls of specific wild birds using generative AI. The generative AI analyzes sound data using a deep learning model and identifies the characteristics of wild bird calls with high accuracy. For example, the generative AI takes collected sound data as input, generates a spectrogram of wild bird calls, and analyzes the spectrogram to identify the wild bird calls. The generative AI learns from past wild bird call data and has the ability to identify the calls of specific wild birds. This allows the identification unit to quickly and accurately identify the calls of specific wild birds from the collected sound data and provide this information to the display unit. Furthermore, the identification unit can identify sound data in real time and instantly grasp the surrounding situation. As a result, the identification unit can identify sound data efficiently and effectively, improving the overall performance of the system.

[0033] The display unit provides the user with information identified by the identification unit. The display unit can provide information using methods such as text display, voice notification, and visual display. Specifically, it can list the identified species and number of wild birds and display them to the user in text format. Voice notification can inform the user of the identified species and number of wild birds audibly. Visual display graphically displays the identified species and number of wild birds, providing the user with information visually. The display unit can estimate the direction of bird calls through acoustic analysis and display it on a map. Specifically, it estimates the direction and distance of the sound source based on sound data collected by the collection unit and displays this information on a map. This allows the user to visually understand the direction from which the bird calls are coming. Furthermore, the display unit can collect user feedback and continuously improve the accuracy and effectiveness of the displayed content. For example, the user can provide feedback on the identified bird information, and the displayed content can be reviewed based on that feedback. The display unit can also reliably transmit information using multiple communication methods. For example, it uses not only smartphone notifications but also voice calls, SMS, and email to reliably deliver important information. This allows the display to provide information to the user quickly and reliably, making birdwatching easier and more enjoyable.

[0034] The display unit can list the types and numbers of identified wild birds. For example, the display unit can list the types and numbers of identified wild birds in text format. The display unit can also display the types and numbers of identified wild birds graphically. The display unit can also announce the types and numbers of identified wild birds via voice. This makes it easier for users to understand information about the wild birds in their surroundings by listing the types and numbers of identified birds.

[0035] The display unit can estimate the direction of a bird's call through acoustic analysis and display it on a map. For example, the display unit can estimate the direction of a bird's call through acoustic analysis and display it on a map with an arrow. The display unit can also estimate the direction of a bird's call through acoustic analysis and display it on a map with an icon. The display unit can also estimate the direction of a bird's call through acoustic analysis and display it on a map using different colors. This allows users to more accurately determine the location of wild birds by displaying the direction of the bird's call on a map.

[0036] The identification unit may include a notification unit that pre-registers the calls of specific wild birds and notifies the user when those birds make a sound. For example, the identification unit can pre-register the calls of specific wild birds and provide an audio notification when those birds make a sound. The identification unit can also pre-register the calls of specific wild birds and provide a text notification when those birds make a sound. The identification unit can also pre-register the calls of specific wild birds and provide a visual notification when those birds make a sound. This makes it easier for users to find specific wild birds by pre-registering their calls and providing notifications when those birds make a sound.

[0037] The sound collection unit can collect ambient sounds using a microphone array connected to a smartphone or earphones. For example, the sound collection unit can collect ambient sounds using a microphone array connected to a smartphone. The sound collection unit can also collect ambient sounds using a microphone array connected to earphones. The sound collection unit can also collect ambient sounds using multiple microphones. This allows for efficient collection of ambient sounds by using a microphone array connected to a smartphone or earphones.

[0038] The sound collection unit can analyze ambient sounds and prioritize the collection of sounds in specific frequency bands. For example, it can analyze ambient sounds in real time and prioritize the collection of frequency bands that contain bird calls. The sound collection unit can also remove human voices and car noises from ambient sounds and collect only bird calls. If the ambient noise level is high, the sound collection unit can also emphasize and collect sounds in specific frequency bands. This allows for the efficient collection of bird calls by prioritizing the collection of sounds in specific frequency bands.

[0039] The sound collection unit can automatically adjust the optimal microphone placement based on the user's location information when collecting sound. For example, the unit can automatically adjust the optimal microphone placement based on the user's location information. The unit can also update the microphone placement in real time as the user moves. The unit can also select the optimal sound collection point by considering the user's location information and the surrounding terrain information. This allows for efficient sound collection by automatically adjusting the optimal microphone placement based on the user's location information.

[0040] The sound collection unit can select the target for collection by referring to the user's past birdwatching history. For example, the collection unit can prioritize collecting the calls of wild birds that the user has observed in the past. The collection unit can also collect the calls of specific wild birds from the user's past birdwatching history. The collection unit can also collect the calls of wild birds observed in places the user has visited in the past. This allows for the collection of more appropriate sounds by referring to the user's past birdwatching history.

[0041] The sound collection unit can adjust its collection method by considering the surrounding weather information. For example, in rainy weather, the unit can remove the sound of rain and collect bird songs. In sunny weather, the unit can remove the sound of wind and collect bird songs. On snowy days, the unit can remove the sound of snow and collect bird songs. By adjusting the collection method by considering the surrounding weather information, more appropriate sounds can be collected.

[0042] The analysis unit can be equipped with a filtering function that automatically removes noise from the collected sound during analysis. For example, the analysis unit can be equipped with a filtering function that automatically removes wind noise from the collected sound. The analysis unit can also be equipped with a filtering function that automatically removes human voices from the collected sound. The analysis unit can also be equipped with a filtering function that automatically removes car noise from the collected sound. This improves the accuracy of the analysis by automatically removing noise from the collected sound.

[0043] The analysis unit can simultaneously analyze the calls of different wild birds and identify multiple calls. For example, the analysis unit can simultaneously analyze the calls of different wild birds and identify them by species. The analysis unit can also simultaneously analyze the calls of multiple wild birds and identify the intensity of the calls. The analysis unit can also simultaneously analyze the calls of different wild birds and identify the direction of the calls. In this way, by simultaneously analyzing the calls of different wild birds, multiple calls can be identified.

[0044] The analysis unit can improve analysis accuracy by considering the time-of-day information of the collected sounds during analysis. For example, the analysis unit can improve analysis accuracy by considering the activity times of wild birds based on the time-of-day information of the collected sounds. The analysis unit can also prioritize the analysis of wild bird calls that occur during specific time periods based on the time-of-day information of the collected sounds. The analysis unit can also analyze nocturnal wild bird calls based on the time-of-day information of the collected sounds. In this way, the analysis accuracy is improved by considering the time-of-day information of the collected sounds.

[0045] The analysis unit can supplement the analysis results by referring to the geographical information of the collected sounds during the analysis. For example, the analysis unit can prioritize the analysis of bird calls inhabiting a specific area based on the geographical information of the collected sounds. The analysis unit can also identify bird calls for each region based on the geographical information of the collected sounds. The analysis unit can also analyze bird calls for a specific region based on the geographical information of the collected sounds. In this way, the analysis results can be supplemented by referring to the geographical information of the collected sounds.

[0046] The identification unit can improve its identification accuracy by learning the call patterns of specific wild birds during identification. For example, the identification unit can improve its identification accuracy by learning the call patterns of specific wild birds. The identification unit can also improve its identification accuracy by learning the call patterns of multiple wild birds. The identification unit can improve its identification accuracy by learning the call patterns of wild birds. As a result, identification accuracy is improved by learning the call patterns of specific wild birds.

[0047] The identification unit can simultaneously identify the calls of multiple wild birds and classify them by species. For example, the identification unit can simultaneously identify the calls of multiple wild birds and classify them by species. The identification unit can also simultaneously identify the calls of multiple wild birds and classify their intensity. Furthermore, the identification unit can simultaneously identify the calls of multiple wild birds and classify their direction. This allows for the simultaneous identification of the calls of multiple wild birds, enabling classification by species.

[0048] The identification unit can improve its identification accuracy by considering the frequency characteristics of the collected sound during identification. For example, the identification unit can identify the calls of wild birds based on the frequency characteristics of the collected sound. The identification unit can also identify the calls of specific wild birds based on the frequency characteristics of the collected sound. The identification unit can also identify the calls of multiple wild birds based on the frequency characteristics of the collected sound. In this way, the identification accuracy is improved by considering the frequency characteristics of the collected sound.

[0049] The identification unit can supplement the identification result by referring to the time-of-day information of the collected sounds during the identification process. For example, the identification unit can identify the calls of wild birds that sing during a specific time period based on the collected time-of-day information of the sounds. The identification unit can also identify the calls of wild birds that sing at night based on the collected time-of-day information of the sounds. The identification unit can also supplement the identification result by considering the activity times of wild birds based on the collected time-of-day information of the sounds. In this way, the identification result can be supplemented by referring to the time-of-day information of the sounds.

[0050] The display unit can update the direction of the identified bird's call in real time while displaying information. For example, if the user moves their smartphone, the display unit can update the direction of the identified bird's call in real time. The display unit can also update the displayed content in real time if the direction of the bird's call changes. The display unit can also update the direction of the identified bird's call in real time based on the user's location information. This allows the user to accurately determine the bird's location by updating the direction of the identified bird's call in real time.

[0051] The display unit can visually show the intensity of the identified bird's call when it is displayed. For example, the display unit can visually show the intensity of the identified bird's call using shades of color. The display unit can also visually show the intensity of the identified bird's call using a bar graph. The display unit can also visually show the intensity of the identified bird's call using numerical values. This allows users to intuitively understand the intensity of the bird's call by visually displaying the intensity of the identified bird's call.

[0052] The display unit can select the optimal display method while considering the user's location information. For example, the display unit can automatically select the optimal display method based on the user's location information. The display unit can also update the display method in real time as the user moves. The display unit can also select the optimal display method considering the user's location information and surrounding terrain information. This allows for the provision of more appropriate information by selecting the optimal display method while considering the user's location information.

[0053] The display unit can show a history of the calls of identified wild birds when it is displayed. For example, the display unit can display a history of identified wild bird calls in chronological order. The display unit can also display a history of identified wild bird calls on a map. The display unit can also display a history of identified wild bird calls in list format. This allows users to review past observation results by displaying a history of identified wild bird calls.

[0054] The notification unit can apply different notification methods depending on the type of bird call identified. For example, if the call of a specific bird is identified, the notification unit can provide an audio notification. If the calls of multiple birds are identified, the notification unit can also provide a vibration notification. If the call of a rare bird is identified, the notification unit can also provide a pop-up notification. This allows the system to provide users with appropriate notifications by applying different notification methods depending on the type of bird call identified.

[0055] The notification unit can select the most appropriate notification method by referring to the user's past notification history when sending a notification. For example, the notification unit can select the most appropriate notification method based on the user's past notification history. The notification unit can also prioritize the application of notification methods that the user has previously preferred. The notification unit can also select a notification method suitable for a specific time period from the user's past notification history. This allows for the selection of a more appropriate notification method by referring to the user's past notification history.

[0056] The notification unit can select the optimal notification method by considering the user's location information when sending a notification. For example, the notification unit can automatically select the optimal notification method based on the user's location information. The notification unit can also update the notification method in real time when the user moves. The notification unit can also select the optimal notification method by considering the user's location information and the surrounding terrain information. This allows for more appropriate notifications by selecting the optimal notification method while considering the user's location information.

[0057] The notification unit can supplement notification content by referring to the history of the identified bird's calls when sending a notification. For example, the notification unit can supplement notification content based on the history of the identified bird's calls. The notification unit can also notify information about a specific bird based on the history of the identified bird's calls. The notification unit can also customize notification content based on the history of the identified bird's calls. This allows for the provision of more appropriate notification content by referring to the history of the identified bird's calls.

[0058] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0059] The Triresearch system can select the optimal sound collection point based on the user's location information in its collection unit. For example, if the user is in a mountainous area, the collection unit can prioritize collecting the calls of wild birds specific to mountainous regions. If the user is by a lake, the collection unit can also prioritize collecting the calls of wild birds that inhabit waterside areas. If the user is in an urban area, the collection unit can perform noise filtering suitable for urban environments and efficiently collect wild bird calls. In this way, by selecting the optimal sound collection point based on the user's location information, sound can be collected efficiently.

[0060] The Trisearch system can improve analysis accuracy by considering the time-of-day information of collected sounds in its analysis unit. For example, it can improve analysis accuracy by considering the activity times of wild birds based on the time-of-day information of collected sounds. It can also prioritize the analysis of wild bird calls that occur during specific time periods. It can also analyze wild bird calls that occur at night. In this way, the analysis accuracy is improved by considering the time-of-day information of collected sounds.

[0061] The Trisearch system analyzes ambient noise in its collection unit and can prioritize the collection of sounds in specific frequency bands. For example, it can analyze ambient noise in real time and prioritize the collection of frequency bands containing bird calls. It can also remove human voices and car noises from the ambient noise and collect only bird calls. If the ambient noise level is high, it can also emphasize and collect sounds in specific frequency bands. This allows for the efficient collection of bird calls by prioritizing the collection of sounds in specific frequency bands.

[0062] The Trisearch system can have a filtering function added to its analysis unit that automatically removes noise from the collected sound. For example, a filtering function can be added to automatically remove wind noise from the collected sound. A filtering function can also be added to automatically remove human voices from the collected sound. A filtering function can also be added to automatically remove car noise from the collected sound. This improves the accuracy of the analysis by automatically removing noise from the collected sound.

[0063] The Trisearch system, in its sound collection unit, can automatically adjust the optimal microphone placement based on the user's location information during sound collection. For example, it can automatically adjust the optimal microphone placement based on the user's location. It can also update the microphone placement in real time as the user moves. It can also select the optimal sound collection point by considering the user's location information and the surrounding terrain information. As a result, by automatically adjusting the optimal microphone placement based on the user's location information, sound can be collected efficiently.

[0064] The birdwatching system can display a history of identified bird calls on its display unit. For example, it can display a history of identified bird calls in chronological order. It can also display a history of identified bird calls on a map. It can also display a history of identified bird calls in a list format. This allows users to review past observation results by viewing a history of identified bird calls.

[0065] The following briefly describes the processing flow for example form 1.

[0066] Step 1: The collection unit collects ambient sounds. The collection unit can collect ambient sounds using, for example, a microphone array connected to a smartphone or earphones. The collection unit can collect natural sounds, artificial sounds, sounds in specific frequency bands, etc. Step 2: The analysis unit analyzes the sound collected by the collection unit. The analysis unit can analyze the sound using methods such as speech recognition algorithms, frequency analysis, and time-domain analysis. The analysis unit can also analyze the collected sound using generative AI. Step 3: The identification unit identifies the calls of specific wild birds from the sounds analyzed by the analysis unit. The identification unit can identify the calls of specific wild birds using methods such as machine learning algorithms and pattern matching. The identification unit can also identify the calls of specific wild birds using generative AI. Step 4: The display unit provides the user with the information identified by the identification unit. The display unit can provide information using methods such as text display, audio notification, and visual display. The display unit can list the types and numbers of identified wild birds. The display unit can estimate the direction of the bird calls by acoustic analysis and display it on a map.

[0067] (Example of form 2) The Bird Research System according to an embodiment of the present invention is a system for supporting birdwatching. The Bird Research System uses earphones, a microphone, and a smartphone as devices. For beginners in particular, distinguishing between the calls of multiple wild birds in nature and finding a specific bird is difficult, and acquiring the necessary skills takes time and experience. The Bird Research System uses generative AI to identify only the calls of the target wild bird from among the noisy calls of wild birds, indicates its direction, and provides functions such as listing the types and numbers of wild birds in the surrounding area. First, the user connects the earphones and microphone to their smartphone and launches the Bird Research app. Next, the generative AI analyzes the surrounding sounds in real time and identifies the calls of wild birds. For example, if the user registers the calls of a specific wild bird in advance, the app will notify the user when that bird calls. The generative AI also accurately identifies the calls of multiple wild birds and identifies their types and numbers. Furthermore, the Bird Research System estimates the direction of the calls through acoustic analysis and displays it on a map, allowing for a more accurate understanding of the location of wild birds. For example, if a user is pointing their smartphone north, the direction of the bird call will also be displayed as north. This allows the user to receive navigation tailored to the direction they are facing. This system can broaden the base of birdwatching, from beginners to intermediate users. Users can intuitively identify and observe wild birds' calls without complex operations. Furthermore, with the support of the generating AI, anyone can easily enjoy birdwatching, and it is expected that this will raise awareness of nature conservation. In this way, the Triresearch system makes it easier for users to enjoy birdwatching and raises awareness of nature conservation.

[0068] The research system according to this embodiment comprises a collection unit, an analysis unit, an identification unit, and a display unit. The collection unit collects ambient sounds. The collection unit can collect ambient sounds using, for example, a microphone array connected to a smartphone or earphones. The collection unit can collect, for example, natural sounds, artificial sounds, sounds in a specific frequency band, etc. The analysis unit analyzes the sounds collected by the collection unit. The analysis unit can analyze the sounds using, for example, methods such as speech recognition algorithms, frequency analysis, and time-domain analysis. The analysis unit can analyze the collected sounds using generative AI. The identification unit identifies the calls of specific wild birds from the sounds analyzed by the analysis unit. The identification unit can identify the calls of specific wild birds using, for example, methods such as machine learning algorithms and pattern matching. The identification unit can identify the calls of specific wild birds using generative AI. The display unit provides the user with the information identified by the identification unit. The display unit can provide information using, for example, text display, voice notification, and visual display. The display unit can list the types and numbers of identified wild birds. The display unit can also estimate the direction of bird calls through acoustic analysis and display it on a map. As a result, the birdwatching research system according to this embodiment can make birdwatching easier for users and raise awareness of nature conservation.

[0069] The sound collection unit collects ambient sounds. For example, the collection unit can collect ambient sounds using a microphone array connected to a smartphone or earphones. Specifically, it uses a microphone built into the smartphone or an externally connected high-sensitivity microphone array to collect ambient sounds with high accuracy. A microphone array can arrange multiple microphones to estimate the direction and distance of a sound source. This allows the collection unit to emphasize sounds from a specific direction and reduce noise. The collection unit can collect, for example, natural sounds, artificial sounds, and sounds in specific frequency bands. Natural sounds include wind, rain, babbling brooks, and bird songs, while artificial sounds include car engine noises, human speech, and machine operation sounds. To collect sounds in a specific frequency band, the collection unit can use filtering technology to extract only the desired sounds. For example, when collecting bird songs, a filter is set to a specific frequency band, emphasizing and collecting sounds in that band. This allows the collection unit to collect ambient environmental sounds with high accuracy and provide them to the analysis unit. Furthermore, the data collection unit can process the collected sound data in real time and transmit it to the analysis unit. This allows the data collection unit to collect sound data efficiently and effectively, improving the overall performance of the system.

[0070] The analysis unit analyzes the sound collected by the collection unit. The analysis unit can analyze the sound using methods such as speech recognition algorithms, frequency analysis, and time-domain analysis. Specifically, it uses speech recognition algorithms to identify speech and specific sounds from the collected sound data. Frequency analysis analyzes the frequency components of sound and extracts sounds characteristic of specific frequency bands. Time-domain analysis analyzes the temporal changes in sound and identifies sound patterns and rhythms. The analysis unit can also analyze the collected sound using generative AI. Generative AI uses deep learning models to analyze sound data and extract sound features with high accuracy. For example, the generative AI takes collected sound data as input, generates a sound spectrogram, and analyzes that spectrogram to extract sound features. The generative AI learns from past sound data and has the ability to identify specific sound patterns and features. This allows the analysis unit to quickly and accurately analyze the collected sound data and provide it to the identification unit. Furthermore, the analysis unit can analyze sound data in real time and instantly grasp the surrounding situation. This allows the analysis unit to analyze sound data efficiently and effectively, improving the overall performance of the system.

[0071] The identification unit identifies the calls of specific wild birds from the sounds analyzed by the analysis unit. The identification unit can identify the calls of specific wild birds using methods such as machine learning algorithms and pattern matching. Specifically, it uses machine learning algorithms to extract the characteristics of wild bird calls from the analyzed sound data and identifies the calls of specific wild birds based on those characteristics. In pattern matching, the analyzed sound data is compared with a database of known wild bird calls and matching patterns are identified. The identification unit can also identify the calls of specific wild birds using generative AI. The generative AI analyzes sound data using a deep learning model and identifies the characteristics of wild bird calls with high accuracy. For example, the generative AI takes collected sound data as input, generates a spectrogram of wild bird calls, and analyzes the spectrogram to identify the wild bird calls. The generative AI learns from past wild bird call data and has the ability to identify the calls of specific wild birds. This allows the identification unit to quickly and accurately identify the calls of specific wild birds from the collected sound data and provide this information to the display unit. Furthermore, the identification unit can identify sound data in real time and instantly grasp the surrounding situation. As a result, the identification unit can identify sound data efficiently and effectively, improving the overall performance of the system.

[0072] The display unit provides the user with information identified by the identification unit. The display unit can provide information using methods such as text display, voice notification, and visual display. Specifically, it can list the identified species and number of wild birds and display them to the user in text format. Voice notification can inform the user of the identified species and number of wild birds audibly. Visual display graphically displays the identified species and number of wild birds, providing the user with information visually. The display unit can estimate the direction of bird calls through acoustic analysis and display it on a map. Specifically, it estimates the direction and distance of the sound source based on sound data collected by the collection unit and displays this information on a map. This allows the user to visually understand the direction from which the bird calls are coming. Furthermore, the display unit can collect user feedback and continuously improve the accuracy and effectiveness of the displayed content. For example, the user can provide feedback on the identified bird information, and the displayed content can be reviewed based on that feedback. The display unit can also reliably transmit information using multiple communication methods. For example, it uses not only smartphone notifications but also voice calls, SMS, and email to reliably deliver important information. This allows the display to provide information to the user quickly and reliably, making birdwatching easier and more enjoyable.

[0073] The display unit can list the types and numbers of identified wild birds. For example, the display unit can list the types and numbers of identified wild birds in text format. The display unit can also display the types and numbers of identified wild birds graphically. The display unit can also announce the types and numbers of identified wild birds via voice. This makes it easier for users to understand information about the wild birds in their surroundings by listing the types and numbers of identified birds.

[0074] The display unit can estimate the direction of a bird's call through acoustic analysis and display it on a map. For example, the display unit can estimate the direction of a bird's call through acoustic analysis and display it on a map with an arrow. The display unit can also estimate the direction of a bird's call through acoustic analysis and display it on a map with an icon. The display unit can also estimate the direction of a bird's call through acoustic analysis and display it on a map using different colors. This allows users to more accurately determine the location of wild birds by displaying the direction of the bird's call on a map.

[0075] The identification unit may include a notification unit that pre-registers the calls of specific wild birds and notifies the user when those birds make a sound. For example, the identification unit can pre-register the calls of specific wild birds and provide an audio notification when those birds make a sound. The identification unit can also pre-register the calls of specific wild birds and provide a text notification when those birds make a sound. The identification unit can also pre-register the calls of specific wild birds and provide a visual notification when those birds make a sound. This makes it easier for users to find specific wild birds by pre-registering their calls and providing notifications when those birds make a sound.

[0076] The sound collection unit can collect ambient sounds using a microphone array connected to a smartphone or earphones. For example, the sound collection unit can collect ambient sounds using a microphone array connected to a smartphone. The sound collection unit can also collect ambient sounds using a microphone array connected to earphones. The sound collection unit can also collect ambient sounds using multiple microphones. This allows for efficient collection of ambient sounds by using a microphone array connected to a smartphone or earphones.

[0077] The sound collection unit can estimate the user's emotions and adjust the timing of sound collection based on the estimated emotions. For example, if the user is excited, the sound collection unit can start collecting sounds immediately. If the user is relaxed, the sound collection unit can also collect sounds at regular intervals. If the user is stressed, the sound collection unit can also temporarily stop collecting sounds. This allows for sound collection at a more appropriate time by adjusting the timing of sound collection according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0078] The sound collection unit can analyze ambient sounds and prioritize the collection of sounds in specific frequency bands. For example, it can analyze ambient sounds in real time and prioritize the collection of frequency bands that contain bird calls. The sound collection unit can also remove human voices and car noises from ambient sounds and collect only bird calls. If the ambient noise level is high, the sound collection unit can also emphasize and collect sounds in specific frequency bands. This allows for the efficient collection of bird calls by prioritizing the collection of sounds in specific frequency bands.

[0079] The sound collection unit can automatically adjust the optimal microphone placement based on the user's location information when collecting sound. For example, the unit can automatically adjust the optimal microphone placement based on the user's location information. The unit can also update the microphone placement in real time as the user moves. The unit can also select the optimal sound collection point by considering the user's location information and the surrounding terrain information. This allows for efficient sound collection by automatically adjusting the optimal microphone placement based on the user's location information.

[0080] The sound collection unit can estimate the user's emotions and determine the priority of sounds to collect based on the estimated emotions. For example, if the user is excited, the sound collection unit can prioritize collecting the calls of a specific wild bird. If the user is relaxed, the sound collection unit can also collect a balanced mix of ambient sounds. If the user is stressed, the sound collection unit can also prioritize collecting quiet sounds. This allows for the collection of more appropriate sounds by prioritizing sounds according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0081] The sound collection unit can select the target for collection by referring to the user's past birdwatching history. For example, the collection unit can prioritize collecting the calls of wild birds that the user has observed in the past. The collection unit can also collect the calls of specific wild birds from the user's past birdwatching history. The collection unit can also collect the calls of wild birds observed in places the user has visited in the past. This allows for the collection of more appropriate sounds by referring to the user's past birdwatching history.

[0082] The sound collection unit can adjust its collection method by considering the surrounding weather information. For example, in rainy weather, the unit can remove the sound of rain and collect bird songs. In sunny weather, the unit can remove the sound of wind and collect bird songs. On snowy days, the unit can remove the sound of snow and collect bird songs. By adjusting the collection method by considering the surrounding weather information, more appropriate sounds can be collected.

[0083] The analysis unit can estimate the user's emotions and adjust the analysis algorithm based on the estimated emotions. For example, if the user is relaxed, the analysis unit can apply an algorithm that performs a detailed analysis. If the user is in a hurry, the analysis unit can also apply an algorithm that performs a rapid analysis. If the user is excited, the analysis unit can also apply an algorithm that provides visually stimulating analysis results. In this way, by adjusting the analysis algorithm according to the user's emotions, more appropriate analysis results can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0084] The analysis unit can be equipped with a filtering function that automatically removes noise from the collected sound during analysis. For example, the analysis unit can be equipped with a filtering function that automatically removes wind noise from the collected sound. The analysis unit can also be equipped with a filtering function that automatically removes human voices from the collected sound. The analysis unit can also be equipped with a filtering function that automatically removes car noise from the collected sound. This improves the accuracy of the analysis by automatically removing noise from the collected sound.

[0085] The analysis unit can simultaneously analyze the calls of different wild birds and identify multiple calls. For example, the analysis unit can simultaneously analyze the calls of different wild birds and identify them by species. The analysis unit can also simultaneously analyze the calls of multiple wild birds and identify the intensity of the calls. The analysis unit can also simultaneously analyze the calls of different wild birds and identify the direction of the calls. In this way, by simultaneously analyzing the calls of different wild birds, multiple calls can be identified.

[0086] The analysis unit can estimate the user's emotions and adjust the display method of the analysis results based on the estimated emotions. For example, if the user is nervous, the analysis unit can provide a simple and highly visible display method. If the user is relaxed, the analysis unit can also provide a display method that includes detailed information. If the user is in a hurry, the analysis unit can also provide a display method that gets straight to the point. By adjusting the display method of the analysis results according to the user's emotions, a more appropriate display becomes possible. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0087] The analysis unit can improve analysis accuracy by considering the time-of-day information of the collected sounds during analysis. For example, the analysis unit can improve analysis accuracy by considering the activity times of wild birds based on the time-of-day information of the collected sounds. The analysis unit can also prioritize the analysis of wild bird calls that occur during specific time periods based on the time-of-day information of the collected sounds. The analysis unit can also analyze nocturnal wild bird calls based on the time-of-day information of the collected sounds. In this way, the analysis accuracy is improved by considering the time-of-day information of the collected sounds.

[0088] The analysis unit can supplement the analysis results by referring to the geographical information of the collected sounds during the analysis. For example, the analysis unit can prioritize the analysis of bird calls inhabiting a specific area based on the geographical information of the collected sounds. The analysis unit can also identify bird calls for each region based on the geographical information of the collected sounds. The analysis unit can also analyze bird calls for a specific region based on the geographical information of the collected sounds. In this way, the analysis results can be supplemented by referring to the geographical information of the collected sounds.

[0089] The identification unit can estimate the user's emotions and adjust the identification algorithm based on the estimated emotions. For example, if the user is relaxed, the identification unit can apply an algorithm that provides detailed identification. If the user is in a hurry, the identification unit can also apply an algorithm that provides rapid identification. If the user is excited, the identification unit can also apply an algorithm that provides visually stimulating identification results. In this way, by adjusting the identification algorithm according to the user's emotions, more appropriate identification results can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0090] The identification unit can improve its identification accuracy by learning the call patterns of specific wild birds during identification. For example, the identification unit can improve its identification accuracy by learning the call patterns of specific wild birds. The identification unit can also improve its identification accuracy by learning the call patterns of multiple wild birds. The identification unit can improve its identification accuracy by learning the call patterns of wild birds. As a result, identification accuracy is improved by learning the call patterns of specific wild birds.

[0091] The identification unit can simultaneously identify the calls of multiple wild birds and classify them by species. For example, the identification unit can simultaneously identify the calls of multiple wild birds and classify them by species. The identification unit can also simultaneously identify the calls of multiple wild birds and classify their intensity. Furthermore, the identification unit can simultaneously identify the calls of multiple wild birds and classify their direction. This allows for the simultaneous identification of the calls of multiple wild birds, enabling classification by species.

[0092] The identification unit can estimate the user's emotions and adjust the display method of the identification results based on the estimated emotions. For example, if the user is nervous, the identification unit can provide a simple and highly visible display method. If the user is relaxed, the identification unit can also provide a display method that includes detailed information. If the user is in a hurry, the identification unit can also provide a display method that gets straight to the point. By adjusting the display method of the identification results according to the user's emotions, a more appropriate display becomes possible. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0093] The identification unit can improve its identification accuracy by considering the frequency characteristics of the collected sound during identification. For example, the identification unit can identify the calls of wild birds based on the frequency characteristics of the collected sound. The identification unit can also identify the calls of specific wild birds based on the frequency characteristics of the collected sound. The identification unit can also identify the calls of multiple wild birds based on the frequency characteristics of the collected sound. In this way, the identification accuracy is improved by considering the frequency characteristics of the collected sound.

[0094] The identification unit can supplement the identification result by referring to the time-of-day information of the collected sounds during the identification process. For example, the identification unit can identify the calls of wild birds that sing during a specific time period based on the collected time-of-day information of the sounds. The identification unit can also identify the calls of wild birds that sing at night based on the collected time-of-day information of the sounds. The identification unit can also supplement the identification result by considering the activity times of wild birds based on the collected time-of-day information of the sounds. In this way, the identification result can be supplemented by referring to the time-of-day information of the sounds.

[0095] The display unit can estimate the user's emotions and adjust the displayed content based on the estimated emotions. For example, if the user is nervous, the display unit can provide simple and highly visible content. If the user is relaxed, the display unit can also provide content that includes detailed information. If the user is in a hurry, the display unit can also provide content that gets straight to the point. This allows for more appropriate display by adjusting the displayed content according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0096] The display unit can update the direction of the identified bird's call in real time while displaying information. For example, if the user moves their smartphone, the display unit can update the direction of the identified bird's call in real time. The display unit can also update the displayed content in real time if the direction of the bird's call changes. The display unit can also update the direction of the identified bird's call in real time based on the user's location information. This allows the user to accurately determine the bird's location by updating the direction of the identified bird's call in real time.

[0097] The display unit can visually show the intensity of the identified bird's call when it is displayed. For example, the display unit can visually show the intensity of the identified bird's call using shades of color. The display unit can also visually show the intensity of the identified bird's call using a bar graph. The display unit can also visually show the intensity of the identified bird's call using numerical values. This allows users to intuitively understand the intensity of the bird's call by visually displaying the intensity of the identified bird's call.

[0098] The display unit can estimate the user's emotions and determine the display priority based on the estimated emotions. For example, if the user is excited, the display unit can prioritize displaying information about a specific wild bird. If the user is relaxed, the display unit can also display a balanced mix of information. If the user is stressed, the display unit can prioritize displaying information about a calm environment. This allows for the provision of more appropriate information by prioritizing the display according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0099] The display unit can select the optimal display method while considering the user's location information. For example, the display unit can automatically select the optimal display method based on the user's location information. The display unit can also update the display method in real time as the user moves. The display unit can also select the optimal display method considering the user's location information and surrounding terrain information. This allows for the provision of more appropriate information by selecting the optimal display method while considering the user's location information.

[0100] The display unit can show a history of the calls of identified wild birds when it is displayed. For example, the display unit can display a history of identified wild bird calls in chronological order. The display unit can also display a history of identified wild bird calls on a map. The display unit can also display a history of identified wild bird calls in list format. This allows users to review past observation results by displaying a history of identified wild bird calls.

[0101] The notification unit can estimate the user's emotions and adjust the timing of notifications based on the estimated emotions. For example, if the user is relaxed, the notification unit can reduce the frequency of notifications. If the user is excited, the notification unit can increase the frequency of notifications. If the user is stressed, the notification unit can temporarily stop notifications. This allows for more appropriate timing of notifications by adjusting the timing according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0102] The notification unit can apply different notification methods depending on the type of bird call identified. For example, if the call of a specific bird is identified, the notification unit can provide an audio notification. If the calls of multiple birds are identified, the notification unit can also provide a vibration notification. If the call of a rare bird is identified, the notification unit can also provide a pop-up notification. This allows the system to provide users with appropriate notifications by applying different notification methods depending on the type of bird call identified.

[0103] The notification unit can select the most appropriate notification method by referring to the user's past notification history when sending a notification. For example, the notification unit can select the most appropriate notification method based on the user's past notification history. The notification unit can also prioritize the application of notification methods that the user has previously preferred. The notification unit can also select a notification method suitable for a specific time period from the user's past notification history. This allows for the selection of a more appropriate notification method by referring to the user's past notification history.

[0104] The notification unit can estimate the user's emotions and adjust the notification content based on the estimated emotions. For example, if the user is stressed, the notification unit can provide a simple and highly visible notification. If the user is relaxed, the notification unit can provide a notification containing detailed information. If the user is in a hurry, the notification unit can provide a notification that gets straight to the point. This allows for more appropriate notifications by adjusting the notification content according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0105] The notification unit can select the optimal notification method by considering the user's location information when sending a notification. For example, the notification unit can automatically select the optimal notification method based on the user's location information. The notification unit can also update the notification method in real time when the user moves. The notification unit can also select the optimal notification method by considering the user's location information and the surrounding terrain information. This allows for more appropriate notifications by selecting the optimal notification method while considering the user's location information.

[0106] The notification unit can supplement notification content by referring to the history of the identified bird's calls when sending a notification. For example, the notification unit can supplement notification content based on the history of the identified bird's calls. The notification unit can also notify information about a specific bird based on the history of the identified bird's calls. The notification unit can also customize notification content based on the history of the identified bird's calls. This allows for the provision of more appropriate notification content by referring to the history of the identified bird's calls.

[0107] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0108] The Triresearch system can estimate the user's emotions and adjust the accuracy of bird call identification based on those emotions. For example, if the user is excited, the identification unit can perform a more detailed analysis and identify multiple bird calls simultaneously. If the user is relaxed, the identification unit can perform a simpler analysis and identify only the main bird calls. If the user is stressed, the identification unit can prioritize identifying specific bird calls and filter out other sounds. This allows the system to provide more relevant information by adjusting the identification accuracy according to the user's emotions.

[0109] The Triresearch system can select the optimal sound collection point based on the user's location information in its collection unit. For example, if the user is in a mountainous area, the collection unit can prioritize collecting the calls of wild birds specific to mountainous regions. If the user is by a lake, the collection unit can also prioritize collecting the calls of wild birds that inhabit waterside areas. If the user is in an urban area, the collection unit can perform noise filtering suitable for urban environments and efficiently collect wild bird calls. In this way, by selecting the optimal sound collection point based on the user's location information, sound can be collected efficiently.

[0110] The Trisearch system can improve analysis accuracy by considering the time-of-day information of collected sounds in its analysis unit. For example, it can improve analysis accuracy by considering the activity times of wild birds based on the time-of-day information of collected sounds. It can also prioritize the analysis of wild bird calls that occur during specific time periods. It can also analyze wild bird calls that occur at night. In this way, the analysis accuracy is improved by considering the time-of-day information of collected sounds.

[0111] The Trisearch system, in its identification unit, can estimate the user's emotions and adjust the display method of the identification results based on the estimated emotions. For example, if the user is nervous, a simple and highly visible display method can be provided. If the user is relaxed, a display method including detailed information can be provided. If the user is in a hurry, a display method that gets straight to the point can be provided. In this way, by adjusting the display method of the identification results according to the user's emotions, a more appropriate display becomes possible.

[0112] The Trisearch system analyzes ambient noise in its collection unit and can prioritize the collection of sounds in specific frequency bands. For example, it can analyze ambient noise in real time and prioritize the collection of frequency bands containing bird calls. It can also remove human voices and car noises from the ambient noise and collect only bird calls. If the ambient noise level is high, it can also emphasize and collect sounds in specific frequency bands. This allows for the efficient collection of bird calls by prioritizing the collection of sounds in specific frequency bands.

[0113] The Trisearch system can have a filtering function added to its analysis unit that automatically removes noise from the collected sound. For example, a filtering function can be added to automatically remove wind noise from the collected sound. A filtering function can also be added to automatically remove human voices from the collected sound. A filtering function can also be added to automatically remove car noise from the collected sound. This improves the accuracy of the analysis by automatically removing noise from the collected sound.

[0114] The Trisearch system can estimate the user's emotions in its identification unit and adjust the identification algorithm based on those emotions. For example, if the user is relaxed, an algorithm that provides detailed identification can be applied. If the user is in a hurry, an algorithm that provides rapid identification can be applied. If the user is excited, an algorithm that provides visually stimulating identification results can be applied. In this way, by adjusting the identification algorithm according to the user's emotions, more appropriate identification results can be provided.

[0115] The Trisearch system, in its sound collection unit, can automatically adjust the optimal microphone placement based on the user's location information during sound collection. For example, it can automatically adjust the optimal microphone placement based on the user's location. It can also update the microphone placement in real time as the user moves. It can also select the optimal sound collection point by considering the user's location information and the surrounding terrain information. As a result, by automatically adjusting the optimal microphone placement based on the user's location information, sound can be collected efficiently.

[0116] The Trisearch system, in its analysis unit, can estimate the user's emotions and adjust the analysis algorithm based on those emotions. For example, if the user is relaxed, a detailed analysis algorithm can be applied. If the user is in a hurry, a rapid analysis algorithm can be applied. If the user is excited, an algorithm that provides visually stimulating analysis results can be applied. In this way, by adjusting the analysis algorithm according to the user's emotions, more appropriate analysis results can be provided.

[0117] The birdwatching system can display a history of identified bird calls on its display unit. For example, it can display a history of identified bird calls in chronological order. It can also display a history of identified bird calls on a map. It can also display a history of identified bird calls in a list format. This allows users to review past observation results by viewing a history of identified bird calls.

[0118] The following briefly describes the processing flow for example form 2.

[0119] Step 1: The collection unit collects ambient sounds. The collection unit can collect ambient sounds using, for example, a microphone array connected to a smartphone or earphones. The collection unit can collect natural sounds, artificial sounds, sounds in specific frequency bands, etc. Step 2: The analysis unit analyzes the sound collected by the collection unit. The analysis unit can analyze the sound using methods such as speech recognition algorithms, frequency analysis, and time-domain analysis. The analysis unit can also analyze the collected sound using generative AI. Step 3: The identification unit identifies the calls of specific wild birds from the sounds analyzed by the analysis unit. The identification unit can identify the calls of specific wild birds using methods such as machine learning algorithms and pattern matching. The identification unit can also identify the calls of specific wild birds using generative AI. Step 4: The display unit provides the user with the information identified by the identification unit. The display unit can provide information using methods such as text display, audio notification, and visual display. The display unit can list the types and numbers of identified wild birds. The display unit can estimate the direction of the bird calls by acoustic analysis and display it on a map.

[0120] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0121] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0122] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0123] Each of the multiple elements described above, including the collection unit, analysis unit, identification unit, and display unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the collection unit collects ambient sound using the microphone array of the smart device 14. The analysis unit is implemented in the identification processing unit 290 of the data processing unit 12 and analyzes the collected sound. The identification unit is implemented in the identification processing unit 290 of the data processing unit 12 and identifies the calls of specific wild birds from the analyzed sound. The display unit is implemented in the control unit 46A of the smart device 14 and provides the identified information to the user. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0124] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0125] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0126] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0127] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0128] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0129] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0130] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0131] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0132] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0133] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0134] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0135] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0136] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0137] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0138] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0139] Each of the multiple elements described above, including the collection unit, analysis unit, identification unit, and display unit, is implemented, for example, in at least one of the smart glasses 214 and the data processing unit 12. For example, the collection unit collects ambient sound using the microphone array of the smart glasses 214. The analysis unit is implemented, for example, in the identification processing unit 290 of the data processing unit 12, and analyzes the collected sound. The identification unit is implemented, for example, in the identification processing unit 290 of the data processing unit 12, and identifies the call of a specific wild bird from the analyzed sound. The display unit is implemented, for example, in the control unit 46A of the smart glasses 214, and provides the identified information to the user. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.

[0140] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0141] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0142] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0143] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0144] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0145] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0146] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0147] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0148] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0149] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0150] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0151] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0152] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0153] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0154] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0155] Each of the multiple elements described above, including the collection unit, analysis unit, identification unit, and display unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the collection unit collects ambient sound using the microphone array of the headset terminal 314. The analysis unit is implemented in the identification processing unit 290 of the data processing unit 12 and analyzes the collected sound. The identification unit is implemented in the identification processing unit 290 of the data processing unit 12 and identifies the call of a specific wild bird from the analyzed sound. The display unit is implemented in the control unit 46A of the headset terminal 314 and provides the identified information to the user. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0156] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0157] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0158] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0159] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0160] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0161] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0162] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0163] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0164] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0165] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0166] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0167] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0168] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0169] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0170] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0171] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0172] Each of the multiple elements described above, including the collection unit, analysis unit, identification unit, and display unit, is implemented in, for example, at least one of the robot 414 and the data processing unit 12. For example, the collection unit collects ambient sounds using the microphone array of the robot 414. The analysis unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, and analyzes the collected sounds. The identification unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, and identifies the calls of specific wild birds from the analyzed sounds. The display unit is implemented, for example, by the control unit 46A of the robot 414, and provides the identified information to the user. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.

[0173] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0174] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0175] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0176] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0177] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0178] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0179] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0180] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0181] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0182] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0183] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0184] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0185] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0186] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0187] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0188] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0189] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0190] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0191] (Note 1) A sound collection unit that collects ambient sounds, An analysis unit analyzes the sound collected by the aforementioned collection unit, An identification unit that identifies the calls of specific wild birds from the sounds analyzed by the aforementioned analysis unit, The system includes a display unit that provides the user with the information identified by the identification unit. A system characterized by the following features. (Note 2) The aforementioned display unit is List the species and number of wild birds identified. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned display unit is The direction of the bird's call is estimated using acoustic analysis and displayed on a map. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned identification unit is It includes a notification unit that allows users to pre-register the calls of specific wild birds and sends a notification when those birds make a sound. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned collection unit is It collects ambient sounds using a microphone array connected to a smartphone or earphones. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned collection unit is It estimates the user's emotions and adjusts the timing of sound collection based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned collection unit is It analyzes ambient sounds and prioritizes collecting sounds within a specific frequency range. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned collection unit is During sound collection, the system automatically adjusts the optimal microphone placement based on the user's location information. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned collection unit is It estimates the user's emotions and determines the priority of sounds to collect based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned collection unit is When collecting sounds, the system selects the target for collection by referring to the user's past birdwatching history. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned collection unit is When collecting sound, the collection method is adjusted considering the surrounding weather information. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned analysis unit, It estimates the user's emotions and adjusts the analysis algorithm based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned analysis unit, During analysis, a filtering function will be added to automatically remove noise from the collected audio. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned analysis unit, During analysis, the calls of different wild birds are analyzed simultaneously, and multiple calls are identified. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned analysis unit, It estimates the user's emotions and adjusts how the analysis results are displayed based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned analysis unit, During analysis, the timing information of the collected sound is taken into consideration to improve the accuracy of the analysis. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned analysis unit, During analysis, the geographical information of the collected sounds is referenced to supplement the analysis results. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned identification unit is The system estimates the user's emotions and adjusts the identification algorithm based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned identification unit is During identification, the system learns the call patterns of specific wild birds to improve identification accuracy. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned identification unit is During identification, the calls of multiple wild birds are identified simultaneously and classified by species. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned identification unit is It estimates the user's emotions and adjusts how the identification results are displayed based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned identification unit is During identification, the frequency characteristics of the collected sound are taken into consideration to improve identification accuracy. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned identification unit is During identification, the time-series information of the collected sounds is used to supplement the identification result. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned display unit is It estimates the user's emotions and adjusts the displayed content based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned display unit is When displayed, the direction of the identified bird's call is updated in real time. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned display unit is When displayed, the intensity of the identified bird calls is visually shown. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned display unit is It estimates the user's emotions and determines the display priority based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned display unit is When displaying content, the system selects the optimal display method by considering the user's location information. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned display unit is When displayed, the history of the identified bird calls will be shown. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned notification unit, It estimates the user's emotions and adjusts the timing of notifications based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 31) The aforementioned notification unit, When a notification is sent, different notification methods are applied depending on the type of bird call that was identified. The system described in Appendix 1, characterized by the features described herein. (Note 32) The aforementioned notification unit, When sending a notification, the system will refer to the user's past notification history to select the most suitable notification method. The system described in Appendix 1, characterized by the features described herein. (Note 33) The aforementioned notification unit, It estimates the user's emotions and adjusts the notification content based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 34) The aforementioned notification unit, When sending a notification, the system will select the most suitable notification method, taking into account the user's location. The system described in Appendix 1, characterized by the features described herein. (Note 35) The aforementioned notification unit, When a notification is sent, the notification content is supplemented by referring to the history of the identified wild bird's calls. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]

[0192] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A sound collection unit that collects ambient sounds, An analysis unit analyzes the sound collected by the aforementioned collection unit, An identification unit that identifies the calls of specific wild birds from the sounds analyzed by the aforementioned analysis unit, The system includes a display unit that provides the user with the information identified by the identification unit. A system characterized by the following features.

2. The aforementioned display unit is List the species and number of wild birds identified. The system according to feature 1.

3. The aforementioned display unit is The direction of the bird's call is estimated using acoustic analysis and displayed on a map. The system according to feature 1.

4. The aforementioned identification unit is It includes a notification unit that allows users to pre-register the calls of specific wild birds and sends a notification when those birds make a sound. The system according to feature 1.

5. The aforementioned collection unit is It collects ambient sounds using a microphone array connected to a smartphone or earphones. The system according to feature 1.

6. The aforementioned collection unit is It estimates the user's emotions and adjusts the timing of sound collection based on the estimated user emotions. The system according to feature 1.

7. The aforementioned collection unit is It analyzes ambient sounds and prioritizes collecting sounds within a specific frequency range. The system according to feature 1.

8. The aforementioned collection unit is During sound collection, the system automatically adjusts the optimal microphone placement based on the user's location information. The system according to feature 1.

9. The aforementioned collection unit is It estimates the user's emotions and determines the priority of sounds to collect based on the estimated user emotions. The system according to feature 1.

10. The aforementioned collection unit is When collecting sounds, the system selects the target for collection by referring to the user's past birdwatching history. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A