system

The system efficiently sorts and makes photo information accessible by converting it into text, enabling easy access and organization using generation AI.

JP2026045616APending Publication Date: 2026-03-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Photos stored in HDDs or the cloud are not efficiently sorted and information is not provided in an easily accessible form.

Method used

A system comprising an upload unit, an analysis unit, and an organization unit that converts photo content into text and organizes it in a user-friendly format using generation AI for easy access.

Benefits of technology

The system allows users to easily access and organize photo information by converting it into text, making it searchable and accessible at any time through smartphones or PCs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026045616000001_ABST
    Figure 2026045616000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to convert the content of a photograph into text and organize the information in a way that is easily accessible to the user. [Solution] The system according to the embodiment comprises an upload unit, an analysis unit, and an organization unit. The upload unit allows the user to upload photos stored in the cloud or on an HDD to the system. The analysis unit analyzes the content of the photos uploaded by the upload unit and converts information about people, places, and events in the photos into text. The organization unit organizes the information converted into text by the analysis unit in a way that is easily accessible to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional technology, photos stored in HDDs or the cloud have not been sufficiently sorted efficiently and information has not been provided in a form that can be easily accessed, leaving room for improvement.

[0005] The system according to the embodiment aims to textify the content of photos and organize information in a form that can be easily accessed by users.

Means for Solving the Problems

[0006] The system according to this embodiment comprises an upload unit, an analysis unit, and an organization unit. The upload unit allows users to upload photos stored in the cloud or on an HDD to the system. The analysis unit analyzes the content of the photos uploaded by the upload unit and converts information about people, places, and events depicted in the photos into text. The organization unit organizes the information converted into text by the analysis unit in a format that is easily accessible to the user. [Effects of the Invention]

[0007] The system according to this embodiment can convert the content of a photograph into text and organize the information in a way that is easily accessible to the user. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The photo organization system according to an embodiment of the present invention is a system that uses a generation AI to convert the content of photos into text and organizes the information in a way that users can easily access at any time. In this photo organization system, the user uploads photos stored in the cloud or on an HDD to the system, the generation AI analyzes the content of the photos, and converts information such as people, places, and events depicted in the photos into text. For example, for a family photo, the content of the photo is specifically converted into text such as "A family photo taken at a beach during the summer holidays of 2023." The converted text information is organized in a way that users can easily access. For example, photos can be categorized by the date, location, and event they were taken for, allowing users to easily search for specific photos. Furthermore, the converted text information is accessible at any time from the user's smartphone or PC, allowing them to rediscover valuable moments and emotions captured in the photos. This system allows potentially forgotten, precious photos to be rediscovered in words, making them easily revisited at any time. For example, if a user wants to search for photos related to a specific event or place, the system automatically displays the relevant photos, allowing the user to easily review them. Thus, the photo organization system allows users to easily convert, organize, and access the content of their photos into text.

[0029] The photo organization system according to this embodiment comprises an upload unit, an analysis unit, and an organization unit. The upload unit allows users to upload photos stored in the cloud or on an HDD to the system. Photos stored in the cloud or on an HDD include, but are not limited to, file formats such as JPEG, PNG, and RAW. The upload unit can upload photos using, for example, drag-and-drop or a file selection dialog. The analysis unit uses a generation AI to analyze the content of the photos uploaded by the upload unit and converts information about people, places, and events in the photos into text. The analysis unit uses, for example, image recognition technology to identify people, places, and events in the photos and converts them into text. Image recognition technology includes, for example, face recognition, object detection, and scene analysis. The analysis unit uses a generation AI to analyze the content of the photos and convert them into text. The generation AI is, for example, a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. The organization unit organizes the information converted into text by the analysis unit in a way that is easily accessible to the user. The organization unit classifies photos by date, location, and event, for example, allowing users to easily search for specific photos. The organization unit can organize information using folders, tags, and search functions. The organization unit makes the text-based information accessible at any time from the user's smartphone or PC. The organization unit allows access to the information using, for example, a smartphone app, a web browser, or a desktop app. As a result, the photo organization system according to this embodiment allows users to easily transcribe, organize, and access the content of photos as text.

[0030] The analysis unit can use image recognition technology to identify people, places, and events in a photograph and convert them into text. Image recognition technology includes, but is not limited to, facial recognition, object detection, and scene analysis. For example, the analysis unit can use facial recognition technology to identify people in a photograph and convert their names and relationships into text. The analysis unit can also use object detection technology to identify places and events in a photograph and convert them into text. For example, the analysis unit can identify buildings and landscapes in a photograph and convert the names of those places into text. The analysis unit can also use scene analysis technology to identify events in a photograph and convert them into text. For example, the analysis unit can analyze the actions of people in a photograph and the decorations in the background to identify the type of event. This allows for accurate text conversion of the photograph's content using image recognition technology. Some or all of the above-described processes in the analysis unit may be performed using, for example, a generative AI, or without a generative AI. For example, the analysis unit can input the photograph's image data into a generative AI, analyze the photograph's content using image recognition technology, and convert it into text.

[0031] The sorting unit can classify photos by date or location taken, or by event, making it easy for users to search for specific photos. Classification includes, but is not limited to, date, location, and type of event. The sorting unit can, for example, analyze the metadata of photos to identify the date and location taken. It can also analyze the content of photos to identify the type of event. For example, it can analyze the people and background decorations in a photo to identify the type of event. The sorting unit classifies photos based on the identified date, location, and type of event taken, making it easy for users to search for specific photos. For example, the sorting unit can organize photos into folders by date taken, making it easy for users to search for photos from a specific date. It can also tag photos by location, making it easy for users to search for photos from a specific location. Furthermore, the sorting unit can classify photos by event, making it easy for users to search for photos from a specific event. This makes photo classification easier and allows users to easily search for specific photos. Some or all of the above processing in the sorting unit may be performed, for example, using generative AI, or not using generative AI. For example, the sorting unit can input the metadata of a photograph into a generating AI, which can then identify and classify the date and location of the photograph.

[0032] The organization unit can make the digitized information accessible to users at any time from their smartphones or PCs. Accessibility includes, but is not limited to, smartphone apps, web browsers, and desktop applications. For example, the organization unit can make the digitized information accessible using a smartphone app. It can also make the digitized information accessible using a web browser. Furthermore, it can also make the digitized information accessible using a desktop application. For example, the organization unit allows users to access photo information anytime, anywhere through a smartphone app. It can also allow users to access photo information from any internet-connected device through a web browser. Furthermore, it can allow users to access photo information even in offline environments through a desktop application. This allows users to access photo information anytime, anywhere. Some or all of the above processing in the organization unit may be performed using, for example, a generative AI, or without a generative AI. For example, the organization unit can input the digitized information into a generative AI and organize it in a user-friendly format.

[0033] The analysis unit can identify people in a photograph using facial recognition technology and convert their names and relationships into text. Facial recognition technology includes, but is not limited to, deep learning-based facial recognition algorithms. The analysis unit can, for example, identify people in a photograph using a deep learning-based facial recognition algorithm. The analysis unit can also convert the names and relationships of the identified people into text. For example, the analysis unit can analyze the faces of people in a photograph and identify their names. The analysis unit can also analyze the relationships of people in a photograph and convert them into text. For example, the analysis unit can analyze the faces of people in a photograph and identify that they are members of a family. This allows for accurate text conversion of information about people in a photograph using facial recognition technology. Some or all of the above processing in the analysis unit may be performed using, for example, generative AI, or without generative AI. For example, the analysis unit can input the image data of a photograph into a generative AI, use facial recognition technology to identify people in the photograph, and convert them into text.

[0034] The organization unit can automatically display relevant photos when a user wants to search for photos related to a specific event or location. Relevant photos include, but are not limited to, matching metadata or similar content. The organization unit can, for example, analyze the metadata of photos to identify photos related to a specific event or location. It can also analyze the content of photos to identify photos related to a specific event or location. For example, the organization unit can analyze the metadata of photos to identify photos related to a specific event or location. It can also analyze the content of photos to identify photos related to a specific event or location. For example, the organization unit can analyze the people or background decorations in photos to identify photos related to a specific event or location. The organization unit automatically displays the identified photos related to the event or location, making it easy for the user to search. This allows the user to easily find photos related to a specific event or location. Some or all of the above processing in the organization unit may be performed, for example, using generative AI, or without generative AI. For example, the organization unit can input photo metadata into a generating AI to identify and display photos related to specific events or locations.

[0035] The upload unit can analyze the user's past upload history and select the optimal upload method. The optimal upload method includes, but is not limited to, factors such as upload speed and file size compression. For example, the upload unit may prioritize suggesting upload methods the user has frequently used in the past (e.g., Wi-Fi, mobile data). The upload unit can also send notifications during specific time periods if the user tends to upload during those times. Furthermore, the upload unit can suggest optimal upload settings based on the types of photos the user has previously uploaded. For example, if the user has previously uploaded using Wi-Fi, the upload unit will prioritize suggesting Wi-Fi. The upload unit can also send notifications during specific time periods if the user has previously uploaded during those times. Furthermore, the upload unit can suggest optimal upload settings based on the types of photos the user has previously uploaded. This allows users to upload photos in the most optimal way based on their past upload history. Some or all of the above processing in the upload unit may be performed using, for example, a generative AI, or without one. For example, the upload unit can input the user's upload history data into a generative AI to select the optimal upload method.

[0036] The upload unit can filter photos when they are uploaded based on the user's current projects or areas of interest. Filtering may include, but is not limited to, project tags or keywords related to areas of interest. For example, the upload unit may suggest that the user upload only photos related to their current projects. It can also prioritize uploading photos related to the user's areas of interest. Furthermore, if the user is interested in a particular theme, the upload unit can filter and upload photos related to that theme. This allows for the uploading of photos that align with the user's interests. Some or all of the above processing in the upload unit may be performed using, for example, a generative AI, or without one. For example, the upload unit can input the user's project data and area of ​​interest data into a generative AI and perform filtering.

[0037] The upload unit can prioritize uploading photos that are highly relevant to the user's geographic location when uploading photos. Geographic location information includes, but is not limited to, GPS data and location services. For example, if the user is in a specific location, the upload unit can prioritize uploading photos related to that location. The upload unit can also prioritize uploading photos from the user's travel destination if the user is traveling. Furthermore, if the user is at home, the upload unit can prioritize uploading photos from within the home. This allows for the prioritization of uploading photos that are highly relevant to the user's geographic location. Some or all of the above processing in the upload unit may be performed using, for example, a generative AI, or without a generative AI. For example, the upload unit can input the user's geographic location data into a generative AI, identify highly relevant photos, and upload them.

[0038] The upload unit can analyze a user's social media activity when uploading photos and upload relevant photos. Social media activity includes, but is not limited to, posts, the number of likes, and comments. For example, if a user has posted on social media about a particular event, the upload unit will prioritize uploading photos related to that event. The upload unit can also upload photos related to a specific hashtag if the user is using that hashtag. Furthermore, the upload unit can automatically select and suggest photos that the user wants to share on social media. This allows for the uploading of relevant photos based on the user's social media activity. Some or all of the above processing in the upload unit may be performed using, for example, a generative AI, or not. For example, the upload unit can input the user's social media activity data into a generative AI to identify and upload relevant photos.

[0039] The analysis unit can adjust the level of detail of the analysis based on the importance of the photographs during the analysis. The importance of a photograph includes, but is not limited to, factors such as frequency of shooting and user ratings. For example, the analysis unit can perform a detailed analysis of photographs of important events. It can also perform a simplified analysis of everyday photographs. Furthermore, it can perform a detailed analysis of photographs of particular interest to the user. This allows the level of detail of the analysis to be adjusted according to the importance of the photographs. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or without a generative AI. For example, the analysis unit can input photograph importance data into a generative AI to adjust the level of detail of the analysis.

[0040] The analysis unit can apply different analysis algorithms depending on the category of the photograph during analysis. Categories include, but are not limited to, landscapes, people, and events. For example, the analysis unit can apply a face recognition algorithm to photographs of people. It can also apply a location recognition algorithm to photographs of landscapes. Furthermore, it can apply an event recognition algorithm to photographs of events. This allows the analysis unit to apply the most suitable analysis algorithm depending on the category of the photograph. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or without a generative AI. For example, the analysis unit can input the category data of the photograph into a generative AI and apply different analysis algorithms.

[0041] The analysis unit can determine the priority of analysis based on when the photos were taken. The shooting date includes, but is not limited to, metadata timestamps and calendar information. For example, the analysis unit may prioritize the analysis of recently taken photos. It may also prioritize the analysis of photos taken during a specific event period. Furthermore, the analysis unit may prioritize the analysis of photos from a period of particular interest to the user. This allows the analysis priority to be determined based on when the photos were taken. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or without a generative AI. For example, the analysis unit can input the photo shooting date data into a generative AI to determine the analysis priority.

[0042] The analysis unit can adjust the order of analysis based on the relevance of the photos during the analysis. Relevance includes, but is not limited to, similarity of content and matching metadata. For example, the analysis unit can analyze photos related to the same event together. It can also analyze photos taken in the same location together. Furthermore, it can analyze photos of the same person together. This allows the order of analysis to be adjusted based on the relevance of the photos. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the analysis unit can input photo relevance data into a generative AI to adjust the order of analysis.

[0043] The sorting unit can analyze the user's past search history to select the optimal sorting method during sorting. The optimal sorting method may include, but is not limited to, the user's search history and usage frequency. For example, the sorting unit can classify based on keywords the user has frequently searched in the past. Furthermore, if the user tends to search during specific time periods, the sorting unit can prioritize sorting information related to those times. In addition, the sorting unit can suggest the optimal classification method based on the user's past search history. This allows the sorting unit to select the optimal sorting method based on the user's past search history. Some or all of the above processing in the sorting unit may be performed using, for example, a generative AI, or without one. For example, the sorting unit can input the user's search history data into a generative AI to select the optimal sorting method.

[0044] The sorting unit can customize the sorting method based on the user's current areas of interest during sorting. Areas of interest include, but are not limited to, survey results and past search history. For example, the sorting unit can prioritize sorting photos related to themes the user is currently interested in. It can also prioritize sorting photos related to the user's current projects. Furthermore, if the user is interested in a particular event, the sorting unit can prioritize sorting photos related to that event. This allows the sorting method to be customized based on the user's current areas of interest. Some or all of the above processing in the sorting unit may be performed, for example, using a generative AI, or not using a generative AI. For example, the sorting unit can input user area of ​​interest data into a generative AI to customize the sorting method.

[0045] The sorting unit can select the optimal sorting method based on the user's geographical location information during sorting. Geographical location information includes, but is not limited to, GPS data and location services. For example, if the user is in a specific location, the sorting unit can prioritize sorting photos related to that location. It can also prioritize sorting photos from the travel destination if the user is traveling. Furthermore, if the user is at home, the sorting unit can prioritize sorting photos from within the home. This allows the sorting unit to select the optimal sorting method based on the user's geographical location information. Some or all of the above processing in the sorting unit may be performed using, for example, a generative AI, or without a generative AI. For example, the sorting unit can input the user's geographical location data into a generative AI to select the optimal sorting method.

[0046] The organization unit can analyze the user's social media activity during organization and propose organization methods. Social media activity includes, but is not limited to, posts, the number of likes, and comments. For example, if a user posts on social media about a specific event, the organization unit will prioritize organizing photos related to that event. The organization unit can also organize photos related to a specific hashtag if the user uses that hashtag. Furthermore, the organization unit can automatically select photos that the user wants to share on social media and propose organization methods. This allows the organization unit to propose the optimal organization method based on the user's social media activity. Some or all of the above processing in the organization unit may be performed using, for example, a generative AI, or not. For example, the organization unit can input the user's social media activity data into a generative AI and propose the optimal organization method.

[0047] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0048] The search engine can analyze a user's past search history to provide optimal search results. For example, it can prioritize displaying relevant photos based on keywords the user has frequently searched in the past. It can also prioritize displaying photos related to specific time periods if the user tends to search during those times. Furthermore, the search engine can suggest optimal search filters based on the user's past search history. This improves the user's search experience and allows them to quickly find the photos they need.

[0049] The analysis unit can estimate the age and gender of people in a photograph and classify the photograph based on this estimated information. For example, it can classify photographs into categories such as photographs of children, photographs of adults, and photographs of the elderly. The analysis unit can also classify photographs into categories such as photographs of men and photographs of women based on the gender of the people in the photograph. Furthermore, the analysis unit can tag photographs based on age and gender, allowing users to easily search for photographs related to specific age groups or genders. This enables more detailed and efficient organization of photographs.

[0050] The analysis unit can analyze the clothing and accessories of people in photographs and classify them based on fashion style. For example, it can classify them into categories such as casual, formal, and sportswear. The analysis unit can also identify clothing from specific brands or designers and tag photographs accordingly. Furthermore, the analysis unit can analyze fashion styles according to season and event, making it easy for users to search for photographs related to specific styles. This enables useful photo organization for users interested in fashion.

[0051] The analysis unit can estimate the season and time of day of the scenery and buildings depicted in a photograph and classify the photograph based on that. For example, it can classify photographs by season, such as spring, summer, autumn, and winter. It can also classify them by time of day, such as morning, noon, evening, and night. Furthermore, the analysis unit can identify events and activities associated with specific seasons and time periods and tag photographs accordingly. This makes it easy for users to search for photographs related to specific seasons and time periods.

[0052] The analysis unit can identify animals and pets in photographs and classify them accordingly. For example, it can classify them by animal, such as dogs, cats, and birds. The analysis unit can also identify specific animal species or pet names and tag photographs based on that. Furthermore, the analysis unit can analyze the behavior and facial expressions of animals and pets and classify photographs based on that. This enables photo organization that is beneficial for users interested in pets and animals.

[0053] The following briefly describes the processing flow for example form 1.

[0054] Step 1: The upload section allows the user to upload photos stored in the cloud or on an HDD to the system. Photos stored in the cloud or on an HDD may include, but are not limited to, file formats such as JPEG, PNG, and RAW. The upload section allows users to upload photos using, for example, drag-and-drop or a file selection dialog. Step 2: The analysis unit uses a generation AI to analyze the content of the photos uploaded by the upload unit and converts information about people, places, and events in the photos into text. The analysis unit identifies people, places, and events in the photos using, for example, image recognition technology and converts them into text. Image recognition technology includes, but is not limited to, face recognition, object detection, and scene analysis. The generation AI includes, but is not limited to, text generation AI (e.g., LLM) and multimodal generation AI. Step 3: The organization unit organizes the information transcribed into text by the analysis unit in a way that is easily accessible to users. For example, the organization unit categorizes photos by date, location, and event, making it easy for users to find specific photos. The organization unit can organize information using folders, tags, and search functions. The organization unit makes the transcribed information accessible to users at any time from their smartphones or PCs. For example, the organization unit makes the information accessible using smartphone apps, web browsers, desktop apps, etc.

[0055] (Example of form 2) The photo organization system according to an embodiment of the present invention is a system that uses a generation AI to convert the content of photos into text and organizes the information in a way that users can easily access at any time. In this photo organization system, the user uploads photos stored in the cloud or on an HDD to the system, the generation AI analyzes the content of the photos, and converts information such as people, places, and events depicted in the photos into text. For example, for a family photo, the content of the photo is specifically converted into text such as "A family photo taken at a beach during the summer holidays of 2023." The converted text information is organized in a way that users can easily access. For example, photos can be categorized by the date, location, and event they were taken for, allowing users to easily search for specific photos. Furthermore, the converted text information is accessible at any time from the user's smartphone or PC, allowing them to rediscover valuable moments and emotions captured in the photos. This system allows potentially forgotten, precious photos to be rediscovered in words, making them easily revisited at any time. For example, if a user wants to search for photos related to a specific event or place, the system automatically displays the relevant photos, allowing the user to easily review them. Thus, the photo organization system allows users to easily convert, organize, and access the content of their photos into text.

[0056] The photo organization system according to this embodiment comprises an upload unit, an analysis unit, and an organization unit. The upload unit allows users to upload photos stored in the cloud or on an HDD to the system. Photos stored in the cloud or on an HDD include, but are not limited to, file formats such as JPEG, PNG, and RAW. The upload unit can upload photos using, for example, drag-and-drop or a file selection dialog. The analysis unit uses a generation AI to analyze the content of the photos uploaded by the upload unit and converts information about people, places, and events in the photos into text. The analysis unit uses, for example, image recognition technology to identify people, places, and events in the photos and converts them into text. Image recognition technology includes, for example, face recognition, object detection, and scene analysis. The analysis unit uses a generation AI to analyze the content of the photos and convert them into text. The generation AI is, for example, a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. The organization unit organizes the information converted into text by the analysis unit in a way that is easily accessible to the user. The organization unit classifies photos by date, location, and event, for example, allowing users to easily search for specific photos. The organization unit can organize information using folders, tags, and search functions. The organization unit makes the text-based information accessible at any time from the user's smartphone or PC. The organization unit allows access to the information using, for example, a smartphone app, a web browser, or a desktop app. As a result, the photo organization system according to this embodiment allows users to easily transcribe, organize, and access the content of photos as text.

[0057] The analysis unit can use image recognition technology to identify people, places, and events in a photograph and convert them into text. Image recognition technology includes, but is not limited to, facial recognition, object detection, and scene analysis. For example, the analysis unit can use facial recognition technology to identify people in a photograph and convert their names and relationships into text. The analysis unit can also use object detection technology to identify places and events in a photograph and convert them into text. For example, the analysis unit can identify buildings and landscapes in a photograph and convert the names of those places into text. The analysis unit can also use scene analysis technology to identify events in a photograph and convert them into text. For example, the analysis unit can analyze the actions of people in a photograph and the decorations in the background to identify the type of event. This allows for accurate text conversion of the photograph's content using image recognition technology. Some or all of the above-described processes in the analysis unit may be performed using, for example, a generative AI, or without a generative AI. For example, the analysis unit can input the photograph's image data into a generative AI, analyze the photograph's content using image recognition technology, and convert it into text.

[0058] The sorting unit can classify photos by date or location taken, or by event, making it easy for users to search for specific photos. Classification includes, but is not limited to, date, location, and type of event. The sorting unit can, for example, analyze the metadata of photos to identify the date and location taken. It can also analyze the content of photos to identify the type of event. For example, it can analyze the people and background decorations in a photo to identify the type of event. The sorting unit classifies photos based on the identified date, location, and type of event taken, making it easy for users to search for specific photos. For example, the sorting unit can organize photos into folders by date taken, making it easy for users to search for photos from a specific date. It can also tag photos by location, making it easy for users to search for photos from a specific location. Furthermore, the sorting unit can classify photos by event, making it easy for users to search for photos from a specific event. This makes photo classification easier and allows users to easily search for specific photos. Some or all of the above processing in the sorting unit may be performed, for example, using generative AI, or not using generative AI. For example, the sorting unit can input the metadata of a photograph into a generating AI, which can then identify and classify the date and location of the photograph.

[0059] The organization unit can make the digitized information accessible to users at any time from their smartphones or PCs. Accessibility includes, but is not limited to, smartphone apps, web browsers, and desktop applications. For example, the organization unit can make the digitized information accessible using a smartphone app. It can also make the digitized information accessible using a web browser. Furthermore, it can also make the digitized information accessible using a desktop application. For example, the organization unit allows users to access photo information anytime, anywhere through a smartphone app. It can also allow users to access photo information from any internet-connected device through a web browser. Furthermore, it can allow users to access photo information even in offline environments through a desktop application. This allows users to access photo information anytime, anywhere. Some or all of the above processing in the organization unit may be performed using, for example, a generative AI, or without a generative AI. For example, the organization unit can input the digitized information into a generative AI and organize it in a user-friendly format.

[0060] The analysis unit can identify people in a photograph using facial recognition technology and convert their names and relationships into text. Facial recognition technology includes, but is not limited to, deep learning-based facial recognition algorithms. The analysis unit can, for example, identify people in a photograph using a deep learning-based facial recognition algorithm. The analysis unit can also convert the names and relationships of the identified people into text. For example, the analysis unit can analyze the faces of people in a photograph and identify their names. The analysis unit can also analyze the relationships of people in a photograph and convert them into text. For example, the analysis unit can analyze the faces of people in a photograph and identify that they are members of a family. This allows for accurate text conversion of information about people in a photograph using facial recognition technology. Some or all of the above processing in the analysis unit may be performed using, for example, generative AI, or without generative AI. For example, the analysis unit can input the image data of a photograph into a generative AI, use facial recognition technology to identify people in the photograph, and convert them into text.

[0061] The organization unit can automatically display relevant photos when a user wants to search for photos related to a specific event or location. Relevant photos include, but are not limited to, matching metadata or similar content. The organization unit can, for example, analyze the metadata of photos to identify photos related to a specific event or location. It can also analyze the content of photos to identify photos related to a specific event or location. For example, the organization unit can analyze the metadata of photos to identify photos related to a specific event or location. It can also analyze the content of photos to identify photos related to a specific event or location. For example, the organization unit can analyze the people or background decorations in photos to identify photos related to a specific event or location. The organization unit automatically displays the identified photos related to the event or location, making it easy for the user to search. This allows the user to easily find photos related to a specific event or location. Some or all of the above processing in the organization unit may be performed, for example, using generative AI, or without generative AI. For example, the organization unit can input photo metadata into a generating AI to identify and display photos related to specific events or locations.

[0062] The upload unit can estimate the user's emotions and adjust the timing of photo uploads based on the estimated emotions. Emotion estimation includes, but is not limited to, facial recognition, voice analysis, and text analysis. For example, the upload unit can estimate the user's emotions using facial recognition technology. It can also estimate the user's emotions using voice analysis technology. Furthermore, it can estimate the user's emotions using text analysis technology. For example, the upload unit can capture the user's facial expressions with a camera and estimate their emotions using facial recognition technology. It can also record the user's voice and estimate their emotions using voice analysis technology. Furthermore, it can analyze the user's text messages and estimate their emotions using text analysis technology. The upload unit adjusts the timing of photo uploads based on the estimated emotions. For example, if the user is relaxed, it can send a notification encouraging them to upload photos. If the user is busy, it can suggest postponing the upload. Furthermore, if the user is experiencing an emotional moment, it can recommend uploading photos to record that moment. This allows users to upload photos at the optimal time according to their emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the upload unit may be performed using a generative AI, or not using a generative AI. For example, the upload unit can input the user's facial expression data into a generative AI and have the generative AI perform emotion estimation.

[0063] The upload unit can analyze the user's past upload history and select the optimal upload method. The optimal upload method includes, but is not limited to, factors such as upload speed and file size compression. For example, the upload unit may prioritize suggesting upload methods the user has frequently used in the past (e.g., Wi-Fi, mobile data). The upload unit can also send notifications during specific time periods if the user tends to upload during those times. Furthermore, the upload unit can suggest optimal upload settings based on the types of photos the user has previously uploaded. For example, if the user has previously uploaded using Wi-Fi, the upload unit will prioritize suggesting Wi-Fi. The upload unit can also send notifications during specific time periods if the user has previously uploaded during those times. Furthermore, the upload unit can suggest optimal upload settings based on the types of photos the user has previously uploaded. This allows users to upload photos in the most optimal way based on their past upload history. Some or all of the above processing in the upload unit may be performed using, for example, a generative AI, or without one. For example, the upload unit can input the user's upload history data into a generative AI to select the optimal upload method.

[0064] The upload unit can filter photos when they are uploaded based on the user's current projects or areas of interest. Filtering may include, but is not limited to, project tags or keywords related to areas of interest. For example, the upload unit may suggest that the user upload only photos related to their current projects. It can also prioritize uploading photos related to the user's areas of interest. Furthermore, if the user is interested in a particular theme, the upload unit can filter and upload photos related to that theme. This allows for the uploading of photos that align with the user's interests. Some or all of the above processing in the upload unit may be performed using, for example, a generative AI, or without one. For example, the upload unit can input the user's project data and area of ​​interest data into a generative AI and perform filtering.

[0065] The upload unit can estimate the user's emotions and determine the priority of photos to upload based on the estimated emotions. Prioritization may include, but is not limited to, the intensity of the emotion or the importance of the photo. For example, if the user is experiencing an emotional moment, the upload unit might prioritize uploading photos to record that moment. Conversely, if the user is relaxed, the upload unit might encourage uploading photos to reminisce about past memories. Furthermore, if the user is busy, the upload unit might postpone uploading less important photos. This allows the upload priority of photos to be determined according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. The generation AI may be a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the processing described above in the upload unit may be performed using a generation AI, or not using a generation AI. For example, the upload unit can input user sentiment data into a generation AI to determine the priority of photos to upload.

[0066] The upload unit can prioritize uploading photos that are highly relevant to the user's geographic location when uploading photos. Geographic location information includes, but is not limited to, GPS data and location services. For example, if the user is in a specific location, the upload unit can prioritize uploading photos related to that location. The upload unit can also prioritize uploading photos from the user's travel destination if the user is traveling. Furthermore, if the user is at home, the upload unit can prioritize uploading photos from within the home. This allows for the prioritization of uploading photos that are highly relevant to the user's geographic location. Some or all of the above processing in the upload unit may be performed using, for example, a generative AI, or without a generative AI. For example, the upload unit can input the user's geographic location data into a generative AI, identify highly relevant photos, and upload them.

[0067] The upload unit can analyze a user's social media activity when uploading photos and upload relevant photos. Social media activity includes, but is not limited to, posts, the number of likes, and comments. For example, if a user has posted on social media about a particular event, the upload unit will prioritize uploading photos related to that event. The upload unit can also upload photos related to a specific hashtag if the user is using that hashtag. Furthermore, the upload unit can automatically select and suggest photos that the user wants to share on social media. This allows for the uploading of relevant photos based on the user's social media activity. Some or all of the above processing in the upload unit may be performed using, for example, a generative AI, or not. For example, the upload unit can input the user's social media activity data into a generative AI to identify and upload relevant photos.

[0068] The analysis unit can estimate the user's emotions and adjust the presentation of the analysis based on the estimated emotions. The presentation of the analysis includes, but is not limited to, the length of the text, the level of detail, and the tone of the words used. For example, if the user is relaxed, the analysis unit will provide detailed analysis results. Conversely, if the user is in a hurry, the analysis unit can provide concise analysis results. Furthermore, if the user is experiencing an emotional moment, the analysis unit can highlight information related to that emotion. This allows the presentation of the analysis to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the analysis unit may be performed using, for example, a generative AI, or without using a generative AI. For example, the analysis unit can input user emotion data into a generative AI and adjust the method of representing the analysis.

[0069] The analysis unit can adjust the level of detail of the analysis based on the importance of the photographs during the analysis. The importance of a photograph includes, but is not limited to, factors such as frequency of shooting and user ratings. For example, the analysis unit can perform a detailed analysis of photographs of important events. It can also perform a simplified analysis of everyday photographs. Furthermore, it can perform a detailed analysis of photographs of particular interest to the user. This allows the level of detail of the analysis to be adjusted according to the importance of the photographs. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or without a generative AI. For example, the analysis unit can input photograph importance data into a generative AI to adjust the level of detail of the analysis.

[0070] The analysis unit can apply different analysis algorithms depending on the category of the photograph during analysis. Categories include, but are not limited to, landscapes, people, and events. For example, the analysis unit can apply a face recognition algorithm to photographs of people. It can also apply a location recognition algorithm to photographs of landscapes. Furthermore, it can apply an event recognition algorithm to photographs of events. This allows the analysis unit to apply the most suitable analysis algorithm depending on the category of the photograph. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or without a generative AI. For example, the analysis unit can input the category data of the photograph into a generative AI and apply different analysis algorithms.

[0071] The analysis unit can estimate the user's emotions and adjust the length of the analysis based on the estimated emotions. The length of the analysis may include, but is not limited to, the user's level of interest or the complexity of the photo's content. For example, if the user is in a hurry, the analysis unit provides a short, concise analysis. Conversely, if the user is relaxed, the analysis unit can provide a detailed analysis. Furthermore, if the user is experiencing an emotional moment, the analysis unit can highlight information related to that emotion. This allows the length of the analysis to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the analysis unit may be performed using, for example, a generating AI, or without using a generating AI. For example, the analysis unit can input user emotion data into the generating AI and adjust the length of the analysis.

[0072] The analysis unit can determine the priority of analysis based on when the photos were taken. The shooting date includes, but is not limited to, metadata timestamps and calendar information. For example, the analysis unit may prioritize the analysis of recently taken photos. It may also prioritize the analysis of photos taken during a specific event period. Furthermore, the analysis unit may prioritize the analysis of photos from a period of particular interest to the user. This allows the analysis priority to be determined based on when the photos were taken. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or without a generative AI. For example, the analysis unit can input the photo shooting date data into a generative AI to determine the analysis priority.

[0073] The analysis unit can adjust the order of analysis based on the relevance of the photos during the analysis. Relevance includes, but is not limited to, similarity of content and matching metadata. For example, the analysis unit can analyze photos related to the same event together. It can also analyze photos taken in the same location together. Furthermore, it can analyze photos of the same person together. This allows the order of analysis to be adjusted based on the relevance of the photos. Some or all of the above processing in the analysis unit may be performed using, for example, a generative AI, or not using a generative AI. For example, the analysis unit can input photo relevance data into a generative AI to adjust the order of analysis.

[0074] The organization unit can estimate the user's emotions and adjust the organization method based on the estimated emotions. Organization methods include, but are not limited to, folder organization, tagging, and search functions. For example, the organization unit can perform detailed classification when the user is relaxed. It can also perform concise classification when the user is in a hurry. Furthermore, if the user is experiencing an emotional moment, the organization unit can highlight information related to that emotion. This allows the organization method to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the organization unit may be performed using, for example, generative AI, or without generative AI. For example, the organization unit can input user emotion data into a generating AI and adjust the organization method.

[0075] The sorting unit can analyze the user's past search history to select the optimal sorting method during sorting. The optimal sorting method may include, but is not limited to, the user's search history and usage frequency. For example, the sorting unit can classify based on keywords the user has frequently searched in the past. Furthermore, if the user tends to search during specific time periods, the sorting unit can prioritize sorting information related to those times. In addition, the sorting unit can suggest the optimal classification method based on the user's past search history. This allows the sorting unit to select the optimal sorting method based on the user's past search history. Some or all of the above processing in the sorting unit may be performed using, for example, a generative AI, or without one. For example, the sorting unit can input the user's search history data into a generative AI to select the optimal sorting method.

[0076] The sorting unit can customize the sorting method based on the user's current areas of interest during sorting. Areas of interest include, but are not limited to, survey results and past search history. For example, the sorting unit can prioritize sorting photos related to themes the user is currently interested in. It can also prioritize sorting photos related to the user's current projects. Furthermore, if the user is interested in a particular event, the sorting unit can prioritize sorting photos related to that event. This allows the sorting method to be customized based on the user's current areas of interest. Some or all of the above processing in the sorting unit may be performed, for example, using a generative AI, or not using a generative AI. For example, the sorting unit can input user area of ​​interest data into a generative AI to customize the sorting method.

[0077] The sorting function can estimate the user's emotions and determine sorting priorities based on those estimated emotions. Sorting priorities include, but are not limited to, the intensity of the emotion or the importance of the photos. For example, if the user is experiencing an emotional moment, the sorting function will prioritize sorting photos related to that moment. Alternatively, if the user is relaxed, the sorting function can sort photos to reminisce about past memories. Furthermore, if the user is busy, the sorting function can postpone sorting less important photos. This allows the sorting priority to be determined according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the processing described above in the sorting unit may be performed using, for example, a generative AI, or without a generative AI. For example, the sorting unit can input user emotion data into a generative AI to determine the sorting priority.

[0078] The sorting unit can select the optimal sorting method based on the user's geographical location information during sorting. Geographical location information includes, but is not limited to, GPS data and location services. For example, if the user is in a specific location, the sorting unit can prioritize sorting photos related to that location. It can also prioritize sorting photos from the travel destination if the user is traveling. Furthermore, if the user is at home, the sorting unit can prioritize sorting photos from within the home. This allows the sorting unit to select the optimal sorting method based on the user's geographical location information. Some or all of the above processing in the sorting unit may be performed using, for example, a generative AI, or without a generative AI. For example, the sorting unit can input the user's geographical location data into a generative AI to select the optimal sorting method.

[0079] The organization unit can analyze the user's social media activity during organization and propose organization methods. Social media activity includes, but is not limited to, posts, the number of likes, and comments. For example, if a user posts on social media about a specific event, the organization unit will prioritize organizing photos related to that event. The organization unit can also organize photos related to a specific hashtag if the user uses that hashtag. Furthermore, the organization unit can automatically select photos that the user wants to share on social media and propose organization methods. This allows the organization unit to propose the optimal organization method based on the user's social media activity. Some or all of the above processing in the organization unit may be performed using, for example, a generative AI, or not. For example, the organization unit can input the user's social media activity data into a generative AI and propose the optimal organization method. === Hard Collateral 1-1 === Each of the multiple elements described above, including the upload unit, analysis unit, and organization unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the upload unit is implemented by the control unit 46A of the smart device 14, and the user uploads photos stored in the cloud or HDD to the system. The analysis unit is implemented by the identification processing unit 290 of the data processing unit 12, and analyzes the content of the photos using generation AI and converts it into text. The organization unit is implemented by the control unit 46A of the smart device 14, and organizes the text information in a way that is easily accessible to the user. === Hard Collateral 1-2 === Each of the multiple elements described above, including the upload unit, analysis unit, and organization unit, is implemented, for example, in at least one of the smart glasses 214 and the data processing unit 12. For example, the upload unit is implemented by the control unit 46A of the smart glasses 214, and the user uploads photos stored in the cloud or HDD to the system. The analysis unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, and analyzes the content of the photos using generation AI and converts it into text. The organization unit is implemented, for example, by the control unit 46A of the smart glasses 214, and organizes the text information in a way that is easily accessible to the user. === Hard Collateral 1-3 === Each of the multiple elements described above, including the upload unit, analysis unit, and organization unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the upload unit is implemented by the control unit 46A of the headset terminal 314, and the user uploads photos stored in the cloud or HDD to the system. The analysis unit is implemented by the identification processing unit 290 of the data processing unit 12, and the user analyzes the content of the photos using generation AI and converts it into text. The organization unit is implemented by the control unit 46A of the headset terminal 314, and the user organizes the text information in a format that is easily accessible to the user. === Hard Collateral 1-4 === Each of the multiple elements described above, including the upload unit, analysis unit, and organization unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the upload unit is implemented by the control unit 46A of the robot 414, and the user uploads photos stored in the cloud or on an HDD to the system. The analysis unit is implemented by the identification processing unit 290 of the data processing unit 12, and the user analyzes the content of the photos using a generation AI and converts it into text. The organization unit is implemented by the control unit 46A of the robot 414, and the user organizes the text information in a way that is easily accessible to the user.

[0080] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0081] The analysis unit can estimate the emotions of people in a photograph and tag the photograph based on the estimated emotions. For example, if a person in a photograph is smiling, it can be tagged "happy," and if they look sad, it can be tagged "sad." The analysis unit can also comprehensively judge the emotions of multiple people in a photograph and assign a tag that represents the overall atmosphere of the photograph. Furthermore, the analysis unit can classify photographs using emotion tags, making it easy for users to search for photographs related to specific emotions. This enhances the emotional value of photographs and allows users to organize their photographs based on their emotions.

[0082] The search engine can analyze a user's past search history to provide optimal search results. For example, it can prioritize displaying relevant photos based on keywords the user has frequently searched in the past. It can also prioritize displaying photos related to specific time periods if the user tends to search during those times. Furthermore, the search engine can suggest optimal search filters based on the user's past search history. This improves the user's search experience and allows them to quickly find the photos they need.

[0083] The analysis unit can estimate the age and gender of people in a photograph and classify the photograph based on this estimated information. For example, it can classify photographs into categories such as photographs of children, photographs of adults, and photographs of the elderly. The analysis unit can also classify photographs into categories such as photographs of men and photographs of women based on the gender of the people in the photograph. Furthermore, the analysis unit can tag photographs based on age and gender, allowing users to easily search for photographs related to specific age groups or genders. This enables more detailed and efficient organization of photographs.

[0084] The editing function can estimate the user's emotions and adjust how photos are displayed based on those estimates. For example, if the user is relaxed, photos can be displayed in a slideshow format. If the user is in a hurry, photos can be displayed in thumbnail format. Furthermore, if the user is experiencing an emotional moment, photos related to that emotion can be highlighted. This allows users to enjoy photos in the most optimal display method according to their emotions.

[0085] The analysis unit can analyze the clothing and accessories of people in photographs and classify them based on fashion style. For example, it can classify them into categories such as casual, formal, and sportswear. The analysis unit can also identify clothing from specific brands or designers and tag photographs accordingly. Furthermore, the analysis unit can analyze fashion styles according to season and event, making it easy for users to search for photographs related to specific styles. This enables useful photo organization for users interested in fashion.

[0086] The organization function can estimate the user's emotions and suggest a photo organization method based on those emotions. For example, if the user is experiencing an emotional moment, it will prioritize organizing photos related to that emotion. If the user is relaxed, it can suggest a detailed classification. Furthermore, if the user is busy, it can suggest a concise classification. This provides the optimal organization method according to the user's emotions, improving the efficiency of photo organization.

[0087] The analysis unit can estimate the season and time of day of the scenery and buildings depicted in a photograph and classify the photograph based on that. For example, it can classify photographs by season, such as spring, summer, autumn, and winter. It can also classify them by time of day, such as morning, noon, evening, and night. Furthermore, the analysis unit can identify events and activities associated with specific seasons and time periods and tag photographs accordingly. This makes it easy for users to search for photographs related to specific seasons and time periods.

[0088] The sorting function can estimate the user's emotions and adjust the display order of photos based on those estimates. For example, if the user is experiencing an emotional moment, photos related to that emotion will be displayed first. If the user is relaxed, older photos can be prioritized to allow them to reminisce about past memories. Furthermore, if the user is busy, photos of high importance can be prioritized. This allows users to enjoy photos in an optimal display order tailored to their emotions.

[0089] The analysis unit can identify animals and pets in photographs and classify them accordingly. For example, it can classify them by animal, such as dogs, cats, and birds. The analysis unit can also identify specific animal species or pet names and tag photographs based on that. Furthermore, the analysis unit can analyze the behavior and facial expressions of animals and pets and classify photographs based on that. This enables photo organization that is beneficial for users interested in pets and animals.

[0090] The sorting function can estimate the user's emotions and determine the priority of sorting photos based on those emotions. For example, if the user is experiencing an emotional moment, it will prioritize sorting photos related to that emotion. If the user is relaxed, it can sort photos to reminisce about past memories. Furthermore, if the user is busy, it can postpone sorting less important photos. This provides an optimal sorting priority according to the user's emotions, improving the efficiency of photo organization.

[0091] The following briefly describes the processing flow for example form 2.

[0092] Step 1: The upload section allows the user to upload photos stored in the cloud or on an HDD to the system. Photos stored in the cloud or on an HDD may include, but are not limited to, file formats such as JPEG, PNG, and RAW. The upload section allows users to upload photos using, for example, drag-and-drop or a file selection dialog. Step 2: The analysis unit uses a generation AI to analyze the content of the photos uploaded by the upload unit and converts information about people, places, and events in the photos into text. The analysis unit identifies people, places, and events in the photos using, for example, image recognition technology and converts them into text. Image recognition technology includes, but is not limited to, face recognition, object detection, and scene analysis. The generation AI includes, but is not limited to, text generation AI (e.g., LLM) and multimodal generation AI. Step 3: The organization unit organizes the information transcribed into text by the analysis unit in a way that is easily accessible to users. For example, the organization unit categorizes photos by date, location, and event, making it easy for users to find specific photos. The organization unit can organize information using folders, tags, and search functions. The organization unit makes the transcribed information accessible to users at any time from their smartphones or PCs. For example, the organization unit makes the information accessible using smartphone apps, web browsers, desktop apps, etc.

[0093] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0094] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (for example, still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or Naive Bayes, and can perform a variety of operations, but is not limited to these examples. Furthermore, AI may also be an AI agent. Also, when the operations described above are performed by AI, the operations may be performed partially or entirely by AI, but is not limited to these examples. Additionally, operations performed by AI, including generative AI, may be replaced by rule-based operations, and rule-based operations may be replaced by operations performed by AI, including generative AI.

[0095] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0096] The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0097] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0098] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0099] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0100] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0101] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0102] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0103] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0104] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0105] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0106] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0107] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0108] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0109] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0110] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0111] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0112] The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0113] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0114] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0115] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0116] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0117] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0118] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0119] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0120] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0121] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0122] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0123] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0124] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0125] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0126] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0127] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0128] The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0129] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0130] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0131] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0132] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0133] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0134] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0135] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0136] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0137] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0138] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0139] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0140] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0141] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0142] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0143] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0144] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0145] The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0146] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0147] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0148] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0149] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0150] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0151] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0152] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0153] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0154] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0155] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0156] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0157] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0158] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0159] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0160] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0161] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0162] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0163] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0164] [Explanation of symbols]

[0165] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. An upload unit that allows users to upload photos stored in the cloud or on an HDD to the system, The analysis unit analyzes the content of the photos uploaded by the aforementioned upload unit and converts information about the people, places, and events in the photos into text. The system includes an organizing unit that organizes the information converted into text by the analysis unit into a format that is easily accessible to the user. A system characterized by the following features.

2. The aforementioned analysis unit, Using image recognition technology, we identify people, places, and events in photographs and convert them into text. The system according to feature 1.

3. The aforementioned editing unit, The photos are categorized by date or location taken, or by event, making it easy for users to find specific photos. The system according to feature 1.

4. The aforementioned editing unit, This allows users to access digitized information anytime from their smartphones or PCs. The system according to feature 1.

5. The aforementioned analysis unit, Using facial recognition technology, we identify people in photographs and convert their names and relationships into text. The system according to feature 1.

6. The aforementioned editing unit, If a user wants to search for photos related to a specific event or location, the system will automatically display relevant photos. The system according to feature 1.

7. The aforementioned upload unit, It estimates the user's emotions and adjusts the timing of photo uploads based on those emotions. The system according to feature 1.

8. The aforementioned upload unit, Analyze the user's past upload history and select the optimal upload method. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A